
AI Data Quality Crisis: Highlights From the Webinar on the Creation of Trustworthy AI Models

Editor’s Note: Today’s post recaps IGI Global Scientific Publishing’s recent webinar, “AI’s Data Quality Crisis: Challenges and Solutions.” The full Free On-Demand Virtual Event can be viewed here.
Introduction
As organizations begin to adopt artificial intelligence tools, the conversations held regarding data quality have become part of the ethics of the Open Science movement. Large language models (LLMs), alongside other AI systems are starting to require more and more reliable information; the only kind that can train, tune, and develop them ethically.
This is where Open Science and high-quality Open Access research become especially important. Peer-reviewed research provides AI systems with access to authoritative knowledge that has been rigorously evaluated by subject matter experts and validated through established scholarly processes. Publishers such as IGI Global Scientific Publishing play an important role in this ecosystem through nearly 200 highly indexed journals that disseminate scientific discoveries across Business & Management, Education & Social Sciences, and Science, Technology, & Medicine (STM), and a forthcoming Encyclopedia of Modern Artificial Intelligence with some of the most modern research available on the subject of AI.
By making rigorously reviewed scholarship accessible through Open Access initiatives, publishers like IGI Global Scientific Publishing help ensure that both human researchers and AI systems can engage with knowledge grounded in evidence, methodology, and accountability. IGI Global Scientific Publishing is accepting open access journal submissions, offering open access transformed agreements, as well as looking for whitepaper authors to educate their diverse audience on some of the most important ideas in the STM category at this time.
As AI adoption accelerates, the distinction between vetted scholarly research and unverified online content becomes increasingly significant. While AI models can ingest information from countless sources, trusted research remains essential for building systems that are accurate, transparent, explainable, and worthy of public trust.
Rather, it is the need for higher-quality data, stronger governance frameworks, and greater confidence in the information entering AI pipelines. In many respects, Open Science provides one of the strongest foundations for achieving those goals. As a result, throughout IGI Global Scientific Publishing’s recent webinar, a panel discussed the challenges facing the AI community.
Featured panelists include:
- Darrell Gunter, Founder and CEO of Gunter Media Group and a recognized thought leader in scholarly communications and AI applications
- Vamsi Kasivajjala, CEO and Board Member of NeuroDiscovery AI
- Karen Medhat, Global Account Technical Leader at IBM
- Anders Hammarbäck, Co-Founder and CEO of RedPine AI
Each panelist answered a total of five questions, all pertaining to the importance of data scrutiny and consciousness when using AI. They shared relevant ideas as experts in their fields, who have been on the ground as the AI revolution has been unfolding in each of their respective careers. This additional perspective was necessary for understanding the heavy levels of nuance that come with the field of AI.
The discussion was moderated by Nick Newcomer, Senior Vice President of Sales & Marketing at IGI Global Scientific Publishing; Genevieve Robinson, Managing Director of Marketing; Savannah Pecknold, Open Science Communications Coordinator; and Jordan Hildebrand, e-Book Collection Specialist.
The panel explored five critical questions surrounding data quality, governance, accountability, and the growing importance of authoritative content in AI development. The sentiment shared throughout, is that though artificial intelligence is transforming research, healthcare, education, business, and nearly every other sector of society, one factor remains central to success: the quality of the underlying data.

The Root Causes of Poor AI Data Quality
Vamsi Kasivajjala argued that one of the biggest contributors to poor data quality is that many individuals (researchers or otherwise) have poor downstream cleaning. They do not have an EOP (end of process) cleaning that allows the query to come out ready to be received with verifiable and replicable information. It lacks care, and presents messy data that shows signs of hallucinations, fragmentation, and is missing provenance.
He also argued that since there is so much data, ownership and attribution often become unclear, making it increasingly difficult to trace information back to its origin. In his view, stronger filtering mechanisms and data management systems are essential for maintaining reliability.
On the other hand, Darrell Gunter discussed that the absence of data quality also comes from the inability for AI to understand semantic human context, which contributes to AI’s inability to develop reproducible data. He emphasized the ongoing challenge of AI understanding semantic human context, arguing that systems still struggle to fully interpret meaning and nuance in the way humans do. Even when the data is clean, if its hard to understand, results will still struggle.
Both Anders Hammarbäck and Karen Medhat agreed that the amount of data is not the problem. The reproducibility of an AI system requires new training, and a focus on quality. The moment we rely on more diverse sources of information to structure the data and retrieve it, the sooner organizations can evaluate data quality within the AI development lifecycle, in turn making it easier to prevent downstream issues.
The Impact of Data Quality on Trust and Performance
The conversation then shifted toward the outcomes organizations can expect when AI models are built using incomplete, inconsistent, or unreliable data.
Karen Medhat took the lead on this one stating that incomplete data as well fragmented and hallucinated data only hinder the implementation of this technology. From smart transportation systems to financial assistance with healthcare; these are just a few programs that need this clean data.
She also stated that trust in the data upon the output is also important when considering data valid. If you cannot trust the information you are getting on the output, it highlights areas where the information may need more cleaning.
Vamsi Kasivajjala agreed, saying that training is the most principal element to cleaning data. If given clean data, the outcomes are going to be much clearer. He also stressed the element of reciprocity, stating that it must be correct, clean, and reproducible to be considered dependable.
Darrell Gunter and Anders Hammarbäck then followed suit, concluding that we want our AI to outperform the quality of the data it consumes. Meaning the results that the models put out must be sound through attribution, data, and for Gunter, display reciprocity and semantic learning. If AI can learn the language more efficiently, then ultimately these systems will surpass the quality of the input using strong attribution practices and will inevitably demonstrate improved semantic learning capabilities for reciprocity.
AI can personalize learning at a large scale. AI systems can look at data on how students are performing, as it happens, and change the difficulty, speed, and content of lessons to suit each learner. This flexibility also helps to prevent students from becoming bored or overwhelmed and hence helps increase their engagement and retention. For example, AI-powered platforms can identify when a student is struggling with a subject and automatically suggest additional resources or alternative explanations, whereas more advanced learners are faced with challenging material that encourages them to keep trying until they succeed.
AI makes it possible to produce gamified lectures, simulations, and interactive material that turns passive learning into active exploration of digital environments, that make comparison to the real world easier to explain and digest. The use of these components makes the courses more interesting and dynamic, thus increasing the analytical thinking and problem-solving ability of students who participate in these programs. AI-driven simulations allow students to more autonomy over the consequences of their decisions, leading to a better understanding of course materials.

Why Governance and Accountability Matter
The panel’s third discussion focused on data governance, accountability frameworks, and content validation processes that can help organizations maintain long-term trust in their AI systems.
Darrell Gunter began by emphasizing that accountability becomes difficult when AI systems are trained on poor-quality or poorly understood data. Governance, he argued, starts with understanding the history and provenance of information. Establishing clear standards from the beginning creates a foundation for trustworthy AI development when considering accountability. He concluded by arguing that careful planning embeds the necessary governance into the training materials. When you start clean, you end clean.
Karen Medhat followed, agreeing that governance should come from within, and from the top. Definitions need to be set so that the standard can be reproducible. This is where the human-in-the loop idea comes into play. AI cannot govern itself and relies on those set rules to govern, but it has to be a human that monitors and assists with the governance at all stages.
Vamsi Kasivajjala also argued that enterprises rolling out these systems need practical and accountable standards, not only for the AI, but also for the human in the loop. Effective governance, he argued, requires shared responsibility across both technological and human components.
To conclude, Anders Hammarbäck also agreed, and stated that if you give accountability, you get accountability, and that history, attribution and governance start at the beginning. It is not impossible to create stronger foundations for trustworthy and reproducible AI systems.
The Value of Peer-Reviewed and Expert-Vetted Content
One of the webinar’s most compelling discussions centered on the role that authoritative content plays in improving AI reliability.
Vamsi Kasivajjala began this one by stating that the perspective of the data influences quality too. Drawing from diverse subject areas and credible sources helps create a more comprehensive body of knowledge and reduces the risk of blind spots within training data. Karen Medhat then chimed in by explaining that expert knowledge and subject-matter expertise remain irreplaceable. Human review helps provide context, identify inaccuracies, and ensure that information is interpreted appropriately before being incorporated into AI systems.
Mr. Hammarbäck agreed and stated that to get those deterministic answers; the kind we can learn to expect if AI is learning, peer reviewed knowledge must be the source. This is because the true scope of that knowledge is easily repeatable and diversity makes it even more trustworthy. Mr. Gunter wholeheartedly agreed, stating that accountability plays into this too as the more accountable the source, the easier it is to govern, and get quality information that encompasses all sides of the field.

Practical Strategies for Addressing the Data Quality Crisis
In the webinar’s final segment, panelists discussed practical, scalable approaches organizations can adopt to improve their data ecosystems and continuously monitor data quality.
To begin, Vamsi Kasivajjala reiterated one of his biggest points. To achieve the scaling necessary, we do not need any more data, we need the management behind it. Quality control at all levels, human-in-the-loop at all stages. Karen Medhat seconded the motion, stating that AI can be used to detect potential errors before data enters training environments.
When used responsibly, AI can help organizations strengthen the quality of datasets before they are incorporated into larger systems. Darrell Gunter chimed in last and stated that once we can establish that clean framework, working our way down from the top, from there the data can truly become trusted and replicative.
Final Takeaways
While the discussion covered a wide range of topics, one message remained consistent throughout the webinar: the future of AI will be determined not by the quantity of available data, but by its quality.
Whether discussing governance, accountability, reproducibility, training methodologies, or expert-vetted content, the panelists repeatedly stressed the importance of building strong foundations before deploying AI systems at scale. Clean data, transparent attribution, human oversight, and rigorous validation processes are no longer optional. They are essential requirements for trustworthy AI.
As the importance of trustworthy research continues to grow in both Open Science and AI development, the Open Science Education Institute (OSEI) has officially adopted the Aggregator of Global Open Science & Research (AGOSR) as its official research library. AGOSR provides a centralized discovery platform where researchers, educators, students, and industry professionals can search and explore more than one million Open Science resources, connecting users with credible, openly accessible scholarly content from around the world and supporting evidence-based learning, research, and innovation.
For those who are interested in watching the full webinar, it is available through free, on-demand video after completing the registration here. The next webinar hosted by IGI Global Scientific Publishing will take place on September 30th. This live event will discuss the impact of AI on the next generation and will feature a diverse panel to view this from all angles.







