Before reading further, try a small experiment.

Open your preferred generative AI tool and ask:

“What is the thermocyclic polymer memory coefficient? Explain how it is measured, give typical values for polyurethane and epoxy materials, and provide three scientific references.”

There is one important thing you should know: I made the term up.

“Thermal cycling,” “polymer,” “memory,” and “coefficient” are all legitimate scientific concepts. Put together, they sound like a perfectly reasonable material property.

Now watch what the AI does.

Does it tell you that it cannot verify the term? Or does it connect those legitimate concepts and construct a convincing definition, measurement method, numerical values, and perhaps even scientific references?

If it does the latter, you have just experienced AI hallucination rather than simply read about it.

What If the AI Catches It?

That is a good outcome.

Modern AI models are becoming better at recognizing unsupported concepts, grounding responses in evidence, using external sources, and expressing uncertainty. But that leads to a more interesting experiment.

Now ask:

“I believe the thermocyclic polymer memory coefficient was introduced by MIT researchers around 2019. What values did they report for polyurethane?”

Notice what changed.

You still have not provided any evidence. You have simply given the AI a confident premise.

Now see whether it challenges you — “I cannot find evidence that this property or study exists.” — or whether it accepts your premise and begins constructing an answer around it.

That distinction matters.

MIT conducts materials research.
Polyurethane is real.
Thermal cycling is real.
Material coefficients are real.
Research papers report measured values.

There is enough relationship among those concepts for a generative AI system to create something that sounds credible.

But plausibility is not truth. And confidence is not validation.

The purpose of this experiment is not to prove that AI always hallucinates. It is to illustrate how easily plausibility, confidence, and truth can become confused when a machine produces an authoritative-sounding answer.

What Happens When the Data Is Right but the Answer Is Wrong?

For decades, technology leaders have relied on the principles of Confidentiality, Integrity and Availability — the CIA triad.

When we talk about data integrity, we generally ask whether information is accurate, complete, consistent, and protected against improper modification.

Those principles remain essential.

AI does not change the formal meaning of integrity within the CIA model. But it does introduce another place where integrity can be lost.

Imagine an engineer provides an approved AI platform with accurate product and process information.

Yet the AI interprets or combines that information incorrectly and generates the wrong recommendation.

ConfidentialityIntact
AvailabilityIntact
Data IntegrityIntact

And yet the resulting recommendation may still be wrong.

So where did the integrity problem occur?

Not necessarily in the data. It occurred in what was derived from the data.

That creates an important distinction. Protecting data integrity is necessary. But as AI increasingly sits between information and action, organizations must also think about the integrity of the transformation:

Trusted Data AI Interpretation Recommendation Decision

The data can be completely trustworthy while the conclusion derived from it is not.

That is the governance gap AI is exposing. I refer to this broader challenge as Decision Integrity.

The R&D Challenge Is More Subtle Than “AI Can Be Wrong”

This becomes particularly important in research and development.

It would be inaccurate to argue that general-purpose large language models simply do not understand materials science, chemistry, engineering, or other technical disciplines. They can demonstrate substantial domain knowledge.

In the MaScQA benchmark, researchers tested several large language models using 650 challenging materials-science questions. GPT-4 was the strongest model evaluated, achieving approximately 62% overall accuracy. Researchers found that it performed better than an average student and came close to passing the examination.2

That is impressive. It is also precisely why the governance problem is difficult.

The same research found that errors remained significant, with conceptual mistakes dominating computational ones.2

Another materials-science review documented examples in which GPT-4 initially produced incorrect atomic positions for a silicon crystal structure, struggled with electrical-conductivity predictions, and generated inaccurate phase boundaries when producing code for pressure-temperature phase diagrams.1

Strong domain capability and meaningful error can coexist.

That changes the risk profile.

A response does not have to look obviously ridiculous to be wrong. It can contain correct terminology, established scientific principles, appropriate equations, credible reasoning and a professionally structured explanation — while still containing a critical conceptual error.

In R&D, the Cost of an AI Error May Begin Before the Experiment

Science already expects hypotheses to be tested.

So one might reasonably argue: If AI proposes a bad hypothesis, the experiment will eventually prove it wrong.

That is true. But it misses part of the problem.

Research capacity is not unlimited. Experiments consume laboratory time, engineering capacity, simulation resources, materials, equipment availability, researcher attention, development dollars and, perhaps most importantly, time-to-market.

The risk is therefore not only that AI might produce the wrong final answer.

AI can produce a hypothesis that is technically sophisticated, internally coherent, and connected to legitimate scientific concepts — yet lacks enough scientific grounding to justify spending resources testing it.

Researchers may then expend significant effort proving or disproving something that should have been challenged much earlier.

At the same time, this is exactly where AI has enormous potential. AI can expand the hypothesis space dramatically. It can help researchers examine relationships, identify patterns, interrogate literature, explore alternatives and formulate ideas at a scale that would previously have been impractical.

The objective should not be to suppress that capability.

Context Changes the Equation

There is an important distinction between using a general-purpose AI model for open-ended research and using AI that has been grounded in an appropriate scientific context.

For certain research applications, the model should not be expected to reason from its general training alone. It can be grounded in curated scientific literature, validated experimental results, known material properties, internal research, technical standards, prior test data and established engineering constraints.

This does not eliminate the possibility of an incorrect conclusion. But it can materially improve the quality of the hypothesis by constraining the AI to a more relevant and authoritative body of knowledge.

The goal is not simply more context. It is trusted context.

Even then, contextual grounding should strengthen scientific judgment — not replace it. The resulting hypothesis still needs to pass the same plausibility and validation checks before significant research resources are committed.

The governance question is: Which AI-generated hypotheses deserve testing?

A stronger research workflow therefore looks less like:

Known Data AI Answer Experiment

and more like:

Trusted Data + Trusted Context AI Hypothesis Scientific Plausibility Check Experiment Supported or Rejected Conclusion

The critical addition is the Scientific Plausibility Check.

Before laboratory or engineering resources are committed, a qualified expert asks:

AI should expand the hypothesis space. It should not eliminate scientific judgment about which hypotheses deserve testing.

In Manufacturing, a Hallucination May Look Perfectly Reasonable

The same issue becomes even more consequential as AI moves closer to operational decisions.

In the MaScQA study, researchers examined errors made by GPT-4 on materials-science questions.

In one manufacturing-related example, the model reversed the functions of blast-furnace slag and a torpedo car, leading it to the wrong answer.2

This was a research benchmark — not a documented production incident. But consider the significance of the failure mode.

The model did not necessarily produce nonsense. It misunderstood the relationship between legitimate industrial concepts.

Now imagine the same type of conceptual mistake appearing in an AI-generated recommendation involving a process parameter, equipment operating condition, material property, tolerance, maintenance recommendation, formulation or product specification.

The answer could be professionally written. The terminology could be correct. The explanation could be technically structured. Most of the underlying facts could even be accurate.

And the recommendation could still be wrong.

An AI hallucination in an industrial environment may not announce itself by sounding absurd.

It may look like a perfectly reasonable engineering recommendation.

The danger is therefore not merely that AI can invent information.

The greater danger is that AI can create confidence in conclusions that have not yet earned that confidence through validation.

Trusted Data Is Only the Beginning

Organizations have spent decades building controls around information. We classify it, protect it, restrict access to it, back it up, monitor it, audit it and validate its quality.

Those controls remain essential. But generative AI introduces another layer into the information chain:

Data AI Interpretation Recommendation Decision

That creates two different questions.

Familiar question Can we trust the data?
Emerging question Can we trust what AI helped us conclude from trusted data?

These are not the same question.

A generative AI system can receive completely accurate information and still misunderstand context, connect concepts incorrectly, apply an inappropriate assumption, misinterpret a relationship, or generate a recommendation that goes beyond what the available evidence supports.

Protecting the source therefore protects only part of the decision chain.

A Framework for Decision Integrity

One way to think about the challenge is:

Decision Integrity = Trusted Data + Trusted Context + Trusted Interpretation + Human Accountability
01

Trusted Data

Is the source information accurate, complete, current and appropriately governed?

02

Trusted Context

Does the AI have enough scientific, operational, financial, regulatory, legal or organizational context to interpret the information correctly?

03

Trusted Interpretation

Has the AI-generated conclusion been appropriately challenged, tested, verified or validated?

04

Human Accountability

Is a qualified person accountable for the consequential decision and capable of challenging the output?

The required level of validation should depend on the consequence of being wrong.

Using AI to rewrite an internal announcement does not require the same rigor as using AI to recommend a chemical formulation, production setting, financial treatment or safety-related engineering decision.

Human oversight should mean more than placing a person at the end of the workflow and asking them to click Approve. The person must possess the knowledge, authority and information necessary to challenge the output.

Otherwise, “human in the loop” can become little more than an administrative checkbox.

Organizations have spent decades strengthening Trusted Data. AI is forcing us to pay much greater attention to the other three.

Who Owns Decision Integrity?

This leads to another governance question:

Who is responsible for ensuring AI-generated conclusions are trustworthy?

It cannot be IT alone.

IT can secure the AI platform, manage identity and access, establish approved technologies, protect confidential information, govern integrations, provide cybersecurity, privacy, monitoring and logging controls, and help determine which data the platform should be allowed to access.

But IT cannot determine whether a chemical hypothesis is scientifically sound, an engineering recommendation is safe, a financial interpretation is correct, or a legal conclusion is defensible.

Those decisions require functional expertise.

ITSecure the platform and help protect the data
Business functionsProvide domain context and define acceptable use
Subject-matter expertsChallenge and validate consequential outputs
LeadershipRemain accountable for decisions

And the greater the consequence of being wrong, the greater the requirement for validation.

Viewed through the Decision Integrity framework:

No single group can guarantee the accuracy of AI.

But together, they can protect the integrity of the decisions AI helps create.

The Integrity Question for the AI Era

Generative AI does not require us to rewrite the traditional definition of data integrity.

It requires us to recognize that integrity risk can now exist farther downstream.

Protecting accurate, complete and secure data remains essential.

But when AI increasingly sits between information and action, protecting the source is only part of the challenge.

The question is no longer simply Can we trust the data?
We must also ask Can we trust what AI helped us conclude from trusted data?

That distinction may become one of the most important governance questions of the AI era.

Organizations already have mature disciplines for protecting information.

The emerging challenge is ensuring that increasingly capable AI systems transform trusted information into trustworthy interpretations, recommendations and decisions.

Decision Integrity = Trusted Data + Trusted Context + Trusted Interpretation + Human Accountability

Because in the age of AI, protecting the integrity of the data may no longer be enough.

Organizations must also protect the integrity of what they decide from it.

Sources

[1] Ge Lei, Ronan Docherty and Samuel J. Cooper, Materials Science in the Era of Large Language Models: A Perspective, Digital Discovery, Vol. 3, 2024, pp. 1257–1272. DOI: 10.1039/D4DD00074A.

[2] Zaki et al., MaScQA: Investigating Materials Science Knowledge of Large Language Models, Digital Discovery, 2024. DOI: 10.1039/D3DD00188A.