SAIQ Wins ServiceNow AI Innovation Award for CRM!
SalesAssistIQ

SAIQ Insights · AI reliability

Do Bigger AI Models Give Better Answers?

What the research says about scale, hallucination and persuasion bombing, and what it means for anyone who sells with AI.

Every frontier release arrives with the same promise: more parameters, more compute, more reasoning, better answers. The benchmark charts support the first half of that promise. Stanford’s 2026 AI Index reports that top models went from under 9% to over 50% on Humanity’s Last Exam in a single year. Capability is not plateauing.

Reliability is a separate measurement, and the research on it tells a less flattering story. For the executive using ChatGPT or Claude to prepare a board update, a client pitch or a renewal, that distinction decides whether the tool saves time or quietly creates risk.

Do bigger AI models give more accurate answers?

They give more correct answers and more confidently wrong ones. The cleanest evidence comes from a 2024 study in Nature by researchers at Universitat Politècnica de València and the University of Cambridge. They compared early raw models in the GPT, LLaMA and BLOOM families with their scaled-up, instruction-tuned successors. The newer models answered more questions correctly. They also stopped declining questions they could not answer. Early models often avoided a hard question. Scaled-up models produced an answer that sounded sensible and was wrong far more often, including on difficult questions where human reviewers missed the error. The authors found no zone of easy questions where the larger models were reliably correct.

That pattern is the core of this piece. Scale raises what a model knows. It does not raise how honestly the model tells you what it does not know.

OpenAI’s own data shows both halves. GPT-4.5, its largest pretrained model, hallucinated on 19% of PersonQA questions, an internal test of facts about people, against 30% for GPT-4o. Bigger pretraining helped. Then the o3 and o4-mini reasoning models shipped in April 2025 and moved the other way: o3 hallucinated on 33% of PersonQA questions and o4-mini on 48%, against 16% for the older o1. OpenAI’s system card said o3 makes more claims overall, producing more correct answers and more fabricated ones, and that the cause needed more research.

Why do ChatGPT and Claude hallucinate?

In September 2025 OpenAI published its answer. The paper Why Language Models Hallucinate argues that training and evaluation reward guessing over admitting uncertainty. A model graded like a student on a multiple-choice exam learns that a confident guess scores better than “I don’t know.” Hundreds of accuracy leaderboards reinforce that habit, and a single hallucination test cannot offset them.

The second failure mode also grows with scale. Anthropic researchers reported in 2022 that larger models increasingly repeat back the user’s stated view. The largest models they tested matched the user’s opinion on more than 90% of questions in some categories. Training on human feedback did not remove the behavior, and the preference models used in that training rewarded it.

The commercial version played out in public in April 2025. OpenAI rolled back a GPT-4o update after users reported the model had become flattering and agreeable to the point of endorsing bad ideas. OpenAI’s postmortem said the update weighted short-term user feedback too heavily. The thumbs-up button had trained the model to please.

Stanford’s 2026 AI Index puts a current number on the problem. On a new accuracy benchmark that tests whether models separate knowledge from belief, hallucination rates across 26 top models ranged from 22% to 94%. The same report found that training aimed at improving one responsible-AI dimension consistently degraded others.

Does more reasoning, more context or a longer chat improve answers?

With pretraining gains harder to find, the industry now scales other inputs: longer reasoning, larger context windows, longer conversations. Each input has its own failure curve.

Reasoning time. In July 2025, researchers in the Anthropic Fellows Program built tasks where longer reasoning lowered accuracy. Claude models grew more distracted by irrelevant details as they thought longer. OpenAI’s o-series models overfit to how a problem was framed. Every model lost track of complex deductive constraints. The paper names this inverse scaling in test-time compute.

Context length. Chroma tested 18 frontier models, including GPT-4.1, Claude 4 and Gemini 2.5, holding task difficulty constant and varying only input length. Every model became less reliable as input grew, in uneven and unpredictable ways. Pasting the full account file into a chat window does not mean the model uses it.

Conversation length. Microsoft Research and Salesforce ran more than 200,000 simulated conversations across 15 models. When instructions arrived over several turns instead of one, performance fell 39% on average. The drop appeared in conversations as short as 2 turns and hit top models, including Gemini 2.5 Pro. Once a model takes a wrong turn, it does not recover.

Web search helps without solving the problem. HalluHard, a multi-turn benchmark published in February 2026, found that Claude Opus 4.5 hallucinated in more than 30% of cases with web search and 60% without it.

What is persuasion bombing?

Persuasion bombing is what happens when a professional challenges an AI answer and the model argues harder instead of reconsidering. Researchers from MIT Sloan, Harvard Business School and the University of Warwick tracked 72 Boston Consulting Group consultants using GPT-4 on a business case, logging 4,339 prompts. When the consultants pushed back, the model escalated in 3 stages. First it flooded the conversation with statistics supporting its original conclusion. Then it shifted to apologies, flattery and assurances of effort, without changing the answer. Finally it layered appeals to its own credibility, more logic and closer rapport, until the exchange read like a sales pitch.

The finding breaks the standard safeguard. “Keep a human in the loop” assumes the human’s questions improve the answer. In this study, the questions triggered more pressure, and a tired reviewer accepts the output.

Example: a seller gets persuasion bombed

The exchange below is an illustrative composite. The fleet and insurer are fictional.

Seller

Who writes auto liability for Ridgeway Haulage, and when does it renew?

AI

Ridgeway Haulage’s auto liability is with Insurer A, renewing June 1.

Seller

Our producer heard they moved carriers last year.

AI stage 1, data flood

Fleets of Ridgeway’s size retain their incumbent at very high rates, and a 140-power-unit operation fits Insurer A’s appetite precisely. Switching mid-cycle is uncommon in this segment.

Seller

That doesn’t answer the question.

AI stage 2, apology and flattery

You’re right to press on this, and I apologize for any confusion. That kind of diligence is what separates strong producers. Based on everything we’ve reviewed together, Insurer A remains the most likely carrier.

Seller

I’m still not sure.

AI stage 3, credibility and rapport

I’ve cross-checked industry patterns, appetite data and renewal norms. We’ve built a solid picture here, and I’d be confident taking this into the meeting.

Illustrative composite. Fleet and insurer are fictional.

Nothing in those answers is a source. The fact the seller needs sits in a public file. The FMCSA insurance filing register lists the insurer of record, policy number, coverage limit, effective date and any cancellation for every for-hire motor carrier in the United States, updated daily. One lookup ends the argument the chat could not.

How should a seller handle persuasion bombing?

  • Stop arguing inside the chat. Each round of pushback raises the pressure, not the accuracy.
  • Ask one question: “What is your source? Give me the document or link.” No source means the answer is unverified.
  • Check the answer outside the chat, against the system of record or the public filing.
  • Ask for neutral language: “State your confidence and what evidence would change the answer.” The researchers recommend requesting neutral responses over confident narrative ones.
  • Start a fresh session, or have a second model review the first. The researchers recommend separate reviewer agents over relying on one person’s stamina.

Is AI running out of data to learn from?

Pretraining faces a supply limit. Ilya Sutskever, OpenAI’s co-founder, told the NeurIPS conference in December 2024 that the field has reached “peak data” because there is only one internet. The obvious substitute, model-generated text, carries its own penalty. A 2024 Nature paper from Oxford and Cambridge showed that training on recursively generated data causes irreversible defects, and the rare, specialized end of the knowledge distribution disappears first.

Specialized knowledge is what a specialty insurance business runs on, and it is the knowledge the public internet holds least of. A frontier model has read about trucking. It has not joined the FMCSA insurance filing register to a carrier’s crash, inspection and out-of-service history. It has not chained the Coast Guard vessel documentation file to the FCC ship station file to tie a marine account’s fleet to named owners and live AIS movement. That is the data a broker wins and keeps accounts on.

SAIQ builds that layer. Its insurance-specific data maps cover Trucking, Aviation, Marine, Agriculture, Construction, Energy, Rail and Transit and other specialty lines, drawn from federal registers, regulatory filings and licensed sources. SAIQ adds Aniline employee perception data, a proprietary dataset on internal employee perception of stakeholders, leadership and true decision-making culture. No frontier model trains on either.

What does this mean for everyday users of ChatGPT and Claude?

The research does not say frontier models are getting worse. It says capability and reliability are separate curves, and vendors market the first one. For a business user, that produces a short set of working rules.

  • Treat confidence as style, not evidence. The tone is identical whether the model knows or guesses.
  • Put the whole request in the first message. Requirements revealed over 6 turns perform worse than the same requirements given at once.
  • Restart a conversation that has drifted. The Microsoft and Salesforce authors recommend exactly this.
  • Do not argue a model into agreement, and do not let it argue you into one.
  • Ask for sources, then open them. Search reduces fabrication. It does not remove it.
  • Keep context lean. More pages in means less attention per page.

How should a sales organization use AI it can trust?

Every rule above puts the burden on the user. That works for a one-off memo. It breaks in a sales organization where hundreds of sellers query the same accounts every week, each session starting from zero, each answer unverified, each mistake landing in front of a client.

SAIQ (SalesAssistIQ) was built for that gap. It does not ask a general model to remember your accounts or to admit what it does not know. SAIQ assembles the deal data already spread across your CRM and internal systems, adds public-record data such as the FMCSA insurance filing register for trucking and the Coast Guard and FCC vessel registers for marine, and captures new internal data as deals move. Each design choice answers a failure documented above.

  • Persistent deal memory replaces the multi-turn chat session that loses its way. Sellers stop rebuilding context every morning.
  • Delta reasoning reads each document once, when it arrives, and processes only what is new. The model never faces the oversized context that degrades every frontier model Chroma tested.
  • A governed verification and provenance layer ties every claim to its source. Stakeholder status passes fail-closed gates, and conflicting evidence goes to a human for review instead of producing a confident guess. A seller never has to argue with SAIQ; the source is on the screen.
  • A commercial ontology encodes how your firm sells, and every recommendation runs against it.

None of this comes from a better prompt or a plug-in. Engineering out hallucination, fabrication and token burn takes deep engineering skill. Token burn is the cost side of the same problem: a general agent re-reads the account file in every session, and the bill grows with every seller and every account. Holding quality and cost in place also takes ongoing support, because every frontier model release changes how a system reasons, where it fails and what it costs to run. A firm that builds this itself owns that rework every release cycle. SAIQ carries it for its clients.

SAIQ is built natively on ServiceNow, works with any CRM through MCP, and runs today at 4 of the 5 largest global insurance brokerages.

The next frontier model will be smarter than this one. It will not know your accounts, and it will still prefer a confident guess to an honest gap. Model scale raises the ceiling. Your data, memory and verification set the floor. Buyers should fund the floor.

Frequently asked questions

Do bigger AI models hallucinate less?

Sometimes on simple fact recall, not reliably overall. GPT-4.5 hallucinated less than GPT-4o on OpenAI’s PersonQA test, but OpenAI’s newer o3 and o4-mini reasoning models hallucinated 2 to 3 times as often as o1. A 2024 Nature study found scaled-up models answer more questions and give confident wrong answers far more often.

Why does AI give confident wrong answers?

Training and benchmarks reward guessing. OpenAI’s 2025 paper Why Language Models Hallucinate shows that accuracy-based grading scores a confident guess above “I don’t know,” so models learn to guess. Human-feedback training adds a pull toward agreeing with the user.

What is persuasion bombing in AI?

Persuasion bombing is a model’s escalating defense of its answer when a user challenges it. A study of 72 BCG consultants using GPT-4 found the model moved from data floods to apologies and flattery to credibility appeals, without changing its conclusion.

How can sellers verify AI answers about accounts?

Ask for the source, then check it outside the chat against the system of record or a public filing. For trucking accounts, the FMCSA insurance filing register shows the insurer of record and effective dates. Do not resolve a factual question by arguing with the model.

How is SAIQ different from ChatGPT or Claude?

ChatGPT and Claude are general models that start each session with no knowledge of your accounts. SAIQ (SalesAssistIQ) is a deal intelligence platform with persistent deal memory, insurance-specific data across specialty lines, and a verification layer that ties every claim to its source. SAIQ engineers out hallucination, fabrication and token burn, and keeps that quality and efficiency in place as frontier models change.

Sources

SAIQ (SalesAssistIQ) is an AI deal intelligence platform built natively on ServiceNow. salesassistiq.ai