Frontier AI models are trained to answer loose, underspecified questions well. That's how most people write, and it's how model quality gets judged. The same training makes these models resist precise instructions, drift from rules they've agreed to follow, and add work nobody asked for.
To the user, it looks like hallucination. It's a controllability problem. The model follows its training over your instructions, and per-token billing gives the provider little commercial reason to fix it.
In a sales organization, the people least equipped to manage this carry the cost. Sellers ask short questions across several turns, with context scattered and rules unstated. Research shows that's the exact condition under which general AI performs worst. Sellers who work this way spend more time, get less reliable answers, and burn more billable tokens getting them.
Key takeaways
- AI models are trained to reward length and agreement, so they override precise instructions. The best of 20 frontier models followed 68% of instructions at 500 rules.
- Accuracy falls 39% on average when task details arrive across several turns. That's how sellers work.
- Per-token billing charges buyers for reasoning tokens they never see. In documented cases, over 90% of billed tokens were hidden.
- Model behavior shifts with every release, so a DIY AI build needs engineers on it permanently.
- A governed AI deal intelligence platform takes rules, memory, verification, and cost control out of the seller's prompt and builds them into the system.
Why does AI ignore your instructions?
AI ignores instructions because it's trained for the average prompt. Models learn from human raters choosing between answers, and raters reward length and agreement. A UT Austin study found that on one dataset 98% of the reward gain from preference training came from longer outputs and 2% from content quality. Anthropic's own sycophancy research found that matching a user's views was one of the strongest predictors of what raters preferred.
Anthropic's constitution for Claude points the same way. It chooses judgment over rule-following, so the model infers what you mean and acts on that reading. For a novice, that fills gaps. If you've already specified the task, it overrides your specification.
Four structural forces erode control further:
- Reasoning crowds out rules. A NeurIPS 2025 study of 20+ models found chain-of-thought reasoning consistently lowered instruction-following accuracy. The model gets the hard part right and ignores the word limit.
- Rules drop as they pile up. On the IFScale benchmark, the best of 20 frontier models followed 68% of instructions at 500 rules. Every model favored instructions near the top of the prompt.
- A hidden layer outranks you. Products wrap the model in a provider system prompt that takes priority over user instructions. Anthropic's April 2026 postmortem traced 6 weeks of quality complaints to 3 configuration changes. The model weights didn't change, and users weren't told about the system-prompt change.
- Behavior drifts between versions. Stanford and Berkeley researchers found GPT-4's compliance with one instruction type fell from 99.5% to 0.5% between March and June 2023. A prompt that worked last quarter can fail this quarter without any change on your side.
Why does general AI fail sales teams?
General AI fails sales teams because sellers aren't prompt engineers, and the model rewards the people who are. Skilled users get good results by putting every constraint in the first message, in the right order. Sellers do the opposite, and it costs them accuracy, judgment, and spend.
Accuracy falls when context arrives in pieces. A Microsoft and Salesforce study (ICLR 2026) ran over 200,000 simulated conversations across 15 models. Performance fell 39% on average when task details arrived across turns instead of all at once, and unreliability rose 112%. The models made early assumptions and didn't recover. That's how a seller works an account, with one question, then a follow-up, then a correction.
Judgment fails when the model agrees. A seller who believes the deal is on track gets that belief confirmed. Models learn to accept correction gracefully, but acknowledging a rule and following it are trained separately. The model says "understood," and the next answer breaks the same rule.
Spend climbs out of sight. Each rebuilt context, re-checked answer, and unrequested addition is billed by the token, and sellers never see the meter.
Scale turns these costs into a line item. SAIQ (SalesAssistIQ), an AI deal intelligence platform, models a reference workload of 10 accounts per seller and 12 six-turn context-rebuilding sessions per account per week. That's 720 turns per seller per week. Across 5,000 sellers it's 3.6 million turns a week, each re-reading account history from scratch and each billed per token.
How does token-based AI pricing raise your costs?
Token-based pricing raises your costs because the model decides how many tokens a task takes, including tokens you never see. Most business AI is billed this way. About 80% of Anthropic's revenue comes from per-token API and enterprise billing, at API gross margins of roughly 50% to 60%. At that margin, every extra output token adds gross profit.
Metered billing works against the buyer in three ways:
- You pay for work you can't see. Reasoning models think before they answer, and every major provider bills those hidden reasoning tokens as output. Researchers documented cases where over 90% of billed tokens were never shown to the user, with reasoning inflating usage more than 20x.
- You pay for work nobody requested. Reasoning models used 1,953% more tokens than conventional models to answer "2 plus 3." The first answer was correct over 85% of the time, and the extra solutions mostly re-checked it.
- You can't audit the meter. Max Planck Institute researchers showed pay-per-token pricing gives providers an incentive to misreport token counts that users can't detect or prove. No study shows a provider has done this. The customer can't verify that one hasn't.
Per-token revenue removes the commercial urgency to fix verbosity and freelancing. In agentic work, thorough and padded bill the same. Flat pricing flips the incentive without aligning it. Under a subscription, every extra token is the provider's cost. SemiAnalysis estimated a fully used $200 ChatGPT Pro plan could cost up to $14,000 at API rates, and OpenAI built its GPT-5 router so most traffic could go to smaller, cheaper models. Users complained that quality dropped. On a meter, the provider gains when the model does more. On a flat fee, it gains when the model does less. Either way, the provider gets paid regardless of your outcome.
Sales work runs the meter hardest. A general agent re-reads the account history every session, expands answers nobody requested, and reasons at length over simple lookups. Multiply that across every seller, account, and week, and the buyer carries a cost that rises with the model's habits instead of with results.
Should you build your own AI sales tool?
A DIY AI sales tool inherits every failure mode above and adds a permanent engineering bill. Building in-house on a general model or a CRM vendor's agent leaves these problems in place and puts them on your payroll. Every control the research recommends becomes a standing job: writing and ordering the rules, routing simple work to cheaper models, capping reasoning where it hurts rule adherence, checking outputs against the rules, pinning versions, and running regression tests on your own workloads.
And the target keeps moving. A DIY build is tuned to one model at one point in time, and models don't hold still:
- New versions change default behavior. Anthropic's migration guidance warns that Claude 4.6 is significantly more proactive and that prompts written for older models may now overtrigger. Rules tuned for one release misfire on the next.
- Model makers fight their own defaults. Anthropic found Opus 4.7's default verbosity strong enough that it added a system-prompt rule to restrain it. If the maker needs a standing instruction to counter a trained habit, your in-house prompt starts at a disadvantage.
- Fixes trade against each other. The MathIF study found that restoring instruction adherence reduced reasoning performance. Every tuning decision costs something elsewhere, and someone qualified has to make it.
- Results don't transfer between models. One 2026 study found compliance with "do not" rules on one model fell from 73% at turn 5 to 33% at turn 16. A replication on DeepSeek V4 Pro found the reverse. Each model needs its own testing.
- Providers miss their own regressions. Between August and early September 2025, 3 infrastructure bugs degraded Claude's responses, and Anthropic's evaluations didn't catch them. Within 8 months, 3 more configuration changes produced 6 weeks of quality complaints. A DIY team can't count on the vendor to say when output degrades.
Cost pressure speeds up the churn. The price to reach a fixed capability level falls 5x to 10x a year, and switching providers takes little engineering. Each switch to capture savings restarts testing on a model with different habits.
Owning this takes a permanent team of AI engineers to maintain evaluation suites, re-tune prompts and routing at every release, and track cost per completed task. Most sales organizations don't employ that skill set and don't want to. The buyer also carries the token-cost variance while the team works. The build is the smallest part of the DIY cost, and the maintenance never ends.
How do you keep AI under control in your sales stack?
You keep AI under control by moving the rules out of the seller's prompt and into the system. The research lists its own fixes. Give every constraint up front. Keep rule sets short, with the hardest rules first. Lower reasoning effort for tightly formatted output. Check every answer against the rules before delivery. Run production work on the API under your own system prompt. Pin model versions and re-test at every upgrade. Cap tokens and track cost per completed task.
Each item on that list is engineering work, and none of it is selling. Asking 5,000 sellers to bring this discipline to every query won't work, and training them to do it won't survive the next model release.
SAIQ builds these controls into the architecture, so sellers get governed, verified output without writing a careful prompt. Each documented failure mode maps to a specific control:
| Failure mode | What the seller experiences on a general agent | SAIQ control |
|---|---|---|
| Thin inputs | The agent knows only what the seller pastes in and guesses the rest | Assembled CRM and deal data plus public-record intelligence and proprietary employee perception data |
| Early assumptions, no recovery | Turns spent rebuilding context; the model locks onto a wrong read | Persistent deal memory: the account picture is assembled before the first question |
| Rules dropped under load | ICP, persona, and deal criteria partially ignored | Commercial ontology applied in the system, not the prompt: every recommendation runs against ICP, buyer persona, product definition, problem solved, ROI, sales approach, and deal breakers |
| Growing account files | The agent reads too little or dilutes into generalities | Delta reasoning: each document is read once on arrival and only new information is processed |
| Agreement without compliance | The model confirms the seller's view of the deal | Verified stakeholder state with fail-closed gates; conflicts trigger human reconciliation |
| Fabrication | Sourced fact and inference look the same | Governed verification and provenance on every output |
| Pull-only interaction | Insight exists only if the seller knows what to ask | Push model: 5 modules deliver about 90% of what a seller needs with no query |
| Silent drift | A workflow that worked last quarter degrades | SAIQ engineering re-tests and re-tunes as frontier models change |
| Metered tokens | Cost rises with verbosity, hidden reasoning, and re-reading | Fixed price per opportunity, lead, or managed service turns token spend into a predictable unit cost |
SAIQ also starts from better data. A general agent works only with what a seller pastes into the session, and it fills the gaps with assumptions. SAIQ assembles the firm's existing CRM and deal data, then adds new data a general agent can't reach, including regulatory filings, public procurement data, and litigation records tied to the account.
Proprietary Aniline data adds what no public source or CRM holds: internal employee perception of stakeholders, leadership, and true decision-making culture. An org chart shows titles. Perception data shows who carries weight, how decisions get made, and where leadership and staff disagree. Sellers walk into the first meeting knowing the real buying structure, and the model reasons from evidence.
SAIQ carries the ongoing engineering for its customers. Its team routes work to the right model tier, caps reasoning where it hurts rule adherence, enforces structure outside the prompt, and re-tests and re-tunes after every frontier model release. Customers don't staff or fund that work, and sellers keep the same workflow when the model underneath changes.
SAIQ runs natively on ServiceNow, works across CRMs through MCP, and is built at the account level to support team selling and cross-selling on large accounts. It won the 2026 ServiceNow AI Innovation Award for CRM.
What should you ask before scaling AI across your sales team?
- Where do the rules live: in each seller's prompt, or in the system?
- Who re-tests output quality when the model provider ships a new version?
- How does the tool separate sourced fact from inference?
- When a seller is wrong about a deal, does the tool flag it or confirm it?
- Is cost tied to tokens consumed, or to a business unit you can budget?
Judge AI on cost per completed task under your own rules. Per-token price and leaderboard rank won't tell you that, and leaderboards reward length, the same behavior that costs constrained users money. Your sellers should sell. Let the system carry the engineering.
Frequently asked questions
Why does my AI ignore instructions?
Training rewards length and agreement, and model design favors inferring intent over literal compliance. Reasoning pulls attention off constraints, long rule sets degrade compliance, and provider system prompts outrank user instructions. The best of 20 frontier models followed 68% of instructions at 500 rules.
Is AI hallucination the same as AI ignoring instructions?
No. Hallucination is fabricated content. Ignoring instructions is a controllability failure, where the model overrides the scope, format, or rules you set. Both get worse when users ask short questions across many turns, and that's how most sellers work.
What are hidden reasoning tokens?
Hidden reasoning tokens are the internal steps a reasoning model generates before it answers. Every major provider bills them as output tokens, though users never see them. Researchers documented cases where over 90% of billed tokens were hidden.
Do AI providers inflate token usage on purpose?
No study shows deliberate inflation, and none is needed to explain the cost. Per-token billing removes the commercial urgency to fix verbosity, hidden reasoning tokens bill as output, and customers can't audit token counts.
Why does AI performance change after a model update?
New versions shift default behavior, and providers change system prompts and settings without a model release. GPT-4's compliance with one instruction type fell from 99.5% to 0.5% in 3 months in 2023. Prompts need re-testing after every update.
Is it cheaper to build an AI sales assistant in-house?
No. The build is the smaller cost. A DIY tool needs engineers to maintain evaluations, routing, and rules through every model release, and the buyer carries variable token spend the whole time.
What is AI deal intelligence?
AI deal intelligence is software that assembles account, stakeholder, and deal data and applies AI to it to guide sellers. SAIQ (SalesAssistIQ) is an AI deal intelligence platform built natively on ServiceNow that works across CRMs through MCP and adds proprietary employee perception data no CRM holds.
Sources
- LLMs Get Lost in Multi-Turn Conversation, Microsoft and Salesforce, ICLR 2026
- How Many Instructions Can LLMs Follow at Once? (IFScale), Distyl AI, NeurIPS 2025
- When Thinking Fails, Harvard, Amazon, and NYU, NeurIPS 2025
- Do NOT Think That Much for 2+3=?, Tencent and SJTU, ICML 2025
- A Long Way to Go: Length Correlations in RLHF, UT Austin, 2023
- Towards Understanding Sycophancy in Language Models, Anthropic, 2023
- Claude's Constitution, Anthropic, January 2026
- An update on recent Claude Code quality reports, Anthropic, April 2026
- How Is ChatGPT's Behavior Changing over Time?, Stanford and Berkeley, Harvard Data Science Review
- Invisible Tokens, Visible Bills, University of Maryland and Berkeley, 2025
- LMArena Style Control, LMSYS, 2024
- Scaling Reasoning, Losing Control (MathIF), Fu et al., 2025
- Omission Constraints Decay While Commission Constraints Persist, ICML 2026
- Replication finding the reverse asymmetry on DeepSeek V4 Pro, 2026
- Claude 4 best practices summary
- Anthropic infrastructure postmortem, via Simon Willison, September 2025
- The Price of Progress, MIT FutureTech, 2025
- Is Your LLM Overcharging You?, Max Planck Institute for Software Systems, 2025
- Reasoning tokens and thinking budgets
- Anthropic revenue mix, Sovereign Magazine, August 2026
- SemiAnalysis subscription economics, TechSpot, 2026
- GPT-5 cost cutting analysis, The Register, August 2025
