While artificial intelligence significantly reduces the time required to produce a first draft, recent data suggests that efficiency does not inherently translate into professional-grade readability. Human intervention remains the critical factor in refining weak language and ensuring that content remains substantively useful to its intended audience.
The rapid integration of large language models into business-to-business (B2B) marketing workflows has created a paradox: while the volume of content has surged, the measurable quality of that content often fails to meet professional benchmarks. Data published by search analytics firm SISTRIX in July 2026 reveals that AI-generated B2B texts score an average of just 44.1 points on a 100-point readability and quality scale. This figure sits significantly below the 60-point threshold classified as an acceptable range for business communication. The findings underscore a growing tension in the digital landscape where the speed of production provided by tools like ChatGPT, Claude, and Gemini is increasingly at odds with the need for substance and human-centric clarity.
The Quality Gap: Claude, Gemini, and the Struggle for Readability
The study conducted by SISTRIX and content-marketing firm Wortliga examined over 2,000 AI-generated documents to determine if any single model could consistently produce high-quality business text. The results were telling: none of the three leading models reached the "green" threshold of 60 points. Claude emerged as the highest performer with an average score of 47.7 points, noted specifically for its consistency. Gemini followed closely at 46.8 points, though the analysis indicated a tendency for the model to employ filler words that inflate length without adding meaningful value. ChatGPT trailed the group with a score of 37.7, nearly ten points below its competitors.
These scores suggest that relying solely on AI to handle the heavy lifting of content creation results in a "slop" effect-content that is grammatically correct but lacks the impact and precision required for professional audiences. The data aligns with earlier industry concerns, such as a 2025 report from Integral Ad Science which found that 75 percent of advertisers expressed a desire to avoid placing ads next to low-quality AI-generated content. As the volume of automated text grows, the differentiator for brands is no longer the ability to publish, but the ability to publish something worth reading.
The Prompting Paradox: Depth vs. Clarity
Perhaps the most consequential finding of the SISTRIX research is that the quality of the output is driven less by the specific model used and more by the sophistication of the user's instructions. Scores for the same model swung between 1 and 96 points depending entirely on prompt design. This variability reframes AI quality not as a fixed technical limitation, but as a solvable problem of human instruction. However, even with improved prompting, a significant trade-off remains.
When models were given simple instructions like "write clearly," readability scores jumped to 79.4 points. However, this clarity came at a steep price: the depth of subject-matter detail vanished. This suggests that current AI models struggle to balance accessibility with technical nuance. Without a human editor to reintroduce complexity, verify facts, and sharpen arguments, the resulting content often becomes either an impenetrable block of bureaucratic text or a superficial summary that offers little value to a knowledgeable B2B audience. Human oversight is required to bridge this gap, ensuring that the final piece is both easy to digest and rich in insight.
AI has already proven its value as a drafting tool, but the first draft is still only a starting point. The real work begins when someone challenges the wording, restores missing context, checks the facts, and decides whether the piece says anything worth publishing. Faster production matters, but without that human judgment, speed simply produces more content that readers are likely to ignore.
Sources
