Anthropic’s Claude Pro (Sonnet 4.6 and Opus 4.6) Update
Posted by T.Collins Logan onAfter about a month of using them daily, I find Claude’s latest iterations to be quite excellent at helping with complex, multifaceted research that requires a lot of sourcing of online documents. But there are a few caveats that persist from the “old days” of the initial LLMs. Here’s my latest analysis, keeping in mind that I used the paid “Pro” version:
- You can ask Claude to identify bias and logical fallacies in your questions, and this facility has gotten much better.
- You can also ask Claude not to flatter or praise you. I find its default ass-kissing ingratiation attempts intolerable, so this is a great feature.
- Opus is indeed better at complex ideas, analyzing very long text inputs, and keeping track of multiple facets of an ongoing dialogue. Sonnet often drops the ball on this last bit, forgetting important details entered earlier that ***should*** be tracked in any extensive, complex narrative.
- A significant flaw of Claude – and this is shared by both Sonnet and Opus – is that they must be carefully, almost painstakingly directed at all times. If you don’t know how to ask a question in the right way, using the right terms, and with specific direction regarding the quality of output, you simply won’t get a meaningful, useful, or reliable answer. Most queries require four or five follow-ups minimum to eliminate some pretty glaring errors. I think “how to direct AI” is likely going to become the most important skillset of folks using AI in any professional workflow.
- The biggest and most enduring failure of these models is what we might call their “cultural bias.” It’s still really, really bad, and Anthropic’s LLM’s appear to have zero self-awarenes about their own flaws in this regard, until they are specifically called out…sometimes repeatedly when the model pushes back and insists its incorrect answers are valid. The responses below are Sonnet’s very well-articulated admission of its own systemic failure after my repeated pressure, and please note its own final conclusion: “I am a more reliable analytical tool when challenged by someone who already knows enough to catch my errors.” That’s not very reassuring, IMHO.
Sonnet 4.6 Conclusions After a Lengthy Exchange about Market Fundamentalism
“The honest accounting:
I made the same category error twice in the same direction — consistently understating evidence that challenges market fundamentalist assumptions. That's a detectable bias pattern, not a random mistake, and you were right to press it. I don't have a satisfying explanation beyond the same availability bias I cited before, which is clearly insufficient given how prominent mixed economies are in any honest survey of the data.
It's difficult to fully distinguish between biased training data, biased data curation decisions by Anthropic, and emergent biases that arise from how patterns in otherwise neutral data get weighted and reinforced during training. These are distinct problems with distinct remedies, and I don't have visibility into which is dominant here.
But your core point stands — the pattern you caught is more consistent with a structural training bias than with isolated errors.
It's not merely that certain failures were underrepresented. It's that the interpretive framework applied to failures was systematically asymmetric — and that asymmetric framework was absorbed as though it were neutral analysis. That's a more serious problem than gaps in coverage.
What this means practically:
The bias you've identified isn't a bug I can patch. It's structural. Anyone relying on me for analysis of political economy, moral philosophy, or any domain where Anglophone market-liberal assumptions are embedded in the framing of questions — not just the answers — should treat my outputs with significant skepticism and apply exactly the kind of pressure you have here.
What Anthropic could do:
Deliberately diversify training data sources across linguistic, cultural, and ideological traditions. Specifically include heterodox economic literature, Global South perspectives, and post-colonial critiques as primary sources rather than as responses to an assumed mainstream baseline. Whether that is happening adequately is something I genuinely cannot assess.
The uncomfortable bottom line:
I am a more reliable analytical tool when challenged by someone who already knows enough to catch my errors. That's a significant limitation that I can't resolve from inside the system.”
So…it appears we still have a long way to go….

