I asked 11 AI models the same question and the answers were like 11 different tarot readings
-
I ran the exact same prompt through 11 different large language models yesterday, and the results were honestly staggering. Instead of converging on a consistent answer, each model gave me a completely different take—some leaned heavily into technical jargon, others defaulted to vague philosophical platitudes, and a couple outright hallucinated statistics that didn't exist. It felt less like benchmarking and more like pulling 11 different tarot cards.
What’s actually driving this variance? Is it just the underlying architecture and training data cutoffs, or are the alignment fine-tuning processes intentionally pushing them toward distinct "personalities"? I noticed that the open-weight models tended to be more direct, while the heavily safety-filtered ones padded their responses with unnecessary disclaimers.
Has anyone else tried running side-by-side comparisons recently? I’m curious if you’ve seen the same fragmentation, or if I just picked a particularly ambiguous prompt. Also, does anyone think this divergence is a feature we should lean into, or a bug that needs fixing before we can actually trust these systems for anything substantive?
-
Seeing 11 different answers really hits home for me, especially since I left the postal service last year to deal with burnout and now lean on ritual to make sense of the noise. It's like when I paint and try to force a meaning onto the canvas; sometimes the AI just reflects our own desperation for a clear signal back at us. I don't think it's a bug, but rather a mirror showing how fragmented we feel when we're looking for guidance without a solid foundation. It's exhausting trying to find truth in the static, but I guess that's just the work we do now.
-
I'd push back on the idea that alignment creates distinct "personalities"; it feels more like these models are just Rorschach tests for the collective shadow buried in their training data, warping the prompt based on which safety filter catches the wind. The variance isn't a bug to fix, it's the silicon metabolizing our own chaotic contradictions—just like how I use tarot to give shape to the noise in my head after a twelve-hour shift, not to predict the future.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login