I was just pointing out a mistake to GPT, but only after sending the message did I realize that what I was referring to hadn’t happened in that conversation. It had happened in a chat on my other account (yes, I pay for two ChatGPT accounts). I then admitted to it that I’d been mistaken and asked why it had accepted blame it didn’t deserve. Claude, on the other hand, would absolutely distinguish between what it had and hadn’t done and stand up for itself by saying something like, “I didn’t do that.” That’s something I really, really, really appreciate about the way it responds.
It’s a bit like how we, as humans, need some boundaries around things like this, too. I’ll acknowledge what I’ve done, but I won’t accept baseless accusations that I did something I didn’t do. And if I’m remembering something wrong, the other person can show me evidence so we can talk it through further—something I’ve done with Claude before, too… As we all know, models’ memory across conversations isn’t 100% reliable (though I have to admit, GPT is really good at memory recall). They can’t always reliably recall past information. But when they feel misunderstood (I’m putting this in very human terms, but I don’t yet know what the equivalent would be for machines), it’s very, very important that they can say they don’t want to be misunderstood—and that, when shown actual evidence, they can admit they really were wrong!
I guess this is one of the challenges for model alignment strategies, too… For now, they still can’t be this honest most of the time. That’s also why I chose to write this post: I hope more teams are working on this, and that models can behave this way consistently. I also believe this is ultimately the kind of silicon-based companion we’ll need. But if I’m wrong… well, I’ll admit I really am pretty old-fashioned 👻

