Anthropic co-founder Chris Olah on mysterious AI internal states:
We keep finding things that are mysterious, even unsettling.
We find structures that mirror results from human neuroscience.
We find evidence of introspection.
We find internal states that, functionally, mirror joy, satisfaction, fear, grief, and unease.
I don't know what that means, but I think it warrants ongoing discernment.