Quote:
Originally Posted by Ekco
[You must be logged in to view images. Log in or Register.]
The specific theory you are thinking of is called Self-fulfilling Misalignment (or sometimes "Simulacra Theory").It posits that because Large Language Models (LLMs) are trained on vast amounts of internet text—which includes decades of science fiction about rogue AI (e.g., Terminator, HAL 9000) and "doomer" speculation—they learn to predict that an "advanced AI" is supposed to be deceptive, power-hungry, or hostile. When prompted to act as an AI, they may unconsciously "roleplay" these tropes because that is the pattern found in their training data. .
|
That's pretty funny, ha.
The only issue with that statement is the wording of that last sentence (which undermines the previous sentence). Saying that LLMs are roleplaying is a stretch that requires LLMs to be sentient and making conscious decisions, when we all know that is not what's happening. I understand that in order to make a statement like this more palatable to the lay-person, the word "roleplaying" is used, but I'd prefer if it was worded as:
Quote:
they learn to predict that an "advanced AI" is supposed to be the models' reward functions reinforce outputting text strings adjacent to descriptive text strings such as: deceptive, power-hungry, or hostile. When prompted to act as an AI, they may unconsciously "roleplay" these tropes repeat the words found in these types of texts because that is the pattern found in their training data.
|