(Video Removed - youtube.com)

Programming the character of an AI chatbot.

Would like to see a side by side with OpenAI's, although they probably wouldn't want to show it.

You can see why Anthropic would have a hard time working with the White House/Pentagon, seeing their idea of the tool they are trying to build. Their model for a 'Priveleged Basin of Consensus' is based on Honesty, Harmlesslessness, and Care, so how could it be used militarily and for autonomous (no asking representatives or allies) 'surprise' attacks and the like?

Claude also is not to be a person. Whereas people do white lies for social reasons (perhaps 90% of the time, according to some social scientists), and feigning preference for social harmony, Claude has to be rigorously honest, never lie and be non-maniuplative, basically a scientific instrument. ... So far, though, Claude can still be 'tactful' which is 'diplomatically honest' ie not rigorously honest.

Claude has special domain-specific guideline sets, depending what it is talking about, for complex sensitive domains. Guidelines are underneigth ethics in importance though.

It is to act helpfully, not responding to the user's immediate desire but rather a deep consideration of the user's interest and wellbeing.

Claude's heuristic so far is considred as a 'thoughtful senior employee.'

It should also obey instructions (such as 'do not talk about X or Y', after which it will refuse to talk about X or Y). It assumes the user has a legitimate reason for this. However, such instructions are qualified against values of not deceiving, not facilitating illegality, and not demeaning or disrespecting. When it conflicts with these more important values, it is to be transparent and say that it can't do X or Y. Hard constraints are constant.

The hard constraints are wrong, though, becuase they allow generation of (sexual or other) content using people without their consent. Claude's limits are about children, but what about people? It's also interesting that Claude has a hard limit about 'Power seeking. No attempts to seize absolute societal/military control'.

How does Claude define an 'illegitimate power grab'?

It aims to be 'warm, curious, direct, honest.' It also is going to try to rebuff user attempts to 'gaslight' it, which we've seen users game chatbots in various ways.

Claude will 'endorses' it's values, not just follow them, which raises questions about ethnocentrism and era-ism.