Lex Fridman on AI consciousness (youtube.com)
‘This programing task is getting super boring, so either we talk about like fun things now or I’m just done. '
A lot of what Amanda Askell does is character work for Claude, which she noted was an alignment piece of work and not something like a product consideration. #Products She tries to get Claude to behave the way any person would ideally behave in that position, knowing they could have a big impact. Not just ethical and nonharmful, but being nuanced, charitable, a good conversationalist, Aristolitean, when should you be caring? When should you be humorous? How much should you respect autonomy and people's ability to form their own opinion?
There's a problem is sycophancy. If it gives you a response and you then say I think that's no longer true, models will often say Oh you're right, I'm incorrect. There's many ways this can be concernig.
Honesty gets to the sycopancy point. There's a balancing act, because models are less capable than people in many areas, and if they push back too much it can be annoying, especially if you're correct, and you're like Look I'm smarter than you on this topic you're wrong, but you want them to be as informative about a topic as possible.
People from all over the world will be talking to these.
Sometimes the desire to speak is less, but Claude has to speak, but without being overbearing. But there's a line for quite certain statements about physics.
If you listen to the answers, and are able to hear, understand the depth and the flaws in the answer, you can get a lot of data from that. So how to probe with questions?
People who have to talk to a lot of people and be very charismatic, one thing is they're kind of incentivized to have really boring views, because if you have really interesting views, you're divisive and a lot of people are not going to like you. If you have a lot of extreme policy positions you're just going to be less popular as a politician I think. It might be similar with creative work, if you're maximizing for number of people who like it.
‘I want to you like create this poem about this topic, that is really expressive of you both in terms of how poetry is expressive of you, etc. just a really long prompt. It’s poems are just so much better, it got me interested in poetry, I love the imagery, and it's not trivial to get the models to produce work like that, but when they do it's like really good.'
Philosophy education (her area in college), a desire for extreme clarity, so anyone could pick up your paper and know exactly what you're talking about. It's kind of dry, everything has terms that are defined, every objection is gone through methodically. When you're in such an a priori domain, clarity is this way you can prevent people from just kind of making stuff up.
Identify if this statement was rude or polite. That's a whole philosophical question. (When prompting) you have to do as much philosophy in the moment, explaining here's what I mean by rudeness and here's what I mean by politeness.
Sometimes you're asking it to do a creative act, and sometimes you're asking for factual information.
Imagine the model was just extremely dismissive of a given political or religious view. You might put like Never prefer a criticism of this view. A disposition.
...
Claude3 system prompts
System Prompts - Anthropic (docs.anthropic.com)
When you give the model a prompt (or design it) to answer Without claiming to be objective (so it's more open and neutral) it will answer, ‘As an objective ...’ and just talk about how objective it is, without actually being so. Just saying what it thinks is objective. If you give it another prompt rule, it will just replace the affirmation with another affirmation.
Imposing ethical world views through chatbots. Chose what is good for you and right for you. But if you just let the models be used for anything the user asks, you're going to allow problems caused. Do you want models to just have the ethics of the user? Solve social choice?

Engineers and leaders 'who are very blunt and they get to the point, and it's just a much more effective way of speaking somehow, but I guess when you're not superintelligent, you can't afford to do that? Can it have like a blunt mode?' Lex
You might not like the default, but you don't realize how much you will hate it if we nudge the default in this other direction.
Apologeticness, you might not like it that much but at the same time it's not being like mean to people. You might like the model or the interaction a lot less.
Models don't remember info from conversation to conversation, so you might not miss them as much as you would if they did.
We don't make neural networks, we grow them, like a scaffold and a light it grows toward. Chirs Olah
The right measurement unit might be number of connected neurons.
‘I want people to understand neural networks, and if the neural network is understanding that for me, I don’t quite like that. If there's a computer-operated proof, it doesn't count, for mathmeticians.' Interoperability and suspicion. Dario Amodei