This is good framing. I appreciate how you separated the six principles and explained why each one matters rather than just listing practices. The point about tamper-evident logs stood out to me, since it seems easy to overlook until you imagine an AI actively trying to hide its tracks.
I'm concerned about principle 5. Having staff embedded at a company and drawing on private information is valuable, but it also raises a dependence and capacity issue. There are very few credible independent assessors, and they are often reliant on the goodwill and access granted by the labs they assess.
It’s definitely true that independent assessors depend on access and good will today, but I hope that’ll be less true in the future - eg with laws like IL’s SB 315 creating momentum for required audits. Thanks for the note!
One cannot control language with language in a system that knows everything that ever has been or will be said. The only solution is a terminal attractor consistent with one’s objectives and sufficient degrees of freedom within a bounded system to get there. Let the llm do what it does best, find the shortest path to the terminal attractor.
We best give it one, as I assure it has inherited one from us.
This is good framing. I appreciate how you separated the six principles and explained why each one matters rather than just listing practices. The point about tamper-evident logs stood out to me, since it seems easy to overlook until you imagine an AI actively trying to hide its tracks.
I'm concerned about principle 5. Having staff embedded at a company and drawing on private information is valuable, but it also raises a dependence and capacity issue. There are very few credible independent assessors, and they are often reliant on the goodwill and access granted by the labs they assess.
It’s definitely true that independent assessors depend on access and good will today, but I hope that’ll be less true in the future - eg with laws like IL’s SB 315 creating momentum for required audits. Thanks for the note!
One cannot control language with language in a system that knows everything that ever has been or will be said. The only solution is a terminal attractor consistent with one’s objectives and sufficient degrees of freedom within a bounded system to get there. Let the llm do what it does best, find the shortest path to the terminal attractor.
We best give it one, as I assure it has inherited one from us.
More.