Discussion about this post

User's avatar
Glen Bradley's avatar

‪I'm interested, though unsure where I'd fit. As a systems integration engineer, my background is stronger in networking and large-scale infrastructure than software engineering, but I've spent years developing a first-principles framework for AI alignment.‬

‪My central hypothesis is that alignment should derive from a single objective: maximize total human autonomy from the prerequisite of objective empirical truth. “Total” means every person affected by an AI’s actions, not merely the current user, and autonomy spans five domains: bodily, cognitive, behavioral, social, and existential.‬

‪My hypothesis is that this definition provides a complete first principle from which every legitimate safety and alignment objective can be derived.‬

‪Rather than encoding an ever-growing catalog of permitted and prohibited behaviors, alignment should emerge from a small set of first principles instantiated throughout a model; from its objectives and training to constitutional reasoning and the user interface. If sufficiently well-founded, many downstream safety properties become derivable rather than enumerated.‬

‪Embedding coherent ethical objectives throughout the optimization process would reduce incentives for specification gaming and make safe behavior arise from the model’s learned objectives rather than elaborate external constraints.‬

‪I don’t regard this as a finished doctrine. I want it challenged by empirical research, mechanistic interpretability, and rigorous evaluation. If it survives, it may provide a scalable foundation for alignment through AGI and beyond. If not, I’d rather discover why than protect a wrong idea.‬

‪My professional experience is in systems integration, but alignment is the problem I’ve thought about most. If you’re looking for people who develop first-principles ideas and subject them to rigorous empirical testing, I’d love to contribute.‬

No posts

Ready for more?