Joshua Penman
AI researcher
I work on how and what language models learn — what a model actually takes away from the data it is trained on, and how that can be steered.
Research
Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing
Finetuning a language model on documents explicitly labeled as fictional still leaves a model that believes them. Goggles is a learned module that intervenes on the finetuning gradient rather than the data, imparting a chosen epistemic frame to whatever the documents teach — raising correct identification of the content as fictional from roughly 9% to 91% while preserving capability.
Elsewhere
- ScholarGoogle Scholar
- LinkedInlinkedin.com/in/joshuasp
- GitHubgithub.com/JoshuaSP