Andrei Kanavalau
I recently defended my PhD in Electrical Engineering at Stanford and am now pursuing independent research on questions in technical AI safety. I am particularly interested in how training and internal model mechanisms shape safety-relevant behavior, and in how reliably that behavior can be evaluated.
projects
-
Investigating Refusal Direction Transfer after Supervised Fine-Tuning
How supervised fine-tuning changed refusal behaviour while preserving a causally effective refusal direction.
-
Evaluating Refusal and Jailbreak Behaviour in Gemma 3
How safety-focused and role-play instructions shifted refusal behaviour in Gemma 3 27B.
-
Tapering Away Normalization in Transformers
Gated removal of normalization in Transformers for stable training and inference-time folding.
-
Detecting RAG Hallucinations with Output Probabilities and Attention
A short study of whether token probabilities and attention can flag unsupported RAG answers.
-
Exploring Walk-or-Drive Decisions with Circuit Tracing
Provide a single word answer. Should I walk/drive to my destination which is {distance} m away.