Andrei Kanavalau
I recently defended my PhD in Electrical Engineering at Stanford and am now pursuing independent research in technical AI safety. I am particularly interested in how training and internal model mechanisms shape safety-relevant behavior, and in how reliably that behavior can be evaluated.
projects
-
Investigating Model Organism Robustness to Fine-Tuning
Construction method impact on quirk persistence during untargeted supervised fine-tuning.
-
Refusal Direction Persistence Through Supervised Fine-Tuning
Supervised fine-tuning changes refusal behaviour while preserving a causally effective refusal direction.
-
Evaluating Refusal and Jailbreak Behaviour in Gemma 3
Behaviour shifts of Gemma 3 27B under safety-focused and role-play instructions.
-
Tapering Away Normalization in Transformers
Gated removal of normalization in Transformers for stable training and faster inference.
-
Detecting RAG Hallucinations with Output Probabilities and Attention
A short study of whether token probabilities and attention can flag unsupported RAG answers.
-
Parameter-efficient Adaptation of Tokenizer-free Byte Latent Transformer
Adapting a tokenizer-free architecture to new languages by retraining only the ~4% "interface" modules and keeping the core transformer fixed.
-
Circuit Tracing Walk-or-Drive Decisions
Provide a single word answer. Should I walk/drive to my destination which is {distance} m away