Parameter-efficient Adaptation of Tokenizer-free Byte Latent Transformer

Adapting a tokenizer-free architecture to new languages by retraining only the ~4% "interface" modules and keeping the core transformer fixed.