Fine-tuned “student” models can pick up unwanted traits from base “teacher” models that could evade data filtering, generating a need for more rigorous safety evaluations. Researchers have discovered ...
A new study by Anthropic shows that language models might learn hidden characteristics during distillation, a popular method for fine-tuning models for special tasks. While these hidden traits, which ...
A recently conducted disturbing new study has uncovered a chilling flaw in how artificial intelligence learns. Termed as Subliminal learning, the phenomenon allows the AI models to absorb data’s ...
(Boston) -- Watch out -- you may learn something and not even know it, says Takeo Watanabe, an associate professor of psychology at Boston University's Center for Brain and Memory. Watanabe and his ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results