Blog
Notes on RL, red-teaming, privacy, and security. I get curious about a claim or a paper, run the experiment myself to see if it holds, and write down what I actually found.
-
Aug 2026
When Policy Entropy Lies: Diagnosing Diversity Collapse in RL Red-Teaming
Trained a GRPO-based red-teaming agent against a defended LLM (Meta-SecAlign-8B); designed a success-conditioned diversity metric revealing the policy collapsed to roughly one distinct successful attack per document despite a stable entropy curve, exposing a blind spot in the standard RL exploration signal.
The views and interpretations expressed are my own and do not represent or imply endorsement by any current or former employer.