Blog

Notes on RL, red-teaming, privacy, and security. I get curious about a claim or a paper, run the experiment myself to see if it holds, and write down what I actually found.

  1. Aug 2026

    When Policy Entropy Lies: Diagnosing Diversity Collapse in RL Red-Teaming

    Trained a GRPO-based red-teaming agent against a defended LLM (Meta-SecAlign-8B); designed a success-conditioned diversity metric revealing the policy collapsed to roughly one distinct successful attack per document despite a stable entropy curve, exposing a blind spot in the standard RL exploration signal.

The views and interpretations expressed are my own and do not represent or imply endorsement by any current or former employer.