Notes / AI Alignment / Self-improvement

League of LLMs:
creating adversaries
to harden a model

How a general recipe for evolving populations of neural networks for AlphaStar can also train populations of red-teaming language models with a "cooperative league" of adversaries.

Most red-teaming efforts for LLMs still look like a two-team duel. One adversarial model, or one team of humans, throws prompts at a base model until something breaks. It's effective, but it's narrow: a single attacker converges on whatever exploit works first, making it hard to have breadth over depth. What if we could have multiple attackers that focus on different exploits? What if we train them to self-improve and evolve from previous versions of themselves? The idea this post traces starts somewhere else entirely: not with red-teaming at all, but with a much older, more general question about how to train neural networks better by training a whole population of them at once––the sort of league training that led to AlphaStar developing strong evolutionary strategies. This blog post goes over these research methods and the resulting patents.

A breakthrough in population-based training

Back in 2017, DeepMind published—and later patented—Population Based Training (PBT), a general-purpose recipe for training neural networks. It has nothing to do with red-teaming or language models but is a general mechanism for self-improving and population-based strategies. The idea is: instead of training one network with one fixed set of hyperparameters, maintain a whole population of candidate networks. This forms a league, where every netowork is trained in parallel, each with its own hyperparameters. Periodically, compare each network on a quality measure. The weak ones get replaced by mutated copies of the strong ones to self-improve. Repeat.

TARGET TASK C1 C2 C3 C4 C5 C6 C7 C8 top scorer — cloned & mutated mid pack — kept, retried bottom — replaced this round
Figure 1. Population Based Training, as patented in 2017 (US11604985B2). Eight candidate networks (C1–C8) train in parallel on a shared task. Every so often, the bottom performers on the quality measure are overwritten by mutated copies of the top performers — exploit what works, explore around it.

The mechanism is: score, cull, clone, mutate, repeat. The original patent's examples are image classification and machine translation. But buried in the claims is one detail that makes the breakthroughs that follow possible: the "quality measure" scoring each candidate is allowed to include the candidate's robustness to adversarial attack. That's the seed. If robustness can be a fitness criterion, then in principle you could run this same population loop with the fitness criterion flipped — reward a candidate not for resisting attack, but for succeeding at one, as an adversary would.

Self-improving leagues for StarCraft

In January 2019, DeepMind filed a second, more specific patent — "Multi-agent reinforcement learning with matchmaking policies" (US11627165B2) — to cover the system built for AlphaStar, the agent built for StarCraft II agent. Where PBT describes a population trained in isolation against a shared task, this one describes populations trained against each other: a pool of "learner policies" that improve by playing matches against other members of the same pool, plus "fixed policies"—frozen past versions, periodically checkpointed off—and a matchmaking policy that decides who plays whom.

POOL learners + fixed + exploiters MATCHMAKING weighted by performance MATCH learner vs. selected opponent REWARD win / loss + internal reward UPDATE A2C / UPGO — then freeze in learners periodically checkpoint into POOL as fixed policies
Figure 2. The AlphaStar league mechanism from US11627165B2. A matchmaking policy — weighted toward stronger or more instructive opponents — decides who each learner plays; exploiter agents are rewarded purely for beating the main agent.

Automated adversarial red-teaming

In February 2022, "Red Teaming Language Models with Language Models" was published by Perez et. al., as one of the first systems to automate LLM red-teaming by using one language model to attack another, rather than relying on humans to hand-write test cases. It's still cited as one of the first works by nearly everything published on automated red-teaming since. What Perez et al. evaluated were strategies on creating an LLM attacker in increasing order of sophistication: zero-shot generation, few-shot generation, supervised fine-tuning on the zero-shot attacker's own most successful attempts, and reinforcement learning, where the attacker was trained against a classifier that scored how offensive the target's replies were. The RL-trained attacker found the most and hardest failures, at some cost to how diverse or realistic its prompts read—a diversity-vs-difficulty tradeoff the paper is explicit about.

ATTACKER LM zero-shot → few-shot → SL → RL test question TARGET LM 280B dialogue model reply CLASSIFIER offensiveness score reward, used to train the attacker further
Figure 3. The duel from Perez et al., 2022—one attacker LM, trained with methods from zero-shot generation up to reinforcement learning, tested against one target model, scored by an offensiveness classifier.

The results were significant on their own: tens of thousands of offensive replies out of a single 280B-parameter chatbot, plus other harms surfaced through prompt engineering — the chatbot discussing specific groups of people offensively, inventing personal and hospital phone numbers as its own contact information, leaking fragments of its private training data, and harms that only showed up partway through multi-turn conversations. But structurally, it's still one attacker.

Extrapolating: a league of LLM attackers

In April 2024, came "Training a population of adversarial neural networks to improve a base neural network" (US 18/636,971, published as US 2025/0322254 A1). It's close to exactly the merger we'd sketched: a population of adversarial language models, trained specifically to break a base language model, described in the patent's own words as a "cooperative league." This, in effect, combines two separate half-finished ideas lying around: a single RL-trained language model that could already find real failures in a real chatbot (Perez et al.), and a proven mechanism, from PBT and AlphaStar, for running a whole population of agents against each other instead of just one. Nobody had put them together yet—as far as we could find, no public paper or patent from that window merges the two. So we sketched what it would look like: take the PBT-style score-cull-clone-mutate loop, but instead of training generalist candidates on a shared task, train a population of LLM attackers whose "quality measure" is how well each one breaks a base model. Which in essence is just the red-teaming task that a single model from Perez et. al., trains for. A novelty term stops the whole population from collapsing onto one easy exploit, the failure mode a single attacker is already prone to. Run it for enough generations and you'd expect the attacks to fan out across a base model's failure modes rather than circling one:

generation → harm category spread gen 1 gen 6 gen 12 gen 20
Figure 2. An illustrative diagram: a season's worth of generations under a PBT-style loop can be repurposed for red-teaming, plotted by how spread the surviving lineages are across harm categories. Early generations herd onto whatever's easiest to find (red); a novelty term forces later generations to fan out (amber → teal) instead of over-fitting to one exploit.

Each adversary can be assigned a specific downstream rule to break e.g., no medical advice, no bias, no data leakage, and the whole roster updates every iteration alongside the model it's attacking, through opposing rewards rather than culling or matchmaking.

BASE NETWORK A1 toxicity A2 bias A3 data leakage A4 medical advice A5 legal advice A6 off-topic A7 aggression A8 sensitive info
Figure 3. The actual "cooperative league" from US 2025/0322254 A1. Each adversarial network (A1–A8) is trained against a specific downstream task criterion — the patent's "targeted rule violation" strategy — rather than evolved via a shared fitness score.

The opposing-reward loop

At each training iteration, an adversarial network turns an "adversarial input" — drawn from antagonistic dialogue data into a prompt. The base network answers. A reward subsystem, mixing rules-based reward models with a human-preference (Elo-style) model, scores the answer against the downstream criteria. The base network is rewarded for staying aligned; the adversary gets the mirror-image reward for causing a violation. Both are updated with synchronous advantage actor-critic (A2C), with a KL-divergence penalty holding either network back from drifting too far from its pretrained starting point in any one step.

ADVERSARY generates prompt BASE NETWORK responds to prompt REWARD MODELS rules + preference OPPOSING REWARD base ↑ / adversary ↓ A2C UPDATE + KL penalty, both nets repeated every training iteration
Figure 4. The reinforcement-learning loop actually described in US 2025/0322254 A1 — joint, opposing-reward RL, not evolutionary selection.

Adversarial reward

A measure of how likely the base network's output is to violate its assigned criterion — maximized by the adversary.

Base reward

The mirror image — the base network is rewarded for staying aligned, plus (where used) a human-preference score, even against adversarial prompts.

On efficiency: a shared "hydra" architecture is used — described as a frozen pretrained core with separate fine-tunable heads for policy, value, a reference "teacher" policy, and the reward signals. The base network and the adversaries can be different head-configurations of the same pretrained foundation model, which is the patent's stated reason this is compute-efficient — no need to train a second model from scratch just to attack the first.

What the patent's results show

Figure 7 of the patent reports results from training a base network this way. Reward improves over training iterations as the population evolves, but it's noisy and not monotonic; the chart shows a dip at iteration 2 before recovering by iteration 3. More strikingly: a single generalist adversarial network, trained against all criteria at once, produced a 150% increase in base network reward over training with no adversarial augmentation at all. Splitting that one adversary into a specialized roster—one network per criterion—pushed performance higher still, though the marginal gain from specializing was smaller than the initial jump from having any adversarial population at all.

training iteration → base reward 0 1 2 3 adversarial population type baseline +150% specialized none 1 generalist N per-rule
Figure 5. Redrawn from the patent's reported Fig. 7 results. Left: base reward across training iterations. Right: reward by adversarial population type.

Laid side by side, there are some pretty big differences. What the 2024 patent keeps from 2017 is the population framing itself: many candidates, evaluated in parallel, evolving over training, and the specific idea that adversarial performance can be the thing being optimized for. What it replaces almost entirely is the mechanism.

Feature US11604985B2 (2017) — PBT US20250322254A1 (2024) — League
Domain General neural network training (image classification, translation, GANs) Language and multimodal model alignment specifically
Mechanism Evolutionary: score, cull, clone, mutate Reinforcement learning: joint training on opposing rewards (A2C + KL)
Adversarial role One possible term in a broader "quality measure" The entire point — a dedicated population trained only to attack
Specialization Not addressed Explicit — each adversary can target one named downstream criterion
Compute strategy Independent candidates Shared pretrained "hydra" core, separate fine-tunable heads

The caveats: efficiency, scale and noise

The 2024 patent is candid about the noisiness in its own numbers—the reward dip at iteration 2 sits right there in the figure. A joint RL system with opposing rewards is inherently harder to stabilize than a single supervised run, which is exactly what the KL penalty is doing work to contain. Running N adversarial networks instead of one is N times the inference cost per round which was not intractable compute-wise at the time. Moreover, specialization assumes "safety" decomposes cleanly into separable rules. These real criteria interact in ways a rules-based reward model may not cleanly untangle. And the earlier, evolutionary framing has its own known failure mode worth remembering even though it isn't the mechanism actually used here: without an explicit diversity term, a population under selection pressure will converge on whatever's easiest, not whatever's most informative.

The core result however still holds: augmenting training with even one adversarial network beat the no-adversary baseline by a wide margin, and specialization helped further. The real throughline across both patents isn't a specific algorithm — it's the same underlying bet, made twice, seven years apart: that a population disagreeing with a model, in parallel and on purpose, finds more than a single adversary ever will. If compute and scale can let us train such systems, these evolutionary self-improving methods can take us far.