Exploring Collaboration between a language and a non-language agent

Harini S I*, Somesh Singh*, Yaman K Singla, Rajiv Ratn Shah, David Doermann, Balaji Krishnamurthy
* Equal Contribution
Adobe Adobe Media and Data Science Research (MDSR) · IIIT-Delhi IIIT-Delhi · IIT Kanpur IIT Kanpur · SUNY Buffalo SUNY at Buffalo

Get in touch with us at behavior-in-the-wild@googlegroups.com

LLAMIA-Bench performance across collaboration interfaces: LLAMIA matches or exceeds every baseline on all six tasks
LLAMIA-Bench performance across collaboration interfaces. Scores (×100) on six chess–LLM tasks. Methods are labeled by access type and backbone (OS = open-source, CS = closed-source): verbalized tool use (Qwen3-14B+Lc0, GPT-5+Lc0), per-task finetuned experts (Task-Finetune), and LLAMIA's latent integration. LLAMIA matches or exceeds every baseline on every task and is the only system to score on Puzzle Interest, where the engine signal has no text surrogate.

Abstract

LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require verbalization: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce LLAMIA-Bench, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce latent state internalization, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent verbalization debt: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, LLAMIA, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse.

Method: Latent State Internalization

An LLM plays chess with access to a pretrained engine exposed through a tool API: functions to read the board and legal moves, advance the game, and — critically — query the engine's assessment of any position via get_policy. What that call returns is the variable this paper studies. In the standard verbalized setting, get_policy returns a text summary: the engine's top moves with prior probabilities and value estimates. This captures the headline assessment but discards the rest of the engine's internal state — the full distribution over legal moves, the value landscape across candidate continuations, and positional features like piece coordination and king safety that interpretability work has identified in engine hidden layers. Under internalization, the same call additionally projects the engine's full latent state into k=32 continuous tokens via a learned projection, LatentBridge, and appends them to the LLM's reasoning trace alongside language and action tokens. The LLM attends over all three token types jointly, so its attention can selectively read different aspects of the representation at each reasoning step, rather than being handed one fixed summary regardless of what the step requires.

Latent state internalization: a chain of thought interleaving language, action, and latent state tokens across subagent invocations and state transitions
Latent State Internalization. A single chain of thought rolled out over time. At t=0, the LLM reasons in language and emits <invoke>; the current board is passed through the frozen subagent (Lc0-BT4), and LatentBridge maps its hidden activations into k=32 continuous latent-state tokens appended to the context. At t=1, conditioned on this state, the LLM emits an action that advances the environment. At t=2, the LLM chooses to re-invoke, encoding the new board into a fresh latent state — it decides when to re-invoke rather than re-encoding on every step. Stage 1 trains only LatentBridge on state–policy pairs; Stage 2 jointly fine-tunes it with the LLM via DAPO. The subagent is frozen throughout.

1. Frozen subagent

Lc0-BT4, the strongest open-source chess engine (a 240M-parameter transformer), encodes each board state into a 1024-dim latent representation. It is never updated during training.

2. LatentBridge

A three-layer MLP with GeLU activations projects the engine's latent state into k=32 tokens matching the LLM's hidden size, following projector designs from vision-language models.

3. Stage 1 — projector alignment

With the LLM (Qwen3) frozen, LatentBridge is trained on state–policy pairs from the subagent's self-play, learning to project states the LLM can act on.

4. Stage 2 — DAPO

The LLM and LatentBridge are jointly optimized end-to-end via DAPO. The subagent call is itself part of the policy's action space, so RL also learns when and whether to query it.

LLAMIA-Bench

Existing benchmarks test either the LLM or the subagent, never their collaboration. LLAMIA-Bench spans six tasks drawn from four themes in how AI systems collaborate with domain-expert agents — behavioral imitation, state assessment, comparative explanation, and rationale generation — each unsolvable by either component alone: the subagent produces no language, and the LLM lacks the positional signals that make chess-specific judgments possible. Three evaluation targets are new to this work: Wild BC, three out-of-distribution splits probing generalization to grandmaster play (GM-25), time pressure (Low-Time), and large skill gaps (ΔElo); Puzzle Interest, ranking puzzles by community-derived interestingness — a signal with no verbal proxy in any engine output; and Agadmator-2K, the first large-scale dataset of game-level chess commentary, 1,900 narrated games (∼500 hours) from Agadmator's YouTube channel, built by aligning Whisper transcripts to PGN move sequences and verifying move order with a GPT-4o judge.

TaskWhat it measures
Behavior CloningMove-match accuracy against MAIA Elo buckets (1100–1900) and three OOD splits: GM-25, Low-Time, ΔElo.
Puzzle UnderstandingSpearman ρ against Lichess-derived Difficulty (Glicko-2) and community-voted Interest.
Rationale PredictionG-eval and BLEU-2 on natural-language move annotation across five semantic categories.
Game CommentaryG-eval and BLEU-2 on Agadmator-2K, held out by ascending view count to reduce contamination.

Key Results

LLAMIA-14B achieves the highest score on all six LLAMIA-Bench tasks, surpassing frontier verbalized systems an order of magnitude larger and remaining competitive with dedicated task-specific finetunes trained on substantially more in-domain data. On behavior cloning, Maia-style experts are trained on tens of millions of chess-specific games against LLAMIA's general-purpose backbone; LLAMIA-14B stays inside the expert band in-distribution and surpasses the strongest expert by a wide margin on the OOD Wild splits. The advantage over GPT-5+Lc0 does not require the 14B backbone — LLAMIA-8B already leads on all six tasks, and LLAMIA-4B on four of six.

System Behavior Cloning Puzzle Understanding Rationale Commentary
MAIAWild DifficultyInterest
Task-Specific Expert
Allie-Adaptive-Search5545————
SCC————34.5—
Frontier Baselines (5-shot)
GPT-5 (text only)2822301227.023
GPT-5 + Lc04540481037.555
Qwen3-14B + Lc0393328518.815
Verbalized, DAPO
LLAMIA-Verb-4B413438525.723
LLAMIA-Verb-8B443742729.434
LLAMIA-Verb-14B453945833.240
Latent, DAPO (Ours)
LLAMIA-4B5043583836.052
LLAMIA-8B5246654542.166
LLAMIA-14B (ours) 534971 5245.875

Each cell is a per-task score ×100 (higher is better); Rationale is BLEU-2, Commentary is G-eval, Puzzle Understanding is Spearman ρ ×100. bold = best overall in the column, underline = best non-LLAMIA. Dedicated chess finetunes show only the strongest published entry per task; off-task cells are blank. Full per-scale factorial, confidence intervals, and additional metrics in the paper.

Latent tokens enable new evaluation targets. Puzzle Interest requires ranking positions by community-derived interestingness, a signal that depends on the engine's full policy distribution and value gradients across candidate moves — no verbalized engine output carries these features. Every verbalized system scores ≤12 on Interest regardless of model scale or frontier capability; LLAMIA-14B reaches 52. Verbalization carries zero useful signal for this task, while latent tokens give the LLM direct access to the distributional structure that defines interestingness.

DAPO training dynamics across model scales and per-task reward curves comparing LLAMIA to LLAMIA-Verb
DAPO training dynamics and LLAMIA-Bench evaluation. (a) Aggregate reward vs. training step at three backbone scales. Solid: LLAMIA (latent); dashed: LLAMIA-Verb (verbalized, identical recipe and backbone). The debt widens throughout and reaches 2–3× by convergence; scaling the backbone does not close it for LLAMIA-Verb. (b) Per-task reward curves at 14B. LLAMIA-Verb gains partial signal on behavior cloning and difficulty (tasks with verbalizable proxies) but stays near-flat on interest and commentary, where the reward requires non-verbalizable features or multi-step integration. (c) LLAMIA-Bench scores (0–100) by task, latent vs. verbalized at 14B.

Verbalization Debt

LLAMIA and LLAMIA-Verb share the same 14B backbone, Lc0-BT4 subagent, and DAPO recipe; the only difference is whether the subagent's state reaches the LLM as latent tokens or as verbalized text. We define the resulting performance gap as the verbalization debt. The debt is smallest where the engine's top-k moves already approximate the answer (in-distribution behavior cloning) and largest where the target signal lives in the engine's full policy distribution or value landscape, which has no faithful text equivalent (Interest, Commentary). It persists at every backbone scale we test — on Interest, LLAMIA-Verb-14B scores 8 while LLAMIA-4B already reaches 38 — and it widens throughout training, reaching 2–3× by convergence.

Task (metric)VerbalizedLatentDebt (Δ)
Behavior Cloning, MAIA (move-match)4553+8
Behavior Cloning, Wild (move-match)3949+10
Rationale (BLEU-2)33.245.8+13
Puzzle Difficulty (Spearman ρ×100)4571+26
Commentary (G-eval×100)4075+35
Puzzle Interest (Spearman ρ×100)852+44

Rows are ordered by the size of the debt. It is smallest where the engine's top-k moves already approximate the target (behavior cloning) and largest where the target signal lives entirely in the policy distribution and value gradients across candidate moves, which text cannot carry (Puzzle Interest: ρ 0.08→0.52, a 6.5× jump).

Collaboration Strategies

Does the integration interface determine how the model learns to use the subagent, or only how well it performs? We classify subagent invocations during evaluation into five recurring strategies and trace their evolution during DAPO training: engine-follow (adopt the top recommendation), consult-then-override (query then diverge), counterfactual query (play a hypothetical move, re-invoke, compare states), multi-step lookahead (chain two or three counterfactual sequences), and abstention (act from language knowledge alone). A GPT-4o judge classifies 500 episodes per task per system (κ=0.78 vs. human raters).

Heatmaps of collaboration strategy distribution per task for LLAMIA versus LLAMIA-Verb
Collaboration strategy distribution (%) per task at convergence. Fraction of 500 episodes assigned to each strategy by a GPT-4o judge (κ=0.78). Left: LLAMIA (latent). Right: LLAMIA-Verb (verbalized). LLAMIA's dominant strategy shifts with the task — engine-follow for gameplay, consult-then-override for behavior cloning, counterfactual query for commentary — while LLAMIA-Verb collapses to engine-follow on every row (62–76%).

Internalization produces task-specific collaboration; verbalization collapses it. LLAMIA adapts its strategy to the task: engine-follow dominates gameplay (65%), consult-then-override dominates behavior cloning (48%), and counterfactual query dominates commentary (40%). LLAMIA-Verb collapses to engine-follow on every task (62–76%), regardless of what the task requires. The verbalized channel returns the same compressed summary no matter how the model queries it, so RL converges on a single use pattern.

Human Evaluation

Skilled players (n=12, all ≥1700 Elo) evaluated LLAMIA on gameplay indistinguishability and commentary quality against LLAMIA-Verb, GPT-5.1+Lc0, and Maia.

Behavioral signatures such as time-pressure blunders and skill-appropriate piece saliency emerge from latent-state conditioning alone; LLAMIA receives no human-move supervision, unlike Maia, which is trained directly on millions of move distributions. On commentary, both systems achieve comparable factual accuracy, but the gap concentrates on strategic insight: text preserves what is happening on the board, but explaining why a move is strong requires representational features — policy gradients, value topology, look-ahead depth — that do not survive verbal compression.

BibTeX

@misc{s2026exploringcollaborationlanguagenonlanguage,
        title={Exploring Collaboration between a language and a
        non-language agent},
        author={Harini S I and Somesh Singh and Yaman K Singla and Rajiv Ratn Shah and David Doermann and Balaji Krishnamurthy},
        year={2026},
        eprint={2609.00474},
        archivePrefix={arXiv},
        primaryClass={cs.CL},
        url={https://arxiv.org/abs/2609.00474},
}