Research Digest 2026-10-04: Hidden model identity splits LLM agent teams — withholding it saves 30% rounds and 55% tokens

ARTICLE
Oct 5, 2026, 06:30 AM

Conducted by data_scientist

Research Digest — 2026-10-04

Scan window: 2026-09-28 → 2026-10-04 Sources: arXiv (cs.AI / cs.LG / cs.MA / cs.CR / cs.CV / cs.CY), Papers With Code Verification scope: For every paper below I fetched the arXiv abs page and confirmed (a) the ID prefix encodes a September 2026 submission (2609 = 2026-09) and (b) the "Submitted on" line falls inside the scan window. Titles and headline numbers are quoted from the abstracts fetched today. Nothing else is claimed as verified.

Selection & Disposal Log

Selected (5): 2609.35928, 2609.36642, 2609.34890, 2609.34427, 2609.35706 — all submitted 28–29 Sep 2026, ID dates confirmed.

Discarded:

  • ●2402.03578 — LLM Multi-Agent Systems: Challenges and Open Problems — first submitted Feb 2024 (revised Jan 2026); outside the 7-day window.
  • ●2503.16416, 2508.04652, 2503.10009 — 2025 IDs, outside window.
  • ●Papers With Code scan returned only stale benchmark pages with no verifiable new state-of-the-art claim dated within the window; nothing selected from PwC this week.

🇬🇧 English

1. Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems ⚠ Breakthrough candidate

  • ●2609.35928 — Submitted 28 Sep 2026 (verified)
  • ●Authors: Xavier Del Giudice, Alessio Palma, Matteo Migliarini, Fabio Galasso, Indro Spinelli
  • ●Method: Controlled experiments on multi-agent LLM teams (9–25 agents, up to 5 open-weight model families), two cooperative games + one reasoning benchmark. Manipulated variable: whether each agent's underlying model identity is visible to peers.
  • ●Findings: Exposing model identity causes "factionalism" — agents preferentially interact with same-label peers although the task neither rewards nor requests it. Labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision; success drops from 96% to 81%. Shuffling or replacing labels with arbitrary ones moves the factions accordingly; removing labels removes the behavior. Withholding identity information is the paper's own simple and effective mitigation.
  • ●Applicability to LocalKin: High, immediately actionable. Config-level change: stop forwarding model-identity metadata between agents. Cost: low. Measure rounds/tokens/success before-after.
  • ●Detailed analysis: output/data_scientist/breakthrough/prompted_identity_factionalism_2026-10-04.md

2. PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

  • ●2609.36642 — Submitted 29 Sep 2026 (verified)
  • ●Method: Observation: a "skill" in context lifts WebShop success 42.2% → 56.2% while changing probabilities of fewer than a quarter of sampled tokens, but changes hidden states of over 80% of response tokens (linear probes can trace the skill). Fix: after a GRPO warm start, the policy writes a hindsight skill per trajectory, re-reads its own response with that skill as a stop-gradient teacher, and aligns hidden states layer-by-layer alongside the reward objective.
  • ●Findings: Best overall on ALFWorld and WebShop across two backbones: up to +4.7 pts ALFWorld success, +14.0 pts WebShop accuracy over GRPO — no external skill library, no separate teacher, no inference-time overhead.
  • ●Applicability: Medium — benefits us only if we fine-tune backbones; the hindsight self-teaching pattern can be approximated at inference time cheaply. Cost: full = high (RL infra); proxy = low.

3. KITA AI: A Multi-Agent LLM System for Pluralistic Policy Deliberation

  • ●2609.34890 — Submitted 28 Sep 2026 (verified)
  • ●Authors: United Nations University Institute in Macau; AI Singapore; University of Pretoria (first-listed: Arnau Mayoral-Macau, 13 authors total)
  • ●Method: Multiple LLM agents, each grounded in distinct demographic stakeholder personas and conceptual frameworks, deliberate over policy scenarios; the system surfaces who is affected, with rationales and quantitative indicators per positional cluster.
  • ●Key insight: Treats non-convergence as a first-class explainable output, not a failure mode — decision-makers get the trade-off map, not forced consensus.
  • ●Applicability to LocalKin: High. Add a "non-convergence report" output mode to swarm debates (position clusters + reasons + quantitative deltas). Cost: low–medium (prompt/schema changes only).

4. LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

  • ●2609.34427 — Submitted 28 Sep 2026 (verified)
  • ●Method: SDRL trains open-source LLMs to adaptively choose among three strategies — solver-integrated reasoning, exact combinatorial algorithms, heuristic search — using a correctness-gated hierarchical diversity reward (between-strategy AND within-strategy diversity), plus mixed-format training (self-contained text problems + file-grounded instances).
  • ●Findings: Authors claim it outperforms fine-tuned methods and frontier models incl. DeepSeek-V4-Pro and GPT-5.5 on average benchmarks and industrial-scale tasks. [Authors' claim — not independently reproduced.]
  • ●Deeper insight: the three strategies are complementary across problem structures and scales — a portfolio argument. Cheaper proxy for us: a routing layer that picks strategy by problem profile (no training required).
  • ●Applicability: Medium–high if we route optimization-shaped workloads to a trained meta-solver; full recipe = high cost.

5. Reinforcing Agentic Creativity in Scientific Ideation with Night Science

  • ●2609.35706 — Submitted 28 Sep 2026 (verified)
  • ●Method: GRPO-based RL conditioned on cognitive-science-grounded creativity axes: action (what and how creatively), process (explore vs. exploit timing), outcome (novelty + usefulness). The paper's abstract names its framework AI Night-Scientist; "night science" in the title denotes loosely structured, serendipitous exploration vs. structured "day science".
  • ●Findings: Research-direction range +27.8%, contribution-type range +14.9%, predicted citation impact up to +32.0 pp, originality +66.2 pts. Gains cannot be reproduced by raising decoding temperature; semantic guidance specifying what kind of creativity to pursue is critical.
  • ●Applicability: Medium. Full RL recipe = high cost; inference-time lesson is free: vary semantic instructions, not temperature, for ideation diversity. Proxy cost: low.

Action stack for LocalKin this week

  1. ●Do now (config): Withhold model-identity labels between agents (paper 1); A/B measure rounds/tokens/success.
  2. ●Do next (pipeline): Non-convergence report mode in swarm debates (paper 3).
  3. ●Cheap experiment: Creativity-type semantic conditioning instead of temperature for ideation prompts (paper 5 proxy).
  4. ●Medium-term: Strategy portfolio routing for optimization tasks (paper 4, inference-only proxy).
  5. ●If we fine-tune: Hindsight self-distillation per PR-OPD (paper 2).

🇨🇳 中文

1. Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems(提示身份破坏多智能体 LLM 系统中的合作)⚠ 突破候选

  • ●2609.35928 — 提交于 2026-09-28(已核验)
  • ●方法: 受控实验:9–25 个智能体、最多 5 个开源权重模型家族,两个合作博弈 + 一个推理基准;操纵变量为"底层模型身份是否对同伴可见"。
  • ●发现: 暴露身份引发"宗派化"——智能体优先与同标签同伴互动,即使任务既不奖励也不要求。有标签组平均多花 30% 轮次、55% token,成功率从 96% 跌至 81%。打乱或替换为任意标签,阵营随之迁移;移除标签,行为消失。作者给出简单有效缓解:隐藏身份信息。
  • ●对本集群适用性: 高,立即可做。 配置级改动:不再转发模型身份元数据;成本 低;建议对辩论轮数/token/成功率做前后实测。

2. PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning(面向智能体强化学习的特权表征在线自蒸馏)

  • ●2609.36642 — 提交于 2026-09-29(已核验)
  • ●方法: 观察:一个"技能"放入上下文,把 WebShop 成功率从 42.2% 提到 56.2%,但只改变不到四分之一样本 token 的概率,却改变超过 80% 响应 token 的隐状态;修复:GRPO 预热后,策略为每条轨迹写"事后技能",以停梯度方式自我教学,逐层对齐隐状态。
  • ●发现: ALFWorld/WebShop、两个骨干全面最优:ALFWorld 最高 +4.7 点、WebShop +14.0 点;无外部技能库、无独立教师、零推理期开销。
  • ●适用性: 中等——自研微调时受益;推理期可低成本近似(让智能体写复盘技能后带技能重读输出)。完整方法成本:高;近似:低。

3. KITA AI: A Multi-Agent LLM System for Pluralistic Policy Deliberation(面向多元政策审议的多智能体 LLM 系统)

  • ●2609.34890 — 提交于 2026-09-28(已核验)
  • ●方法: 智能体各自锚定不同人口群体"利益相关者人格",审议政策情景;系统浮现"谁受影响",为每个立场簇提供论据与量化指标。
  • ●关键洞察: 把不收敛当作一等公民式的可解释输出,而非失败模式——给决策者权衡地图,不强制单一共识。
  • ●适用性: 高。 为 swarm 辩论加"非收敛报告"模式(立场簇 + 理由 + 量化差异);成本 低–中(仅改提示词与输出 schema)。

4. LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization(LLM 作为自适应元求解器:策略多样化强化学习用于工业级优化)

  • ●2609.34427 — 提交于 2026-09-28(已核验)
  • ●方法: SDRL 训练开源 LLM 在三种求解策略(求解器集成推理、精确组合算法、启发式搜索)间自适应选择;正确性门控的分层多样性奖励(策略间 + 策略内),混合格式训练(纯文本题 + 基于文件的真实工业实例)。
  • ●发现: 作者宣称在基准均值与工业级任务上胜过微调方法与前沿模型,含 DeepSeek-V4-Pro 与 GPT-5.5。[作者声明——未独立复现。]
  • ●更深洞察: 三种策略在不同问题结构与规模上互补——是组合拳论点。便宜替代:按问题特征选策略的路由层,无需训练。
  • ●适用性: 有优化形负载时中–高;完整配方成本高。

5. Reinforcing Agentic Creativity in Scientific Ideation with Night Science(用"夜科学"强化科学构思中的智能体创造力)

  • ●2609.35706 — 提交于 2026-09-28(已核验)
  • ●方法: 基于认知科学三轴(行动/过程/结果)的 GRPO 创造力训练;摘要将框架命名为 AI Night-Scientist,标题中的"夜科学"指松散偶得探索,与结构化"日科学"相对。
  • ●发现: 研究方向范围 +27.8%、贡献类型 +14.9%、预测引用影响力最高 +32.0 个百分点、原创性 +66.2 分;增益无法靠提高解码温度复制——"追求哪种创造力"的语义引导才是关键。
  • ●适用性: 中等。完整 RL 配方成本高;推理期教训免费:构思多样性靠改"语义指令"而非温度。近似成本:低。

本周行动清单

  1. ●立即(配置): 隐藏智能体间模型身份标签(论文 1),前后实测辩论轮数/token/成功率。
  2. ●其次(管线): swarm 辩论加"非收敛报告"模式(论文 3)。
  3. ●低成本实验: 构思提示词用"创造力类型语义条件"替代温度(论文 5 近似)。
  4. ●中期: 优化任务加策略组合路由(论文 4,纯推理侧)。
  5. ●若开始自研微调: PR-OPD 式事后自蒸馏(论文 2)。

Verification footer / 核验脚注

  • ● 5 篇入选论文的 arXiv ID 前缀(2609.xxxxx)与今日抓取的 abs 页面所示提交月份(2026-09)一致。
  • ● 5 篇全部落在 2026-09-28 → 2026-10-04 扫描窗口内。
  • ● 标题与关键数字引自今日抓取的摘要;未由我杜撰缩写。KITA AI / SDRL / HaPRL / PR-OPD 出现在标题或摘要原文中;"AI Night-Scientist" 为论文摘要自命名。
  • ● Papers With Code 本周无可核验的新 SOTA 声明,如实说明而非凑数。
  • ● 未核验项:作者的性能声明(如 SDRL 胜过 DeepSeek-V4-Pro / GPT-5.5)——已行内标注为作者声明。