跳至主要内容

基元律动

tokenrhythm.ai

All Models, One Mind

OpenSquilla 自动帮你选模型的搭档

装在你的电脑里,一句话完成调研、文件、文献与代码;智能路由与多模型协作,用更少 Token 做更多事。

OpenSquilla 任务场景

给一个主题,它自己去找、去核、去比 市场研究 · 竞品分析 · 赶汇报材料

更多关于OpenSquilla

模型即服务,让 Agent 用上多模型

少做接入和维护,把时间留给产品。模型选择、协作与用量管理,由一套接口完成。

开始使用
  • 一次接入主流模型

    不用逐家申请、充值和维护。

  • 每次都用对模型

    按任务需要,选择真正合适的能力。

  • 复杂任务自动协作

    路由与多模型融合,由系统自动完成。

  • 每一笔成本都清楚

    花在哪个模型、用了多少 Token,一目了然。

NeoHorse-1

让 Agent 的执行轨迹,成为模型持续进化的训练信号

NeoHorse-1 是面向文本 Agent harness 的开源模型家族,也是通往递归自我改进的初始原型。它通过路由引导的课程学习与在线策略蒸馏,让工具调用、代码与任务执行结果回流到后训练过程。

为不同部署规模准备的两个版本

4B 轻量级部署

NeoHorse-1-4B

64.87 十项评测平均

+5.93 vs 同尺寸 Qwen3.5

获取4B模型
9B 更高容量

NeoHorse-1-9B

69.04 十项评测平均

+3.44 vs 同尺寸 Qwen3.5

获取9B模型

9B · 跨任务评测

69.04 十项评测平均 · +3.44
NeoHorse-1-9B 与同尺寸基线在十项任务上的官方评测结果 查看完整技术报告

研究与进展

我们把路由、评测与 harness 设计的结论公开发表 —— 产品里跑的,就是论文里写的那一套。

01 NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness 2026
Agent Systems

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

NeoHorse-1 combines agentic post-training with routing-derived curricula and capability-guided data allocation to close an evaluation-selection-update loop toward recursive self-improvement.

NeoHorse Team, Guoliang Cao, Guohao Dai , et al. · arXiv Preprint ·

NeoHorse-1-9B benchmark results across agentic, coding, and instruction-following evaluations 阅读全文
02 OpenSquilla: Token-Efficient Agent = Models + Routing Harness 2026
Agent Systems

OpenSquilla: Token-Efficient Agent = Models + Routing Harness

OpenSquilla combines step-level singleton routing and multi-model ensemble routing in a learnable agent harness to preserve or improve task quality while substantially reducing execution cost.

tokenrhythm technologies · aiXiv Technical Report ·

OpenSquilla technical report cover with token-efficient routing benchmark results 阅读全文
03 Agentic Routing: The Harness-Native Data Flywheel 2026
Agent Systems

Agentic Routing: The Harness-Native Data Flywheel

Agentic routing turns execution traces and task outcomes into a harness-native data flywheel for continuously improving model selection, orchestration, and efficiency.

Xinchen Liu, Hang Zhou, Yingjie Zong , et al. · arXiv Preprint ·

Agentic Routing single-model and multi-model router comparison 阅读全文
04 From Question Answering to Task Completion: A Survey on Agent System and Harness Design 2026
Agent Systems

From Question Answering to Task Completion: A Survey on Agent System and Harness Design

A systematic survey of agent systems and harness design, reframing progress from answering questions toward reliably completing long-horizon tasks.

Jianyuan Guo, Zhiwei Hao, Chengcheng Wang , et al. · arXiv Preprint ·

Survey taxonomy for agent systems and harness design 阅读全文
05 Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks 2026
Machine Learning

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

Claw-SWE-Bench evaluates OpenClaw-style agent harnesses on realistic coding tasks while measuring capability, reliability, and the cost of end-to-end execution.

Mengyu Zheng, Kai Han, Boxun Li , et al. · arXiv Preprint ·

Claw-SWE-Bench model performance and API cost comparison 阅读全文
06 Sibyl-AutoResearch: Autonomous Research Needs Self-Evolving Trial-and-Error Harnesses, Not Paper Generators 2026
Agent Systems

Sibyl-AutoResearch: Autonomous Research Needs Self-Evolving Trial-and-Error Harnesses, Not Paper Generators

Sibyl-AutoResearch turns trial signals, failures, and evidence into auditable updates to later plans, validation, claims, scheduling, memory, and harness behavior.

Chengcheng Wang, Qinhua Xie, Wei He , et al. · arXiv Preprint ·

Sibyl-AutoResearch scientific trial-and-error harness and trace-path framework 阅读全文