ImproveAnyTask: An Autonomous Post-Training Harness for Iterative Model Self-Improvement

Xingbo Yao1,3,* Xiaoman Wang4,* Zhengwu Lei3,5 Tinghui Luo3,5 Yilin Zhang3 Yuefeng Wu3,6 Yijie Xu1 Tianfu Wang1 Qingyuan Zhan3 Ye Guo3 Daoxin Zhang3 Zhe Xu3 Jian Liu3,† Hui Xiong1,2,†

1HKUST(GZ)   2HKUST   3Xiaohongshu Inc.   4ECNU   5ZJU   6NTU
*Equal contribution.   †Corresponding authors.

ImproveAnyTask diagnoses model errors, selects research-backed update directions, and executes task-specific post-training within a fixed compute budget.

Recursive self-improvement (RSI) Autonomous post-training LLM adaptation Agent harness Model self-improvement
11benchmarks
+18.29Base mean gain
+11.97Instruct mean gain
24 hbudget per optimization
Performance comparison across eleven benchmarks in five capability domains.
Qwen3.5-4B-Base results across five domains. ImproveAnyTask improves all 11 tasks, with a maximum gain of 41.96 percentage points on LiveCodeBench v6.

Method

From observed errors to executable updates

The harness organizes task adaptation as a repeated evaluation, research, update, and re-evaluation loop.

01

Error Attribution

Combines aggregate metrics with individual failures to identify the highest-impact problem.

02

Update Direction

Compares research-backed strategies by expected gain, applicability, and reproduction difficulty.

03

Executable Model Updates

Translates the selected strategy into data and training, with execution checks before full post-training.

ImproveAnyTask framework showing error attribution, update direction, model updates, and re-evaluation.
ImproveAnyTask treats task adaptation as iterative model optimization.

Main results

Broad gains under the same budget

We evaluate Base and Instruct models on knowledge, reasoning, instruction following, coding, tool use, and general-agent tasks.

+41.96

Largest improvement

LiveCodeBench v6 on Qwen3.5-4B-Base.

5

Capability domains

Knowledge, reasoning, instruction following, coding, and agents.

2

Model initializations

Qwen3.5-4B-Base and Qwen3.5-4B-Instruct.

Optimization analysis

Iterative gains, recovery, and module effects

Extended optimization records both successful updates and regressions, while ablations replace each major module with a basic alternative.

IFEval optimization trajectory over a 100-hour equivalent compute budget.
IFEval over a 100-hour equivalent budget. The trajectory reaches 85.21% before rejecting a weaker final candidate.
Ablation trajectories for MMLU-Pro and MATH-500.
MMLU-Pro and MATH-500 trajectories. Replacing any major module reduces the final improvement.

Task adaptation without manually redesigning every iteration

ImproveAnyTask connects diagnosis, research, and post-training in one continuous loop, reducing the manual effort required to improve a model for a specific task.