Welcome to my academic homepage. I am Yunlong Lin. I work on long-horizon agents that operate, learn, and improve in open-ended environments, with a focus on agentic reinforcement learning, harness design, and recursive self-improvement (RSI).

Research Interests

My research asks how agents can sustain open-ended work, learn from execution, and improve the systems around them:

(i) Long-horizon agent systems & harnesses: persistent state, memory, planning, tool use, and reliable execution across evolving workflows in JarvisHub, Claw-Eval-Live, and JarvisX-Cowork.

(ii) Agentic RL & environment feedback: trajectory rollouts, rubrics, verifiers, evaluator design, and harness generalization in Seed2.0, JarvisIR, JarvisArt, and Gen-Searcher.

(iii) Self-improving agents / RSI: Auto R&D and AI-for-AI through the joint evolution of skills, data, evaluators, and harnesses in the Seed2.1 Auto R&D Agent ("Seed-for-Seed"), JarvisEvo, and GenEvolve.

I welcome discussions and collaborations on long-horizon agents, agentic RL, harnesses, and recursive self-improvement.

Contact
  • WeChat: lyl20136148
  • Email: linyl@stu.xmu.edu.cn
News
  • JarvisHub, our open harness for long-horizon agents with persistent state and traceable feedback, is now live: Project, GitHub, Hugging Face, 机器之心, and 新智元.
  • Claw-Eval-Live, our live benchmark for agents in evolving real-world workflows, has been released.
  • JarvisEvo, our self-evolving agent with coupled actor-evaluator optimization, has been accepted by CVPR 2026.
  • JarvisArt, our long-horizon tool-using agent, has been accepted by NeurIPS 2025; the code and benchmark are open-sourced.
Experience
  • Jan 2026 – Present · Research Intern, ByteDance Seed Long-horizon agent improvement through harness generalization, rubric/evaluator design, and Auto R&D ("Seed-for-Seed").
  • Jun 2025 – Dec 2025 · Qingyun Intern, Tencent Hunyuan Self-evolving agents through long-horizon interaction, reward feedback, and editor-evaluator optimization.
Selected Research
ByteDance Seed
ByteDance Seed: Agent Improvement & Auto R&D
Seed2.1: Auto R&D Agent for AI-for-AI

ContributionWorked on the Auto R&D Agent and the "Seed-for-Seed" prototype, where rubrics and evaluators turn R&D outcomes into feedback for iterative agent self-improvement.

Seed2.0: Agent Improvement through Harness Generalization

ContributionWorked on harness generalization and evaluator design, enabling agents to learn robust long-horizon behavior across changing interfaces and execution environments.

Open-Source Project
JarvisHub canvas-native agent harness
JarvisHub: An Open Harness for Canvas-Native Creative Agents

JarvisHub is an open harness for long-horizon agents. It provides persistent project state, structured actions, shared memory, and traceable environment feedback across open-ended workflows.

CVPR 2026 (Tencent HY)
JarvisEvo editor-evaluator loop
JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization

JarvisEvo studies recursive self-improvement through a coupled actor-evaluator loop. Long-horizon trajectories, reward feedback, and reflection allow the agent and its evaluator to improve together.

Preprint 2026
Claw-Eval-Live workflow overview
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

Claw-Eval-Live is a continuously refreshed benchmark for long-horizon agents in evolving real-world workflows. Verifiable execution traces and live tasks expose whether agents can generalize beyond static harnesses.

Open-Source Project
JarvisX-Cowork demo
JarvisX-Cowork: A Personal AI Creative Assistant for End-to-End Creative Workflows

JarvisX-Cowork is an end-to-end harness for open-ended creative work. Persistent planning, shared memory, and structured tool interfaces help the agent carry state coherently from intent to final deliverables.

Preprint 2026
Gen-Searcher agentic search overview
Gen-Searcher: Reinforcing Agentic Search for Image Generation

Gen-Searcher applies agentic RL to multi-step search and generation workflows. Search trajectories and environment feedback train the agent to gather evidence, revise decisions, and improve final outcomes.

NeurIPS 2025
JarvisArt agent workflow
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

JarvisArt formulates expert-tool orchestration as a long-horizon agent problem. Interpretable subgoals, execution feedback, and iterative control support reliable multi-step workflows inside a professional editing environment.

CVPR 2025
JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration

JarvisIR post-trains a tool-using agent with mixed-rank reward feedback. The restoration stack serves as an execution environment in which the agent learns to plan, select, and coordinate experts for downstream objectives.