alphaXiv

History

Papers Benchmarks

CUHK IMIXR

391

05 Dec 2025

agents computer-science computer-vision-and-pattern-recognition

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

Beihang University

Tsinghua University Meituan CUHK MMLab CUHK IMIXR

The EditThinker framework enhances instruction-following in any image editor by introducing an iterative reasoning process. It leverages a Multimodal Large Language Model to critique, reflect, and refine editing instructions, leading to consistent performance gains across diverse benchmarks and excelling in complex reasoning tasks.

04 Dec 2025

chain-of-thought computer-science artificial-intelligence

DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

South China University of Technology

Sun Yat-Sen University

The Chinese University of Hong Kong CUHK MMLab CUHK IMIXR

DraCo introduces an interleaved visual and textual reasoning paradigm for unified Multimodal Large Language Models (MLLMs), generating a low-resolution visual draft for self-correction. This approach significantly enhances text-to-image generation quality, achieving an 8% improvement on GenEval and a 0.91 point increase on ImagineBench for challenging and rare concept prompts.

05 Dec 2025

agents computer-science computer-vision-and-pattern-recognition

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

Beihang University

Tsinghua University Meituan CUHK MMLab CUHK IMIXR

Instruction-based image editing has emerged as a prominent research area, which, benefiting from image generation foundation models, have achieved high aesthetic quality, making instruction-following capability the primary challenge. Existing approaches improve instruction adherence via supervised or reinforcement learning, yet single-turn success rates remain limited due to inherent stochasticity and a lack of deliberation. In this work, we propose a deliberative editing framework to 'think' while they edit, which simulates the human cognitive loop by iteratively executing a Think-while-Edit cycle: Critiquing results and Refining instructions , followed by Repeating the generation until satisfactory. Specifically, we train a single MLLM, EditThinker, to act as the reasoning engine of this framework, which jointly produce the critique score, reasoning process, and refined instructions. We employ reinforcement learning to align the EditThinker's thinking with its editing, thereby generating more targeted instruction improvements. Extensive experiments on four benchmarks demonstrate that our approach significantly improves the instruction-following capability of any image editing model by a large margin. We will release our data construction framework, datasets, and models to benefit the community.

There are no more papers matching your filters at the moment.

Personalize Your Feed

Install Browser Extension

We're hiring

alphaXiv

Explore

State of the Art

Sign In

Labs

Feedback

Dark mode

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

Personalize Your Feed