VibeEdit.

Image Editing with Canvas Instructions

Jinjing Zhao1,*Fangyun Wei2,*Yitong Wang3Xiuyu Wu4Yunuo Chen5Yang Yue6Sirui Zhang7Wenbo Wang1Hongyang Zhang8Dong Chen2Yan Lu2Chang Xu1,†

Affiliations & contributions

1 University of Sydney   2 Microsoft Research   3 Fudan University   4 Nankai University
5 Shanghai Jiao Tong University   6 Tsinghua University   7 University of Science and Technology of China   8 University of Waterloo

* Equal contribution   † Corresponding author

Canvas instruction
Three chairs, with the middle chair circled and a note requesting velvet blue
Edited image
The middle chair becomes velvet blue while the surrounding room is preserved

Task

Canvas instructions specify what to change and where to edit directly on the image.

Spatial marks identify the target; short notes describe the change. Together, they form a canvas instruction. VibeEdit follows it to add, remove, replace, modify or move objects, without a separate text prompt.

Read the full abstract

In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source–target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.

01

Interface

Circle, scribble, write or drag directly on the image to specify the edit.

Circle

Select an object or an empty region.

Scribble

Mark an object to remove.

Write

Describe the new content or attribute.

Drag

Connect the source and destination.

Change the middle chair to blue

Canvas instruction
Source image with three white chairs change colorto velvet blue
VibeEditReady
Edited image
The middle chair is now velvet blue; the surrounding room is preserved
Waiting for the edit

Recorded result · looped preview

02

Method

VibeEdit separates source and canvas inputs, then learns through supervised training and reinforcement learning.

Canvas-annotated image requesting replacement of a chair with a red suitcaseAnnotated image A = I ⊕ C
Qwen2.5-VL FROZEN
semantic tokens
Clean source image of a hiker standing beside a folding chairClean source I
VAE encoder FROZEN
source tokens
Isolated red circle and replacement note on neutral grayCanvas layer C̄
VAE encoder FROZEN
canvas tokens
SEMANTICS + SOURCE + CANVAS + NOISY TARGET

Qwen-Image-Edit DiT

Trainable LoRA · rank 128

Separate source and canvas conditioning

Qwen2.5-VL reads the annotated image; the VAE encodes the source and canvas separately. Their tokens condition the DiT. Only rank-128 LoRA adapters are trained.

View model pipeline FIGURE 3Open full-resolution PDF
03

Data

We construct source–target image pairs offline and render canvas instructions during training.

1,554,062supervised training pairs
3,520RL conditions · 704 per task
  1. 1

    Select and segment

    Florence-2 proposes objects. GPT-5.6-sol selects an editable target and describes the change. SAM 3 produces the object mask.

  2. 2

    Generate edit pairs

    Generate edited targets and store each pair with masks and structured edit descriptions.

  3. 3

    Filter and annotate

    Filter low-quality pairs with GPT-5.6-sol. Render varied marks and notes from the stored records during training.

Supervised training corpus
Edit typePair generationStored pairs
Addition / removalObjectClear removes an object; reverse the pair for addition.435,496
Attribute modificationQwen-Image-Edit changes a selected attribute.458,154
ReplacementFLUX.1 Fill replaces an object inside an expanded mask.393,440
MovementQwen-Image-Edit moves the object; SAM 3 finds its destination mask.266,972

Counts exclude online reversals and annotation variations. RL conditions are selected for high reward variation across rollouts.

View data pipeline FIGURE 2Open full-resolution PDF
04

Results

Evaluate edit success on the intended target and preservation of the surrounding image.

VibeEdit benchmark

419 human-curated cases, with similar-looking objects that make target selection challenging.

96 Add80 Remove83 Modify78 Replace82 Move

Each case specifies an edit region and intended change. Text baselines receive an external prompt (21.3 words on average); VibeEdit receives the source and canvas.

Metrics

VLM rubric score (%)
A fixed GPT-5.6-sol judge checks edit success, outside preservation and local edit quality with binary questions. Per-case pass rates are averaged within each task, then equally across the five tasks.
Outside-region PSNR (dB)
Measures source–output similarity outside a padded edit region. Higher values indicate better preservation.
Higher is better for both metrics.

Full comparison

VLM rubric score (%) / outside-region PSNR (dB). Higher is better for both metrics.
MethodAddRemoveModifyReplaceMoveOverall

Visual-only baselines receive the annotated image. Text-instructed baselines receive the source and an external instruction. VibeEdit receives the source and canvas layers without a separate text prompt.

Paper

VibeEdit: Image Editing with Canvas Instructions

BibTeX

@misc{zhao2026vibeeditimageeditingcanvas,
  title={VibeEdit: Image Editing with Canvas Instructions},
  author={Jinjing Zhao and Fangyun Wei and Yitong Wang and Xiuyu Wu
          and Yunuo Chen and Yang Yue and Sirui Zhang and Wenbo Wang
          and Hongyang Zhang and Dong Chen and Yan Lu and Chang Xu},
  year={2026},
  eprint={2610.12229},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.12229},
}