Engineering

How We Built a More Controllable Virtual Try-On Model

How We Built a More Controllable Virtual Try-On Model

Engineering cover artwork

Category

Engineering

Published

Most AI image models rely on text prompts to control what they generate. For fashion teams, that creates a practical problem. A prompt can describe a garment, but it cannot reliably specify the exact shape, position, or styling the final image should follow.

Consider a simple request: roll up the left sleeve, tuck in the front of a shirt, or add a jacket without changing the clothing underneath. These instructions are easy to understand but difficult to translate into precise visual results through language alone. Teams often end up repeating the same process: adjust the prompt, generate another image, inspect the result, and try again.

CtrlVTON was developed to give virtual try-on models a more direct form of control. Instead of relying entirely on text, it allows users to define the intended garment placement through a visual mask.


Why text alone is not enough

Natural language works well for describing broad creative direction. It becomes less reliable when the task requires precise decisions about where a garment begins, how it overlaps with another item, or which parts of an outfit should remain unchanged.

A stylist may want a shirt to stop at a particular point on the waist. A merchandising team may need a jacket layered over an existing top without changing the rest of the outfit. A designer may want to adjust the sleeve length while preserving the person, pose, and background.

These are spatial instructions. They depend on shape and placement, not simply on the written description of a garment.


How CtrlVTON uses spatial control

CtrlVTON adds a spatial control layer to a virtual try-on model. Users draw the intended outline of a garment directly on the image, and the model uses that mask to guide how the clothing should appear on the body.

This approach supports three editing operations.

  • Full swap: Replace an existing garment with a different item.

  • Add: Introduce another garment, such as a jacket over a shirt, while preserving the existing outfit.

  • Partial swap: Replace one selected item while maintaining the surrounding garments.

Each operation is identified by a dedicated text token, while the mask provides the spatial information needed to guide the edit. The combination allows users to specify both what should change and where the change should happen.


Editing the garment instead of rewriting the prompt

Spatial control also changes how teams refine an image.

After generating an initial result, users can adjust the mask to change the garment’s placement or silhouette. Moving the hem can suggest a different tuck. Expanding the outline can create a looser fit. Shortening the mask can alter the sleeve length.

Because the edit is tied to a defined region, the process can preserve elements outside that area, including the person, pose, background, and other garments.

For production teams, this creates a clearer relationship between the intended modification and the resulting image. Instead of searching for the right wording, they can directly revise the part of the garment they want to change.


Virtual try-on garment and person reference outputs


Building the training data with VIP-SAM

Training a model to follow garment-specific spatial instructions requires accurate segmentation data.

The challenge becomes more difficult when a person is wearing several similar items. Standard segmentation methods may identify a general clothing category but struggle to determine which specific garment matches a separate reference image.

To address this, we developed VIP-SAM, short for Visual-Instance-Prompt Segmentation.

Given a reference image of a garment, VIP-SAM identifies the corresponding item in a photograph of a person wearing it. This makes it possible to isolate a particular piece even when multiple garments appear in the same outfit.

VIP-SAM achieved strong results on our fashion benchmark and standard segmentation benchmarks, supporting the training process required for more precise virtual try-on control.


How CtrlVTON performed against existing models

We evaluated CtrlVTON against proprietary image editing systems including Nano Banana Pro, GPT-Image 1.5, Seedream 4.5, and FLUX.2 [pro].

Each system received the same inputs: an image of a person, a reference garment, and a target outline defining the intended garment placement.

The comparison focused on two questions: whether the generated result preserved the reference garment and whether it followed the requested spatial layout.

CtrlVTON showed stronger alignment with the target layout while maintaining comparable garment fidelity. In evaluations against open weight virtual try-on models, it also performed strongly on preserving the appearance of the reference garment.

These results suggest that general image quality alone does not capture the capabilities required for controllable fashion editing. A convincing result also needs to follow the intended shape and placement.


A benchmark for controllable virtual try-on

We are also releasing VITON-HD-edit, a public benchmark designed to evaluate editing-based virtual try-on, mask-controllable try-on, and instance-level segmentation.

The benchmark gives researchers a shared way to assess whether a model can preserve garment identity, follow a specified layout, and support targeted edits within an existing outfit.


Why controllability matters

Virtual try-on becomes more useful for fashion and merchandising when teams can specify how a garment should appear rather than repeatedly generating variations until one looks right.

CtrlVTON moves that process toward direct visual instruction. The garment reference defines what should be shown, the editing operation defines what should change, and the mask defines where the change should occur.

The same principle may also apply to other visual production tasks that require precise spatial control, including product placement, character design, and industrial visualization.

All human model and clothing images are either free-license images from the internet, drawn from public datasets including VITON-HD, DressCode, DressCode-MR, Garments2Look, and OmniTry Bench, or generated. All images and brands remain the property of their respective owners.

Model in a red draped dress with a high slit, leaning against a dark wood door frame
Model in a red draped dress with a high slit, leaning against a dark wood door frame

Build the intelligence layer behind your brand.

Start with creative production. Build the product and brand intelligence that powers how your brand operates how the world discovers what you sell.