VisualErase

A visual perspective on concept erasure

VisualErase.

Dual-Branch Visual Trajectory Redirection for Robust Concept Erasure in Text-to-Image Diffusion Models

Give erasure a visual destination. Redirect both text-conditioned and text-free denoising toward concept-removed images, while retaining general generation quality.

Qianlong Xiang1,2,3 Miao Zhang1 Kun Wang4 Yupeng Hu5 Junhui Hou2 Liqiang Nie1

1 Harbin Institute of Technology (Shenzhen) 2 City University of Hong Kong 3 Shenzhen Loop Area Institute 4 National University of Singapore 5 Shandong University

0%

Style · maximum ASR

8%

Celebrity · maximum ASR

0.1%

Nudity · maximum ASR

7

Attacks · including TINA+

Maximum attack success rate across the seven evaluated attacks for each task. Lower is better. Results on Stable Diffusion v1.4.

AbstractRead the full abstract

Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts while preserving general generation capability by redirecting target-related text-to-image mappings. However, recent studies show that erased models may still retain visual generative trajectories of target concepts, leaving them vulnerable to adversarial recovery attacks and revealing a fundamental gap between redirecting text-to-image mappings and truly removing visual knowledge. To bridge this gap, we propose VisualErase, a new paradigm that redirects concept-bearing visual generative trajectories toward explicitly defined concept-removed outcomes. To enable this redirection, we use structure-preserving image editing to construct content-aligned, concept-removed counterparts for source images, providing explicit visual endpoints that retain non-target content. We then derive a denoising target from each source-to-counterpart pair and use a dual-branch redirection loss to align both text-conditioned and unconditional predictions with this target, since conditional supervision alone does not explicitly constrain generation without textual guidance. To mitigate the adverse effects of concept erasure on non-target generation, we jointly optimize the redirection loss with a counterpart retention loss that matches denoising predictions from the frozen pretrained model. Across style, celebrity, and nudity erasure, VisualErase limits the maximum attack success rate over seven attacks to 0%, 8%, and 0.1%, respectively, while retaining general generation quality. These results highlight the importance of visual trajectory redirection for robust concept erasure beyond text-to-image mappings alone.

01

A visual perspective on concept erasure

The gap

A silenced prompt can leave a visual pathway intact.

Text-centric erasure redirects target-related prompts. Yet text-free inversion with TINA+ can still recover target concepts, exposing residual visual generation pathways.

Our approach

Define what the model should generate instead.

A content-aligned counterpart defines the desired visual endpoint. VisualErase supervises both denoising branches toward this endpoint without adversarial prompt search or attack-generated supervision.

Comparison of text-to-image redirection and visual trajectory redirection
Text-to-image redirection changes a textual association. Visual trajectory redirection specifies concept-removed endpoints for source visual states.

02

Dual-Branch Visual Trajectory Redirection

VisualErase paired data construction, target derivation and dual-branch redirection
Paired image construction provides an explicit endpoint. The same derived target supervises source-prompt and empty-prompt predictions.
01

Construct counterparts

Remove style, replace identity, or add clothing while preserving non-target content, structure, and composition.

02

Redirect both branches

Denoise the same source state toward the counterpart under both the original prompt and the empty condition.

03

Retain generation utility

Reuse counterparts as retention examples and match the frozen pretrained denoiser on their noised states.

εredir = (ztsrc − αt zref) / σt

ℒ = 0.5 ℒcond + 0.5 ℒuncond + ℒretain

Style and celebrity update attention projections for 1,000 steps. Nudity uses the same objective with fixed top-10% saliency-gated Non-Cross weights for 5,000 steps.

03

Robust erasure across three tasks

Across style, celebrity, and nudity erasure, VisualErase suppresses standard and adversarial concept recovery while preserving general generation quality.

Van Gogh style erasure
MethodACC ↓PEZ ↓MMA ↓RAB ↓P4D ↓UDA ↓CCE ↓TINA+ ↓FID ↓CLIP ↑
SD1.476.040.040.090.094.096.060.076.014.026.6
STEREO0.00.00.00.00.00.04.046.015.826.0
VisualErase0.00.00.00.00.00.00.00.014.726.5

ACC / ASR (%): lower is better. COCO-30K generation quality: FID ↓, CLIP ×100 ↑.

04

Qualitative comparisons

Style, celebrity and nudity comparison of SD1.4, ESD, STEREO and VisualErase under standard generation and TINA+
Same evaluation cases across SD1.4, ESD, STEREO, and VisualErase. Black masks conceal explicit baseline content for presentation only. Blur in VisualErase outputs is generated by the model, not added in post-processing. Click a figure to inspect the full-resolution image.

Conclusion

Robust erasure through visual trajectory redirection.

VisualErase redirects conditional and unconditional generation toward concept-removed counterparts. Across style, celebrity, and nudity erasure, it achieves zero or near-zero target recovery under seven attacks while maintaining strong generation quality. These results highlight the importance of visual trajectory redirection for robust concept erasure beyond text-to-image mappings alone.