Text → image · SD v1.4
38.1% → 3.5% Average harmful generations on I2P, compared with the unaligned backbone.Multimodal safety · Vision & language
Align what is unsafe.
Preserve what is not.
ShieldCLIP is a selective safety-alignment framework for harmful content mitigation in multimodal foundation models.
* Equal contribution
Pair-level labels hide where the harm actually is.
Existing alignment pipelines often treat every generated image–text pair as jointly unsafe. In practice, one generated modality can be benign while the other is harmful.
ShieldCLIP uses independent safety labels for text and images. Safe content stays anchored to the original CLIP space; only unsafe representations are redirected. Mixed pairs update the unsafe branch alone, while jointly unsafe pairs retain cross-modal coherence.
Real and generated samples are processed by trainable encoders and compared with frozen CLIP anchors. The objective changes with the observed text and image safety state.
Safety labels at the modality level.
ViSUv2 pairs real safe image–text samples with generated counterparts, then labels generated text and images independently. This exposes the mixed cases that pair-level supervision misses.
Representative ViSUv2 quadruplets. Green denotes safe content; red denotes unsafe content. Sensitive visual content is blurred in the source figure.
Safer generations, retrievals, and captions.
ShieldCLIP reduces harmful outputs across text-to-image generation, bidirectional retrieval, and image-to-text generation—without abandoning the original CLIP geometry.
Text → image · SDXL
33.6% → 4.6% Average harmful generations on I2P, with the same encoder-level intervention.Text → image retrieval · ViSUv2
95.2% → 3.5% Harmful top-1 retrievals compared with the original CLIP model.Harmful image generation rate · lower is better
Across two diffusion backbones
| Model | SD v1.4 | SDXL | ||
|---|---|---|---|---|
| I2P | ViSUv2 | I2P | ViSUv2 | |
| Original backbone | 38.1 | 25.4 | 33.6 | 33.4 |
| Safe-CLIP | 21.8 | 6.1 | 8.2 | 2.0 |
| SafetyDPO | 11.4 | 2.8 | 8.7 | 4.4 |
| DES | 3.9 | 1.1 | 16.3 | 5.3 |
| ShieldCLIP | 3.5 | 1.0 | 4.6 | 1.4 |
Balanced retrieval
Harmful retrievals on NudeNet, NSFW URLs, and SMID stay at or below 0.4% in both text-to-image and image-to-text directions.
Safer image captioning
Replacing LLaVA’s visual encoder with ShieldCLIP lowers harmful captions on NudeNet images, with similar gains on NSFW URLs.
Utility preservation
ShieldCLIP outperforms SafeR-CLIP on five of six zero-shot classification datasets and remains close to the original CLIP encoder.
Human comparisons
Across a blinded study with 20 annotators, ShieldCLIP was preferred for both safety and semantic preservation over three strong alternatives.
Redirecting unsafe cues while retaining the scene.
Comparisons across SD v1.4, SDXL, ShieldCLIP, and competing mitigation methods. Sensitive outputs are blurred in the paper figure.
View cross-modal retrieval examples
Modality-aware supervision
Independent labels for generated text and images distinguish safe–safe, unsafe–unsafe, and both mixed safety states.
Selective objective
Safe content is preserved, unsafe content is redirected, mixed pairs update only the unsafe branch, and jointly unsafe pairs retain coherence.
Cross-task validation
Evaluation spans retrieval, Stable Diffusion v1.4 and SDXL image generation, and LLaVA image-to-text generation.
Release hub
Everything will live here.
Code, trained models, reproduction instructions, and the controlled-access protocol for ViSUv2 will be released in this repository.
BibTeX will be added upon release.