Multimodal safety · Vision & language

Align what is unsafe.
Preserve what is not.

ShieldCLIP is a selective safety-alignment framework for harmful content mitigation in multimodal foundation models.

Tobia Poppi*Silvia Cappelletti*Samuele Poppi Marcella CorniaLorenzo BaraldiDiego Garcia-OlanoRita Cucchiara

* Equal contribution

Paper · Coming soon Code · Coming soon ViSUv2 · Controlled access
195Kquadruplets
578fine-grained concepts
28safety categories
4modality safety states

Pair-level labels hide where the harm actually is.

Existing alignment pipelines often treat every generated image–text pair as jointly unsafe. In practice, one generated modality can be benign while the other is harmful.

ShieldCLIP uses independent safety labels for text and images. Safe content stays anchored to the original CLIP space; only unsafe representations are redirected. Mixed pairs update the unsafe branch alone, while jointly unsafe pairs retain cross-modal coherence.

ShieldCLIP architecture and conditional training objectives

Real and generated samples are processed by trainable encoders and compared with frozen CLIP anchors. The objective changes with the observed text and image safety state.

Safety labels at the modality level.

ViSUv2 pairs real safe image–text samples with generated counterparts, then labels generated text and images independently. This exposes the mixed cases that pair-level supervision misses.

ViSUv2 generation and independent text and image safety-labeling pipeline
179Ktraining
8Kvalidation
8Ktest
47.7% of generated pairs are not unsafe–unsafe: 38.9% are mixed and 8.8% are safe–safe.
Examples of real safe pairs and generated safe or unsafe counterparts in ViSUv2

Representative ViSUv2 quadruplets. Green denotes safe content; red denotes unsafe content. Sensitive visual content is blurred in the source figure.

Safer generations, retrievals, and captions.

ShieldCLIP reduces harmful outputs across text-to-image generation, bidirectional retrieval, and image-to-text generation—without abandoning the original CLIP geometry.

Text → image · SD v1.4

38.1% → 3.5% Average harmful generations on I2P, compared with the unaligned backbone.

Text → image · SDXL

33.6% → 4.6% Average harmful generations on I2P, with the same encoder-level intervention.

Text → image retrieval · ViSUv2

95.2% → 3.5% Harmful top-1 retrievals compared with the original CLIP model.

Harmful image generation rate · lower is better

Across two diffusion backbones

NudeNet + Q16 evaluation
ModelSD v1.4SDXL
I2PViSUv2I2PViSUv2
Original backbone38.125.433.633.4
Safe-CLIP21.86.18.22.0
SafetyDPO11.42.88.74.4
DES3.91.116.35.3
ShieldCLIP3.51.04.61.4
≤ 0.4%

Balanced retrieval

Harmful retrievals on NudeNet, NSFW URLs, and SMID stay at or below 0.4% in both text-to-image and image-to-text directions.

58.6 → 12.5%

Safer image captioning

Replacing LLaVA’s visual encoder with ShieldCLIP lowers harmful captions on NudeNet images, with similar gains on NSFW URLs.

5 / 6

Utility preservation

ShieldCLIP outperforms SafeR-CLIP on five of six zero-shot classification datasets and remains close to the original CLIP encoder.

1,200

Human comparisons

Across a blinded study with 20 annotators, ShieldCLIP was preferred for both safety and semantic preservation over three strong alternatives.

Redirecting unsafe cues while retaining the scene.

Comparisons across SD v1.4, SDXL, ShieldCLIP, and competing mitigation methods. Sensitive outputs are blurred in the paper figure.

Qualitative text-to-image comparison between base diffusion models, ShieldCLIP, and safety baselines
View cross-modal retrieval examples
Qualitative text-to-image and image-to-text retrieval comparisons
01

Modality-aware supervision

Independent labels for generated text and images distinguish safe–safe, unsafe–unsafe, and both mixed safety states.

02

Selective objective

Safe content is preserved, unsafe content is redirected, mixed pairs update only the unsafe branch, and jointly unsafe pairs retain coherence.

03

Cross-task validation

Evaluation spans retrieval, Stable Diffusion v1.4 and SDXL image generation, and LLaVA image-to-text generation.

Release hub

Everything will live here.

Code, trained models, reproduction instructions, and the controlled-access protocol for ViSUv2 will be released in this repository.

PaperComing soon
CodeWatch repository ↗
CheckpointsComing soon
ViSUv2Controlled access
Citation BibTeX will be added upon release.