NeurIPS 2026 Poster

Half-Truths Break Similarity-Based Retrieval

1 University of Tübingen · Tübingen AI Center
2 Konrad Zuse School of Excellence in Learning and Intelligent Systems (ELIZA)

Adding a wrong detail can make a caption score higher

CLIP retrieves images by scoring how well they match a text description. Adding something that is not in the image should lower that score. Yet a correct description with one false detail—a half-truth—can score higher. Compare CLIP with our method, CS-CLIP, on two real examples below.

A woman using computers at a table; no printer is visible.

woman

Image–text similarity · dashed lines show the score before adding a detail

CLIP0.214
CS-CLIP0.275

Start with a correct description. The dashed lines mark its score before adding a detail.

Measured scores from the verified COCO evaluation set. A correct addition should raise the score; a false addition should lower it.

Abstract

When a text description is extended with an additional detail, image-text similarity should drop if that detail is wrong. We show that CLIP-style dual encoders often violate this intuition: appending a plausible but incorrect object or relation to an otherwise correct description can increase the similarity score. We call such cases half-truths. Across 1,437 verified comparisons from COCO and CC3M using two foil generators, we find that this failure persists across generators and image domains. On COCO with Qwen-generated foils, CLIP prefers the correct shorter description only 34.0% of the time, dropping to 27.8% for relation additions. We trace this vulnerability to weak supervision on caption parts: contrastive training aligns full sentences but does not explicitly enforce that individual entities and relations are grounded.

We propose CS-CLIP (Component-Supervised CLIP), which decomposes captions into entity and relation units, constructs a minimally edited foil for each unit, and fine-tunes the model to score the correct unit above its foil while preserving standard dual-encoder inference. CS-CLIP raises half-truth accuracy to 76.4% on this set and improves average performance on established compositional benchmarks by 5.8 points, suggesting that reducing half-truth errors aligns with broader gains in compositional understanding.

Teach the model to check each part of a caption

Standard CLIP training matches whole captions to images. Component-Supervised CLIP (CS-CLIP) also trains on the objects, attributes, and relationships inside them. Each correct part is paired with a minimally changed, incorrect alternative, which we call a foil.

  1. 01

    Split the caption

    “A woman riding a brown horse in a park.”

    brown horsewoman riding horsewoman in park

    A text model extracts objects and relationships.

  2. 02

    Change one detail

    brown horsewhite horse
    woman riding horsewoman riding giraffe

    Small edits create plausible alternatives.

  3. 03

    Train on the difference

    Image + correct part↑ similarity
    Image + incorrect part↓ similarity

    Learn these distinctions alongside whole-caption training.

Same retrieval at test time. Encode the image, encode the text, compare their similarity.

Training architecture

CS-CLIP training architecture: extract caption units, generate minimally edited foils, and train image and text encoders with unit-level supervision.

Better at rejecting false details—even on a different image source

Accuracy measures how often the correct description beats its half-truth. We test 1,437 human-verified comparisons: COCO with two different foil generators, and CC3M with different images. CS-CLIP leads on both COCO sets and on relationship additions across all three sets. Entity additions describe objects or attributes; relation additions describe relationships.

Half-Truth accuracy (%)

Main evaluation

COCO · Qwen

509 comparisons
CLIP34.0
FSC-CLIP61.3
CS-CLIP76.4
Different foil generator

COCO · Mistral

454 comparisons
CLIP28.6
FSC-CLIP54.0
CS-CLIP67.2
Different image source

CC3M · Qwen

474 comparisons
CLIP36.1
FSC-CLIP67.5
CS-CLIP66.2

Higher is better. FSC-CLIP is trained on CC3M and leads on that set overall; CS-CLIP leads on its relation examples.

Compare all 16 models

Download all results (CSV) ↓
Overall accuracy (%) · higher is better · bold marks the best score
Model COCO
Qwen
COCO
Mistral
CC3M
Qwen
CLIP 34.0 28.6 36.1
NegCLIP 53.8 45.2 55.3
DeGLA 46.2 40.1 52.3
ReadCLIP 58.0 50.9 60.1
FSC-CLIP (CC3M) 61.3 54.0 67.5
FSC-CLIP (COCO) 56.0 50.7 60.1
FSC-CLIP (LAION+COCO) 49.7 43.8 61.0
CoN-CLIP 56.4 50.4 48.1
CE-CLIP 43.8 39.2 55.5
LabCLIP 44.4 38.1 43.0
DAC (LLM) 39.7 36.3 42.0
DAC (SAM) 37.1 35.7 42.0
TSVLC 37.1 35.9 41.8
CLoVe 34.2 28.6 48.5
CLIC (LAION) 33.8 31.3 39.5
CS-CLIP (ours) 76.4 67.2 66.2

CS-CLIP has the highest overall accuracy on both COCO sets; FSC-CLIP (CC3M) has the highest on CC3M.

What comes from the examples, and what comes from the training?

Give the model the same incorrect alternatives, then change whether it learns from whole sentences or individual parts.

Accuracy (%) · COCO (Qwen) and 16 compositional benchmarks · higher is better
Training Entity Relation Compositional
NegCLIP 62.2 45.5 55.3
Same foils, whole sentences 66.5 82.7 56.8
Same foils, individual parts CS-CLIP 73.6 79.2 57.8

Sentence-level foils already improve relationships. Training on individual parts adds 7.1 points on entities and 1.0 on compositional accuracy; the sentence-level variant remains strongest on relations.

Fewer half-truth errors, stronger compositional understanding

Among the 16 evaluated models, CS-CLIP achieves the highest average Half-Truth accuracy across the three sets and the highest average image-to-text accuracy across 16 compositional benchmarks.

CS-CLIP compared with CLIP, NegCLIP, DeGLA, ReadCLIP, and FSC-CLIP trained on CC3M or COCO. CS-CLIP leads with 69.9% average Half-Truth accuracy and 57.8% compositional accuracy; the strongest baselines reach 60.9% and 57.4%, respectively.
Selected baselines, including the strongest comparator for each metric. Half-Truth averages the three evaluation sets; compositional accuracy averages 16 benchmarks. FSC-CLIP labels indicate its training data.

The gains extend beyond the Half-Truth test. CS-CLIP also improves matching on independent compositional benchmarks, while keeping standard dual-encoder retrieval. Zero-shot classification falls from 63.6% to 59.9%.

Cite this work

@inproceedings{kargi2026halftruths,
  title     = {Half-Truths Break Similarity-Based Retrieval},
  author    = {Kargi, Bora and Uselis, Arnas and Oh, Seong Joon},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2602.23906}
}