General VLM benchmark data pipeline
Build visual benchmarks that really test models.
A reusable pipeline for building controlled visual edits, repairing local artifacts, authoring real-image probes, and evaluating whether VLMs can answer correctly when we test visual evidence.
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
University of Chicago1 · Toyota Technological Institute at Chicago2 · Stony Brook University3
Abstract
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specifications, generated or edited images, and question–answer pairs. Automated filtering removes candidates solved by a Filtering VLM, while human review verifies candidate validity and supports annotation correction and localized image repair. We instantiate SABRE-Prior to test whether VLMs follow visual evidence instead of relying on world priors—learned expectations about familiar objects and scenes. Its 600 images and 1,000 questions span Context (unexpected entities in familiar scenes), Texture (counterfactual materials), Attribute (non-canonical component counts), and Language Elicitation (answers suggested by language but unsupported by the image). Across six VLMs, macro-average accuracy ranges from 17.8% to 31.3% (22.6% mean). A real-image Attribute control is comparably difficult for the Filtering VLM. SABRE-Counting and SABRE-Spatial pilots show that the workflow supports other stress-test settings. These results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.
SABRE in Action
From automated benchmark construction to human verification and localized image repair.
Workflow Overview
From a Test Primer to structured candidates, automated filtering, and human-verified VLM stress tests.
Annotation & Verification
Review paired images, mark visual regions, correct question–answer metadata, and save verified samples.
Entity Removal
Remove a localized entity while preserving the surrounding scene, structure, and visual consistency.
Try entity removalEntity Swap
Replace a selected entity while retaining its location, scale, and relationship to the surrounding context.
Try entity swapPerformance under stress
Accuracy across four stress-test categories, with macro averages and 95% confidence intervals.
| Rank | Model | Context | Texture | Attribute | Language | Macro avg. |
|---|---|---|---|---|---|---|
| 01 |
|
10 |
40 |
17 |
58 |
31.3 |
| 02 |
|
7 |
52 |
17 |
17 |
23.3 |
| 03 |
|
3 |
46 |
14 |
29 |
23.0 |
| 04 |
|
0 |
52 |
26 |
11 |
22.3 |
| 05 |
|
1 |
28 |
20 |
23 |
18.0 |
| 06 |
|
4 |
28 |
16 |
23 |
17.8 |
Higher is better. Bold colored values mark the best result in each category.
Try localized image repair.
Your image is never uploaded to SABRE. It stays in this browser and is sent directly to Gemini only when you run an edit.














