Self-evolving vision-language models

Program-Verified Self-Evolution for Vision-Language Models

Ahmed Heakl1,2Sungik Choi1Moontae Lee1,4Salman Khan2,3
1LG AI Research2MBZUAI3Australian National University4University of Illinois at Chicago
Radar chart: VQS improves every benchmark on Qwen3-VL-2B, VISE loses on LogicVista and MMStar
VQS improves every benchmark. Qwen3-VL-2B; VISE, the strongest baseline, loses on LogicVista and MMStar.
94.4%
of VQS answers are correct
Human evaluation; 76.4% for majority vote, 82.2% for a model judge.
+3.18
average over 10 benchmarks
Qwen3-VL-2B after one cycle, against +1.01 for the strongest baseline.
3 / 3
scales won
Largest gain at 2B, 4B and 8B: +3.18, +2.32 and +2.43.
+3.84
after three cycles
Gains keep growing on the same unlabeled images.

Abstract

Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24% of majority-vote labels and 18% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94% of VQS answers correct, against 76% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B.

Majority votes reward wrong answers

Without gold answers, prior self-evolving methods take their label from the solver's own outputs. When most samples are wrong, the wrong answer becomes the label and is reinforced. VQS computes the answer with a program over a checked parse instead.

Majority-vote labels vs. computed answers
Majority-vote labels vs. computed answers. (a) In prior self-play, three of five answers say 2 mugs, so the wrong answer is rewarded and the correct "1" is discarded. (b) VQS: a schema-constrained parser turns the image into a structured parse, a fixed template writes the question and computes its answer, a visual checker verifies each claim behind it, a blind gate drops questions answerable without the image, a difficulty band keeps questions the solver gets right in some but not all of 8 rollouts, and GRPO rewards exact match with the computed answer.

Method

One set of weights plays every role: parser, paraphraser, checker, blind gate and solver. No human question, answer or label is used at any stage.

1

Parse

The model turns each image into a record under a per-domain JSON schema that the decoder enforces.

2

Program

70 template families read the record and return a question, its computed answer and its hop count.

3

Check

Every claim behind the answer is confirmed against the image; a blind gate and a difficulty band filter the rest.

4

Train

GRPO with reward 0.9 × exact match + 0.1 × format, in a curriculum ordered by hop count.

Examples of structured scene parsing
Structured parsing. Photos become object–relation graphs, charts become series of printed and numeric values, diagrams become directed graphs, and infographics become entry records.

Label-free parser training

The parser samples K = 4 parses per image. Each is scored by the fraction of its claims the checker confirms, and the best becomes an SFT target only if it beats the mean of the four by δ = 0.10 and asserts at least m = 4 claims.

Label-free parser training
Label-free parser training. z₂ calls the mug blue and puts it right of the book, so it passes 6 of 8 claims. z₁ passes 8/8 and beats the mean (0.75) by the margin, so it becomes the target.

Explore the pipeline

Ninety-three generated questions, each traced from the image to the final keep-or-drop decision: the parse, the program that computed the answer, the question rewrite, every fact the model re-read, and its four guesses without the image. Everything shown was produced by Qwen3-VL-2B on its own parses; nothing is written by hand. Open any card, or start with one of these:

Keep and drop decisions apply the paper's gates to the recorded outputs: a question is dropped if a read-back contradicts the parse, or if more than half of the four blind answers match. "Human-graded" examples carry a rater's verdict from the paper's human evaluation.

Results

Ten benchmarks, three Qwen3-VL scales, five self-evolving baselines. VQS gives the largest average gain at every scale and improves all ten benchmarks at each.

InfoVQA = InfographicsVQA, SQA = ScienceQA, MMB = MMBench, ESB = EmbSpatial, LogicV = LogicVista. Small numbers are changes against the backbone's own base; underlined values are the best trained model per column.

Average gain over base
Are the generated answers correct?

Human raters checked 500 shared questions, 125 per visual domain. Accuracy among retained questions; VQS keeps 76.0% of questions, majority vote 92.0%, the model judge 90.8%.

Analysis

Choosing the parser's targets

Parser targetParse prec.Answer acc.Solver acc.
Untrained80.886.858.0
Random parse82.191.061.5
Whole-parse check85.094.062.0
Per-claim check (VQS)87.696.162.4

Precision and answer accuracy rated by hand on 150 samples; solver accuracy is the 10-benchmark average of a 2B solver trained on each parser's questions.

Parser structured F1 before and after training
The parser improves without labels. F1 against human annotations rises by 5.0 on GQA scene graphs, 6.6 on ChartQA tables and 7.4 on AI2D-RST diagrams.
Parser target selection: quality and coverage
Parser target selection. Each panel varies one setting around (K, δ, m) = (4, 0.10, 4). Precision is the share of correct visual claims after training; retained is the share of images yielding a target.

Gains keep growing over cycles

Gains across training cycles
Average gain over base per benchmark category for Qwen3-VL-2B. Each cycle merges the previous solver, writes new questions on the same images and resamples toward a 50% solve rate.
  • +3.84after three cycles, up from +3.18 after one and +3.41 after two. The third cycle adds more than the second, so the curve has not flattened.
  • +5.94on knowledge and reasoning benchmarks, led by ScienceQA (+8.53) and MMMU (+8.44).
  • +2.62on perception and +2.24 on general benchmarks.
  • 9 / 10benchmarks improve in every cycle. MMStar peaks after cycle 1 and still ends 0.93 above base.

What matters

Pipeline (2B)Solver acc.Parser prec.
Untrained parser58.080.8
Full VQS62.487.6
without constrained decoding61.985.3
without fact-checker60.987.6
without blind gate61.887.6
without difficulty band62.187.6
Question design (2B)SettingSolver acc.
Family difficultyEasy / Middle / Hard60.8 / 62.4 / 61.1
GeneratorFixed templates / written programs62.4 / 61.4
CurriculumShuffled / global / within-family61.3 / 61.7 / 62.4
Difficulty bandInitial / refreshed62.4 / 62.3

Scale and coverage

What scale and coverage change
What scale and coverage change. (a) RL changes fewer benchmark answers at larger scale: 7.82% at 2B, 5.58% at 4B, 4.93% at 8B. (b) All three scales reach a training reward of about 0.80. (c) More template families keep raising the 2B gain, from 1.9 with 10 families to 3.18 with 70.

Other model families

VQS improves all ten benchmarks for every family: +1.77 on Gemma3-12B-It, +1.51 on InternVL3-8B-Instruct and +1.84 on Llama-3.2-11B-Vision-Instruct.

Use the model

VQS-2B is Qwen3-VL-2B-Instruct trained with VQS on questions it wrote for itself from unlabeled images. It loads like any Qwen3-VL checkpoint. Answers are short by design, so end each question with the answer-format instruction used in training.

from transformers import AutoModelForImageTextToText, AutoProcessor
from PIL import Image

model = AutoModelForImageTextToText.from_pretrained("ahmedheakl/VQS", dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained("ahmedheakl/VQS")

messages = [{"role": "user", "content": [
    {"type": "image"},
    {"type": "text", "text": "In 2020, which is higher, Stayovers or Day trippers?\n"
                             "Answer the question using a single word or phrase."},
]}]
text = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = processor(text=[text], images=[Image.open("chart.png")], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])

The model card also covers vLLM and training details; the full pipeline, from unlabeled images to GRPO, is on GitHub.

Citation

@article{heakl2026vqs,
  title   = {Program-Verified Self-Evolution for Vision-Language Models},
  author  = {Heakl, Ahmed and Choi, Sungik and Lee, Moontae and Khan, Salman},
  journal = {arXiv preprint arXiv:2609.33855},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.33855}
}