DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
A framework for evidence-grounded document field extraction with compact vision-language models
Abstract
Many document fields must be derived from visual evidence through multiple reasoning steps. DocMIDE trains compact vision-language models to plan, retrieve evidence and derive an answer. Its GRPO training rewards output format, retrieved evidence, intermediate derivation steps and the final value, raising Qwen3.5-4B accuracy from 70.8% to 95.9% on a 4,151-pair implicit extraction benchmark.







