Evidence-Grounded and Uncertainty-Aware Vision-Language Models for Reliable Chest X-Ray Interpretation and Report Generation
Supervisory Team
Introduction
Vision-language models (VLMs) show strong potential for automated chest X-ray interpretation and radiology report generation. However, their reliability remains limited by hallucinated findings, incorrect localisation, poor confidence calibration, and reduced robustness across datasets and rare abnormalities. This PhD project will develop an evidence-grounded and uncertainty-aware VLM framework that links generated findings to relevant image regions and uses uncertainty estimates to identify and control unreliable predictions. The research will focus on improving factual reliability, localisation accuracy, uncertainty calibration, and generalisability in AI-assisted chest X-ray interpretation and report generation.
Aim
- To develop an evidence-grounded and uncertainty-aware vision-language framework for chest X-ray interpretation and report generation, with improved factual reliability, localisation accuracy, uncertainty calibration, and robustness across different datasets.
Objectives
-
Develop methods for spatially grounding radiological findings within chest X-ray images.
-
Develop a VLM-based framework that links generated report statements to corresponding visual evidence.
-
Investigate uncertainty estimation and calibration at the level of individual clinical findings.
-
Develop mechanisms for detecting, flagging, revising, or suppressing unsupported or unreliable generated findings.
-
Evaluate the proposed framework for factual accuracy, hallucination reduction, generalisability, and reliability.
-
Investigate performance across common and rare thoracic abnormalities and under cross-dataset distribution shifts.
Methodology
-
Baseline Development and Benchmarking: Existing chest X-ray classification, localisation, segmentation, and VLM-based report-generation approaches will be implemented or reproduced to establish baseline performance.
-
Visual Evidence Grounding: Methods will be developed to identify clinically relevant anatomical or pathological regions and associate these regions with corresponding clinical findings.
-
Evidence-Grounded Report Generation: Spatially grounded information will be integrated into a medical VLM so that report generation is conditioned on both global image information and relevant local visual evidence.
-
Uncertainty-Aware Hallucination Control: Uncertainty will be combined with visual grounding to determine whether generated findings are sufficiently supported by the image.
-
Evaluation and Generalisation: The complete framework will be compared with baseline and state-of-the-art methods using internal and cross-dataset evaluation, ablation studies, rare-abnormality analysis, and robustness testing.
Timeline
-
Year 1: Literature review, dataset selection and preparation, implementation of baseline methods, benchmarking, and initial development of visual grounding approaches.
-
Year 2: Development of the evidence-grounded VLM framework, region-to-finding alignment methods, uncertainty estimation, and hallucination-control mechanisms.
-
Year 3: Cross-dataset evaluation, rare-abnormality and robustness analysis, comprehensive ablation studies, comparison with state-of-the-art methods, publications, and thesis preparation.
Expected Outcomes
-
A novel evidence-grounded VLM framework for reliable chest X-ray interpretation and report generation.
-
A finding-level uncertainty and hallucination-control mechanism for identifying and managing unsupported clinical statements.
-
A comprehensive evaluation of robustness, factual reliability, and generalisation across publicly available chest X-ray datasets.
-
Improved alignment between generated clinical findings and supporting visual evidence.
-
Reproducible research outputs contributing to trustworthy AI in medical imaging.
-
Peer-reviewed publications contributing to research in medical imaging, computer vision, vision-language models, and trustworthy AI.