CHÚC MỪNG SINH VIÊN LỚP HTTT2024.2 CÓ BÀI BÁO ĐƯỢC CHẤP NHẬN TẠI HỘI NGHỊ QUỐC TẾ CLEF 2026 (SCOPUS)
CLEF 2026 (Conference and Labs of the Evaluation Forum) là hội nghị quốc tế về đánh giá các hệ thống truy xuất và truy cập thông tin, dự kiến được tổ chức tại Jena, Đức, từ ngày 21 – 24/09/2026.
CLEF là diễn đàn quốc tế dành cho các nhà nghiên cứu, chuyên gia và các nhóm nghiên cứu nhằm trình bày, đánh giá và so sánh các phương pháp, hệ thống trong lĩnh vực truy xuất thông tin, xử lý ngôn ngữ và các bài toán liên quan đến hiểu dữ liệu đa phương thức. Hội nghị đồng thời tổ chức nhiều chiến dịch đánh giá (evaluation campaigns), tạo môi trường nghiên cứu và thử nghiệm các phương pháp mới trên các bộ dữ liệu và bài toán thực tế.
Mã ISSN: 1613-0073
Nhà xuất bản: CEUR Workshop Proceedings (CEUR-WS)
Thông tin Hội nghị: https://www.clef-initiative.eu/
Tên bài báo: “UIT_NEWRON at ImageCLEFmed MEDVQA-GI 2026: Canonical Post-Processing for Gastrointestinal Visual Question Answering”
Sinh viên thực hiện:
- Nguyễn Nhật Tân – MSSV: 24521580 – Lớp HTTT2024.2
- Huỳnh Đặng Nhật Bảo – MSSV: 24520157 – Lớp CNTT2024.1
Giảng viên hướng dẫn: TS. Nguyễn Tất Bảo Thiện
Abstract: This paper presents the system design and experimental findings of the UIT_NEWRON team at the ImageCLEFmed MEDVQA-GI-2026 evaluation campaign, conducted on the Kvasir-VQA benchmark for gastrointestinal (GI) endoscopy understanding. For Task 1 (Clinically Relevant Visual Question Answering), we propose a two-stage pipeline comprising: (i) Qwen2.5-VL-7B-Instruct fine-tuned via Quantized Low-Rank Adaptation (QLoRA) in a 4-bit NormalFloat (NF4) configuration for parameter-efficient domain adaptation, and (ii) a deterministic rule-based canonical post-processing framework (Version 7) that maps free-form generative outputs into medically standardised clinical terms. On the public validation set, our system achieved BLEU-1 of 0.6397, ROUGE-L of 0.8003, and METEOR of 0.4354. A significant gap between validation and private test performance motivated a rigorous post-submission failure analysis, which uncovered two critical defects: a branch-ordering error causing incorrect routing of polyp-removal queries, and an aggressive context-gating mechanism that collapsed to hardcoded defaults under distribution shift. Both defects were diagnosed and corrected in Version 8. Beyond the submitted system, this paper makes three concrete contributions: (1) a replicable two-stage pipeline for clinical GI VQA combining parameter-efficient fine-tuning with deterministic output normalisation; (2) an empirical characterisation of branch-ordering sensitivity and context-gate fragility in rule-based post-processors under distribution shift; and (3) actionable remediation strategies — including soft confidence-weighted routing and multi-query majority voting — validated through post-competition analysis.











