CHÚC MỪNG SINH VIÊN LỚP HTTT2023.2 CÓ BÀI BÁO ĐƯỢC CHẤP NHẬN TẠI HỘI NGHỊ QUỐC TẾ CLEF 2026 (SCOPUS)
CLEF 2026 (Conference and Labs of the Evaluation Forum) là hội nghị quốc tế về đánh giá các hệ thống truy xuất và truy cập thông tin, dự kiến được tổ chức tại Jena, Đức, từ ngày 21 – 24/09/2026.
CLEF là diễn đàn quốc tế dành cho các nhà nghiên cứu, chuyên gia và các nhóm nghiên cứu nhằm trình bày, đánh giá và so sánh các phương pháp, hệ thống trong lĩnh vực truy xuất thông tin, xử lý ngôn ngữ và các bài toán liên quan đến hiểu dữ liệu đa phương thức. Hội nghị đồng thời tổ chức nhiều chiến dịch đánh giá (evaluation campaigns), tạo môi trường nghiên cứu và thử nghiệm các phương pháp mới trên các bộ dữ liệu và bài toán thực tế.
Mã ISSN: 1613-0073
Nhà xuất bản: CEUR Workshop Proceedings (CEUR-WS)
Thông tin Hội nghị: https://www.clef-initiative.eu/
Tên bài báo: “PrompterSquad at CheckThat! 2026: Self-Distilled Hard Negatives and a Hybrid Pairwise–Listwise Ranker for Fact-Checking Numerical Claims”
Sinh viên thực hiện: Phạm Thái Sơn – MSSV: 23521361 – Lớp HTTT2023.2 – Tác giả chính
Giảng viên hướng dẫn: ThS. Nguyễn Hồ Duy Trí
Abstract:
This paper describes the PrompterSquad system submitted to Task 2 of the CLEF 2026 CheckThat! Lab, which addresses test-time scaling for fact-checking of numerical claims. Given a claim, a set of retrieved evidence passages, and multiple candidate reasoning traces produced by a Large Language Model (LLM), the task requires participants to (i) rank the reasoning traces by their utility in leading to the correct verdict, and (ii) predict the final verdict from the top-ranked traces.
We frame this as a learning-to-rank problem and fine-tune microsoft/deberta-v3-base as a pointwise scorer using a hybrid loss that combines a pairwise margin objective with a listwise ListMLE term. We further apply LoRA-based parameter-efficient adaptation and a self-distilled hard negative mining procedure that begins after the second epoch.
On the official English test set, our system achieves an average score of 0.4138 (Macro-F1 = 0.5967, Recall@5 = 0.2308), ranking 9th of 9 teams within a spread of only 0.0148 points. A controlled ablation on the held-out validation half identifies the hybrid loss and hard negative mining as the two training-side components contributing meaningful gains; the inference-time refinements (Monte Carlo dropout averaging and multi-k weighted voting), which we deployed in our submission, do not produce a measurable benefit in the ablation and we report this honestly.
Through an oracle upper-bound analysis we further uncover two structural properties of the task: (i) 29.4% of validation claims have no positive coverage—no trace in their pool has an intermediate verdict matching the gold label—imposing a hard ceiling of approximately 0.67 Macro-F1 on any verifier, and explaining the unusually narrow 0.015-point spread among the top-9 teams; and (ii) the Conflicting class is learnably hard rather than intrinsically hard—an oracle ranker achieves 0.68 F1 on Conflicting, more than three times the best participating system, indicating that the bottleneck is verifier learning rather than data ambiguity. We release our code and trained model checkpoints to support reproducibility.











