CHÚC MỪNG NHÓM SINH VIÊN LỚP HTTT2022.1 NGÀNH HỆ THỐNG THÔNG TIN CÓ BÀI BÁO ĐƯỢC CHẤP NHẬN TẠI HỘI NGHỊ KHOA HỌC QUỐC TẾ ICSCC 2026
The 12th International Conference on Smart Computing and Communication (ICSCC 2026) là hội nghị khoa học quốc tế trong lĩnh vực Điện toán thông minh (Smart Computing) và Truyền thông (Communication), là diễn đàn dành cho các nhà nghiên cứu, học giả và chuyên gia trao đổi các kết quả nghiên cứu, tiến bộ công nghệ trong các lĩnh vực điện toán, công nghệ thông tin, truyền thông, điện tử và hệ thống năng lượng. Với chủ đề “From Intelligence to Impact: Building Sustainable Smart Communities”, hội nghị tập trung vào các công nghệ thông minh và ứng dụng thực tiễn nhằm xây dựng cộng đồng thông minh, bền vững.
Hội nghị được tổ chức theo hình thức trực tiếp kết hợp trực tuyến tại Denpasar, Bali, Indonesia từ ngày 23 đến 25/07/2026. Kỷ yếu hội nghị được xuất bản trong IET Conference Proceedings (Online ISSN: 2732-4494) bởi The Institution of Engineering and Technology (IET), Vương quốc Anh, và theo công bố của hội nghị sẽ được lập chỉ mục Scopus, Ei Compendex và IET Inspec.
Thông tin chi tiết về hội nghị có thể tham khảo tại: https://icscc.undiknas.ac.id/
Kỷ yếu các kỳ trước (IEEE Xplore): https://ieeexplore.ieee.org/xpl/conhome/1833024/all-proceedings
Tên bài báo: “A Scalable Selective Stacking Framework for Interpretable Customer Churn Prediction”
Sinh viên thực hiện:
- An Văn Kết (MSSV: 22520595) – Lớp HTTT2022.1
- Lý Quan Long (MSSV: 22520814) – Lớp HTTT2022.1
Giảng viên hướng dẫn: ThS. Nguyễn Hồ Duy Tri, ThS. Nguyễn Hồ Duy Trí
Abstract: Customer churn prediction is critical in enterprise analytics, as retaining existing customers is more cost-effective than acquiring new ones. However, this problem remains challenging due to class imbalance, heterogeneous feature types, and scalability limitations of fixed ensemble structures. Although Apache Spark-based approaches have been applied to individual pipeline stages, a cohesive end-to-end Spark-native framework remains underexplored. This study proposes a selective stacking framework built on Apache Spark and evaluates it through performance analysis and small-cluster scalability experiments. To address class imbalance in mixed-type data, a K-Means-guided SMOTE-NC strategy is introduced. The feature space is partitioned into coherent subgroups, enabling partition-wise oversampling that preserves SMOTE’s interpolation characteristics in distributed settings while supporting nominal features. For ensemble selection, candidate subsets are evaluated using model competence and pairwise disagreement metrics, with the optimal subset selected via exhaustive search and out-of-fold F1. In addition, a two-level post-hoc SHAP (SHapley Additive exPlanations) analysis provides interpretability at both the feature and model levels. Experimental results across benchmark datasets demonstrate competitive performance against base models and full-stacking baselines, with scalability experiments yielding 1.55–1.74× speedups. Statistical reliability is further confirmed through McNemar’s test.











