Khóa luận tập trung giải quyết bài toán phát hiện sớm các dấu hiệu bóp méo, gian lận trên báo cáo tài chính của 221 doanh nghiệp ngành bán lẻ niêm yết tại Việt Nam giai đoạn 2021-2025 bằng cách triển khai một quy trình học máy (Machine Learning) toàn diện theo chuẩn CRISP-DM. Để khắc phục rào cản về việc thiếu hụt các dữ liệu gian lận được công bố chính thức, nghiên cứu đã sử dụng mô hình điểm số Beneish M-Score làm bộ lọc đại diện để gán nhãn phân lớp (bình thường/bất thường) cho toàn bộ tập dữ liệu. Từ đó, 12 đặc trưng tài chính cốt lõi đã được trích xuất và tối ưu hóa để làm đầu vào huấn luyện cho ba thuật toán phân loại: Hồi quy Logistics và Quản lý chuỗi cung ứngistic, Random Forest và XGBoost.
Từ khóa:
Khai phá dữ liệu, Phát hiện bất thường, Báo cáo tài chính
Abstract:
This thesis focuses on solving the problem of early detection of manipulation and anomalies in the financial statements of 221 listed retail companies in Vietnam during the 2021-2025 period by deploying a comprehensive machine learning pipeline based on the CRISP-DM framework. To overcome the barrier of lacking officially published fraud data, the study utilized the Beneish M-Score model as a proxy filter to label the entire dataset into binary classes (normal/anomalous). From this foundation, 12 core financial features were extracted and optimized to serve as training inputs for three classification algorithms: Logistics và Quản lý chuỗi cung ứngistic Regression, Random Forest, and XGBoost. Experimental results proved that XGBoost is the most optimal algorithmic architecture for this dataset, achieving an excellent balance between correctly identifying risks and minimizing false alarms. Notably, through model explainability techniques, the research identified the Cash Flow from Operations to Total Debt ratio (CFO_to_Debt) as the most critical warning signal for the retail sector. In terms of practical application, the entire data processing and detection workflow is proposed to be packaged into an automated risk-warning Dashboard system, providing a practical AI-driven tool to help state regulators, auditing firms, and investors enhance the efficiency of financial monitoring.
Key word:
Data mining, Anomaly detection, Financial statement analysis
| Ngôn ngữ: | vie |
| Tác giả: | Lê Đình Thành (2022604343) |
| Người đóng góp: | GVHD: Dương Thị Hoàn |
| Thông tin nhan đề: | Ứng dụng khai phá dữ liệu trong phát hiện bất thường trên báo cáo tài chính tại các doanh nghiệp ngành bán lẻ niêm yết ở Việt Nam. Data Mining Applications for Anomaly Detection in Financial Statements of Listed Retail Enterprises in Vietnam |
| Nhà xuất bản: | Đại học Công nghiệp Hà Nội |
| Mô tả vật lý: | 86tr. |
| Năm xuất bản: | 2026 |
Sử dụng ứng dụng Libol Bookworm quét QRCode này để mượn và đọc tài liệu)
(Lưu ý: Sử dụng ứng dụng Bookworm để xem đầy đủ tài liệu. Bạn đọc có thể tải Bookworm từ App Store hoặc Google play với từ khóa "Libol Bookworm”)