logo

Machine Learning Algorithms for Classification Tasks after Valuable Data Preprocessing: Evidence from Economic Indicator Data

Authors

  • Azizjon Meliboev

    Digital technology and Mathematics, Kokand University
    Author

Keywords:

machine learning; classification; data preprocessing; economic indicators; monetary aggregates; ROC curve; confusion matrix

Abstract

Machine learning algorithms are increasingly used for classification tasks in artificial intelligence, cybersecurity, economics, and data-driven decision-making. However, the quality of classification results depends not only on the selected algorithm but also on the quality of data preprocessing. This study investigates the role of valuable data preprocessing in improving machine learning-based classification performance using economic indicator data. The dataset contains 147 monthly observations from February 2013 to April 2025 and includes monetary and deposit-related variables such as broad money supply, national currency money supply, narrow money supply, cash in circulation, demand deposits, other national currency deposits, and foreign currency deposits expressed in national currency equivalent. The research follows the IMRAD structure and applies exploratory data analysis, missing value checking, descriptive statistics, distribution analysis, correlation analysis, feature preparation, and classification modeling. Four machine learning algorithms were applied: Decision Tree, Random Forest, Logistic Regression, and Support Vector Machine. The models were evaluated using ROC curves and confusion matrix-based metrics, including accuracy, precision, recall, specificity, and F1-score. The results show that preprocessing produced a clean dataset with no missing values and revealed strong positive correlations among the selected variables. The final classification result achieved an accuracy of approximately 98.83%, precision of 97.87%, recall of 100%, and F1-score of 98.92%. The findings confirm that systematic preprocessing and exploratory analysis are essential for obtaining reliable machine learning classification results. At the same time, because economic indicators have strong time-series characteristics, future studies should apply chronological validation and rolling-window testing to avoid overly optimistic performance estimation.

References

Breiman, L. (2001). Random forests. Machine Learning, 45, 5-32.

Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20, 273-297.

Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861-874.

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., et al. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825-2830.

Han, J., Kamber, M., & Pei, J. (2012). Data Mining: Concepts and Techniques. Morgan Kaufmann.

Geron, A. (2019). Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow. O'Reilly Media.

Quinlan, J. R. (1986). Induction of decision trees. Machine Learning, 1, 81-106.

James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An Introduction to Statistical Learning. Springer.

Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning. Springer.

Bishop, C. M. (2006). Pattern Recognition and Machine Learning. Springer.

Downloads

Additional Files

Published

2026-06-16

How to Cite

Meliboev, A. (2026). Machine Learning Algorithms for Classification Tasks after Valuable Data Preprocessing: Evidence from Economic Indicator Data. TLEP – International Journal of Multidiscipline, 3(6), 84-92. https://www.tlepub.org/index.php/1/article/view/1047