A comprehensive evaluation of IDS datasets and models with a derived framework for informed dataset selection: IDS-EeComm

Ifrah Sanober, Roohie Naaz Mir

Abstract


Benchmark IDS datasets are compromised by class imbalance, outdated attack patterns, and feature redundancy, yet their influence on model performance remains underexplored, particularly for newer datasets like HIKARI-2021. To address this, ten ML and DL models were evaluated across four datasets (NSL-KDD, UNSW-NB15, CIC-IDS2017, HIKARI-2021) under six configurations per dataset, combining balanced and imbalanced settings with three RFE-based feature subsets, using accuracy, precision, recall, F1-score, and AUC, validated through ANOVA, pairwise t-tests, and Tukey's HSD. Statistical testing confirms that ML models are comparatively dataset-independent (p > 0.05), whereas DL models exhibit strong dataset sensitivity (p ≈ 1.2e-18); HIKARI-2021 yields the most consistent performance across all models and metrics, UNSW-NB15 shows the lowest stability, and dataset choice significantly impacts DL inference time (p < 0.0001) but not ML computational cost. These results are bounded by offline, binary classification conditions with fixed hyperparameters, and do not address cross-dataset transfer, adversarial robustness, multi-class settings, or PR-AUC evaluation, meaning conclusions should not be generalised to operational IDS deployments without further validation. Practitioners can use these findings to match dataset-model-feature configurations to deployment constraints (lightweight ML models for resource-constrained environments, DL models for high-resource complex-traffic settings), supported by IDS-ReComm, a novel conceptual framework that translates empirical evidence into structured, rule-based deployment guidance. To the best of the authors' knowledge, this is among the first studies to systematically benchmark HIKARI-2021 against classic IDS datasets across both ML and DL models with rigorous statistical validation, explicit computational efficiency analysis, and a derived decision-support framework.

 

Received 23 July 2026

Accepted 09 September 2026

Published 23 September 2026


Keywords


AUC; Class Imbalance; Computational Efficiency; Deep Learning; Feature Selection; HIKARI-2021; Intrusion Detection; Machine Learning; Statistical Validation; UNSW-NB15

Full Text:

PDF

References


M. Alkasassbeh and S. Al-Haj Baddar, “Intrusion Detection Systems: A State-of-the-Art Taxonomy and Survey,” Arabian Journal for Science and Engineering, vol. 48, no. 8, pp. 10021--10064, Nov. 2022, doi: 10.1007/s13369-022-07412-1.

G. Kocher and G. Kumar, “Machine Learning and Deep Learning Methods for Intrusion Detection Systems: Recent Developments and Challenges,” Soft Computing, Jun. 2021, doi: 10.1007/s00500-021-05893-0.

Md. A. Talukder et al., “Machine Learning-Based Network Intrusion Detection for Big and Imbalanced Data Using Oversampling, Stacking Feature Embedding and Feature Extraction,” Journal of Big Data, vol. 11, no. 1, p. 33, Feb. 2024, doi: 10.1186/s40537-024-00886-w.

T. T. Nguyen and V. J. Reddi, “Deep Reinforcement Learning for Cyber Security,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 8, pp. 1–17, 2021, doi: 10.1109/tnnls.2021.3121870.

Z. Ahmad, A. Shahid Khan, C. Wai Shiang, J. Abdullah, and F. Ahmad, “Network Intrusion Detection System: A Systematic Study of Machine Learning and Deep Learning Approaches,” Transactions on Emerging Telecommunications Technologies, vol. 32, no. 1, pp. 1–29, Oct. 2020, doi: 10.1002/ett.4150.

D. Parsons, “Five Startling Findings in 2023’S ICS Cybersecurity Data,” SANS Institute. Accessed: Sep. 22, 2024. [Online]. Available: https://www.sans.org/blog/five-startling-findings-2023-ics-cybersecurity-data

R. M. Augusto and R. M. Augusto, “Strengthening ICS/OT Cyber Resilience: Learning From 2023’S Cybersecurity Incidents From Dragos’ Report,” Industrial Cyber. Accessed: Sep. 22, 2024. [Online]. Available: https://industrialcyber.co/expert/strengthening-ics-ot-cyber-resilience-learning-from-2023s-cybersecurity-incidents-from-dragos-report/

J. O. Nehinbe, “A Critical Evaluation of Datasets for Investigating IDSs and IPSs Researches,” in 2011 IEEE 10th International Conference on Cybernetic Intelligent Systems (CIS), Sep. 2011.

A. Ferriyan, A. H. Thamrin, K. Takeda, and J. Murai, “Generating Network Intrusion Detection Dataset Based on Real and Encrypted Synthetic Attack Traffic,” Applied Sciences, vol. 11, no. 17, p. 7868, Aug. 2021, doi: 10.3390/app11177868.

M. Tavallaee, E. Bagheri, W. Lu, and A. A. Ghorbani, “A Detailed Analysis of the KDD CUP 99 Data Set,” in 2009 IEEE Symposium on Computational Intelligence for Security and Defense Applications, Jul. 2009, pp. 1–6.

N. Moustafa and J. Slay, “UNSW-NB15: A Comprehensive Data Set for Network Intrusion Detection Systems (UNSW-NB15 Network Data Set),” in 2015 Military Communications and Information Systems Conference (MilCIS), Nov. 2015, pp. 1–6. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/7348942?casa_token=MCoyi5t44IQAAAAA:ZCODekObKojal3RTAApfo0ej9wNT7DtfDPGqq9Ok5CLbEyosH4XYCMu--0c2pObMagIiyUGMwMk

I. Sharafaldin, A. Gharib, A. H. Lashkari, and A. A. Ghorbani, “Towards a Reliable Intrusion Detection Benchmark Dataset,” Software Networking, no. 1, pp. 177–200, 2017, doi: 10.13052/jsn2445-9739.2017.009.

C. Zhang, F. Ruan, L. Yin, X. Chen, L. Zhai, and F. Liu, “A Deep Learning Approach for Network Intrusion Detection Based on NSL-KDD Dataset,” in 2019 IEEE 13th International Conference on Anti-counterfeiting, Security, and Identification (ASID), Oct. 2019, pp. 41–45.

V. Kumar, D. Sinha, A. K. Das, S. C. Pandey, and R. T. Goswami, “An Integrated Rule Based Intrusion Detection System: Analysis on UNSW-NB15 Data Set and the Real Time Online Dataset,” Cluster Computing, vol. 23, no. 2, pp. 1397–1418, Oct. 2019, doi: 10.1007/s10586-019-03008-x.

S. Choudhary and N. Kesswani, “Analysis of KDD-Cup’99, NSL-KDD and UNSW-NB15 Datasets Using Deep Learning in IoT,” Procedia Computer Science, vol. 167, pp. 1561–1573, 2020, doi: 10.1016/j.procs.2020.03.367.

M. Ghurab, G. Gaphari, F. Alshami, R. Alshamy, and S. Othman, “A Detailed Analysis of Benchmark Datasets for Network Intrusion Detection System,” Asian Journal of Research in Computer Science, pp. 14–33, Apr. 2021, doi: 10.9734/ajrcos/2021/v7i430185.

G. Engelen, V. Rimmer, and W. Joosen, “Troubleshooting an Intrusion Detection Dataset: The CICIDS2017 Case Study,” in 2021 IEEE Security and Privacy Workshops (SPW), May 2021, pp. 7–12. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/9474286

L. Liu, G. Engelen, T. Lynar, D. Essam, and W. Joosen, “Error Prevalence in NIDS Datasets: A Case Study on CIC-IDS-2017 and CSE-CIC-IDS-2018,” in 2022 IEEE Conference on Communications and Network Security (CNS), Oct. 2022, pp. 254–262.

M. Lanvin, P.-F. Gimenez, Y. Han, F. Majorczyk, L. Mé, and É. Totel, “Errors in the CICIDS2017 Dataset and the Significant Differences in Detection Performances It Makes,” Lecture Notes in Computer Science, pp. 18–33, 2023, doi: 10.1007/978-3-031-31108-6_2.

R. Fernandes, J. Silva, Ó. Ribeiro, I. Portela, and N. Lopes, “The Impact of Identifiable Features in ML Classification Algorithms With the HIKARI-2021 Dataset,” in 2023 11th International Symposium on Digital Forensics and Security (ISDFS), May 2023, pp. 1–5.

D. Kwon, R.-M. Neagu, P. Rasakonda, J. T. Ryu, and J. Kim, “Evaluating Unbalanced Network Data for Attack Detection,” in Proceedings of the 2023 on Systems and Network Telemetry and Analytics, in SNTA ’23. New York, NY, USA: Association for Computing Machinery, Jul. 2023, pp. 23–26.

Z. Zoghi and G. Serpen, “UNSW‐NB15 Computer Security Dataset: Analysis Through Visualization,” Security and Privacy, vol. 7, no. 1, Jun. 2023, doi: 10.1002/spy2.331.

S. Layeghy, M. Gallagher, and M. Portmann, “Benchmarking the Benchmark — Comparing Synthetic and Real-World Network IDS Datasets,” Journal of Information Security and Applications, vol. 80, p. 103689, Feb. 2024, doi: 10.1016/j.jisa.2023.103689.

M. Catillo, A. Del Vecchio, A. Pecchia, and U. Villano, “Transferability of Machine Learning Models Learned From Public Intrusion Detection Datasets: The CICIDS2017 Case Study,” Software Quality Journal, vol. 30, no. 4, pp. 955–981, Mar. 2022, doi: 10.1007/s11219-022-09587-0.

M. Catillo, A. Del Vecchio, L. Ocone, A. Pecchia, and U. Villano, “USB-IDS-1: A Public Multilayer Dataset of Labeled Network Flows for IDS Evaluation,” in 2021 51st Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W), Jun. 2021, pp. 1–6.

H. Jha, M. Khanna, H. Jhawar, and R. Jindal, “Performance Analysis of Deep Neural Network for Intrusion Detection Systems,” Lecture Notes in Networks and Systems, pp. 445–455, 2023, doi: 10.1007/978-981-99-3758-5_41.

S. J. Stolfo, W. Lee, W. Fan, A. Prodromidis, and P. K. Chan, “Index of /Databases/Kddcup99,” Uci.edu. [Online]. Available: https://kdd.ics.uci.edu/databases/kddcup99/

A. Ferriyan, A. H. Thamrin, J. Murai, and K. Takeda, “HIKARI-2021: Generating Network Intrusion Detection Dataset Based on Real and Encrypted Synthetic Attack Traffic,” zenodo.org, 2024, doi: 10.5281/zenodo.5199540.

N. Moustafa and J. Slay, “The UNSW-NB15 Dataset | UNSW Research,” research.unsw.edu.au. Accessed: Jun. 13, 2024. [Online]. Available: https://research.unsw.edu.au/projects/unsw-nb15-dataset

I. Sharafaldin, A. H. Lashkari, and A. A. Ghorbani, “Index of /CICDataset/CIC-IDS-2017/Dataset,” 174.165.80. Accessed: Jun. 13, 2024. [Online]. Available: https://205.174.165.80/CICDataset/CIC-IDS-2017/Dataset/

D. Elreedy and A. F. Atiya, “A Comprehensive Analysis of Synthetic Minority Oversampling Technique (SMOTE) for Handling Class Imbalance,” Information Sciences, vol. 505, pp. 32–64, Dec. 2019, doi: 10.1016/j.ins.2019.07.070.

X. Deng, Q. Liu, Y. Deng, and S. Mahadevan, “An Improved Method to Construct Basic Probability Assignment Based on the Confusion Matrix for Classification Problem,” Information Sciences, vol. 340–341, pp. 250–261, May 2016, doi: 10.1016/j.ins.2016.01.033.

F. Gorunescu, Data Mining: Concepts, Models and Techniques. Berlin, Heidelberg: Springer Berlin Heidelberg, 2011.

G. Sameera, R. V. Vardhan, and K. V. S. Sarma, “Binary Classification Using Multivariate Receiver Operating Characteristic Curve for Continuous Data,” Journal of Biopharmaceutical Statistics, vol. 26, no. 3, pp. 421–431, May 2015, doi: 10.1080/10543406.2015.1052479.




DOI: https://dx.doi.org/10.21622/ACE.2026.06.2.2333

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Ifrah Sanober, Roohie Naaz Mir


Advances in Computing and Engineering

E-ISSN: 2735-5985

P-ISSN: 2735-5977

 

Published by:

Academy Publishing Center (APC)

Arab Academy for Science, Technology and Maritime Transport (AASTMT)

Alexandria, Egypt

ace@aast.edu