A Trigger Aware, Event Centric, and Uncertainty Calibrated Neuro-Symbolic Framework for Actionable Cyber Threat Intelligence from Indonesian Online News
DOI:
https://doi.org/10.65780/bima.v1i5.31Keywords:
cyber threat intelligence;, Indonesian news;, event clustering;, trigger-aware learning;, conformal predictionAbstract
Online news can provide timely cyberthreat signals, but duplicative reporting, fragmented event descriptions, resource-constrained Indonesian language text, and uncalibrated model confidence limit its operational use. This study presents A Trigger-Aware, Event-Centric, and Uncertainty-Calibrated Neuro-Symbolic Framework for Actionable Cyber Threat Intelligence from Indonesian Online News (TRACE-CTI-ID), a proof-of-concept framework that integrates exact deduplication, event-centric clustering, trigger-aware semantic representation, neuro-symbolic fusion, ordinal risk estimation, conformal uncertainty, mitigation mapping, and an event-centric knowledge graph. The experiment used 711 Liputan6 records collected on March 14, 2025. Exact deduplication reduced the corpus to 79 unique headlines, which were automatically consolidated into 26 events. Splitting the separate events resulted in 30 training articles, 5 calibration articles, and 44 test articles with zero event leakage. The calibrated neuro-symbolic model achieved a micro-F1 of 0.283 and a macro-F1 of 0.441, outperforming the baseline TF-IDF of 0.074 and 0.013, respectively. However, ordinal severity prediction remained weak with an accuracy of 0.136, a macro-F1 of 0.138, and a mean absolute error of 2.023. Conformal coverage was also unstable, and the abstention mechanism did not direct uncertain articles to human review. These findings demonstrate the technical feasibility of the integrated pipeline while also demonstrating that silver labeled, title only, and single source data are insufficient for final operational validation. Therefore, the key contribution is a transparent, leak aware evaluation architecture and protocol that can be strengthened through full text collection from multiple sources and independent expert annotation.
Downloads
References
[1] B. Jordan, R. Piazza, and T. Darley, eds., “STIX Version 2.1,” OASIS Standard, Jun. 2021.
[2] MITRE, “MITRE ATT&CK: A knowledge base of adversary tactics and techniques,” 2026. [Online]. Available: https://attack.mitre.org/. Accessed: Jul. 22, 2026.
[3] C. Pascoe, S. Quinn, and K. Scarfone, “The NIST Cybersecurity Framework (CSF) 2.0,” NIST CSWP 29, Feb. 2024, doi: 10.6028/NIST.CSWP.29.
[4] J. Liu, J. Yan, J. Jiang, Y. He, X. Wang, Z. Jiang, P. Yang, and N. Li, “TriCTI: An actionable cyber threat intelligence discovery system via trigger-enhanced neural network,” Cybersecurity, vol. 5, no. 1, Art. no. 8, 2022, doi: 10.1186/s42400-022-00110-3.
[5] Y. You, J. Jiang, Z. Jiang, P. Yang, B. Liu, H. Feng, X. Wang, and N. Li, “TIM: Threat context-enhanced TTP intelligence mining on unstructured threat data,” Cybersecurity, vol. 5, Art. no. 3, 2022, doi: 10.1186/s42400-021-00106-5.
[6] H. Jo, Y. Lee, and S. Shin, “Vulcan: Automatic extraction and analysis of cyber threat intelligence from unstructured text,” Computers & Security, vol. 120, Art. no. 102763, Sep. 2022, doi: 10.1016/j.cose.2022.102763.
[7] S. Silvestri, S. Islam, D. Amelin, G. Weiler, S. Papastergiou, et al., “Cyber threat assessment and management for securing healthcare ecosystems using natural language processing,” International Journal of Information Security, vol. 23, pp. 31-50, 2024, doi: 10.1007/s10207-023-00769-w.
[8] G. Wang, P. Liu, J. Huang, H. Bin, X. Wang, and H. Zhu, “KnowCTI: Knowledge-based cyber threat intelligence entity and relation extraction,” Computers & Security, vol. 141, Art. no. 103824, Jun. 2024, doi: 10.1016/j.cose.2024.103824.
[9] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. 34th Int. Conf. Machine Learning, PMLR, vol. 70, 2017, pp. 1321-1330.
[10] M. Campos, A. Farinhas, C. Zerva, M. A. T. Figueiredo, and A. F. T. Martins, “Conformal prediction for natural language processing: A survey,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 1497-1516, 2024, doi: 10.1162/tacl_a_00715.
[11] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP,” in Proc. 28th Int. Conf. Computational Linguistics, 2020, pp. 757-770, doi: 10.18653/v1/2020.coling-main.66.
[12] B. Wilie et al., “IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding,” in Proc. 1st AACL and 10th IJCNLP, 2020, pp. 843-857, doi: 10.18653/v1/2020.aacl-main.85.
[13] A. Conneau et al., “Unsupervised cross-lingual representation learning at scale,” in Proc. 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 8440-8451, doi: 10.18653/v1/2020.acl-main.747.
[14] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. EMNLP-IJCNLP, 2019, pp. 3982-3992, doi: 10.18653/v1/D19-1410.
[15] OASIS Cyber Threat Intelligence Technical Committee, “TAXII Version 2.1,” OASIS Standard, Jun. 2021.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 BIMA (Bulletin of Intelligent Machines and Algorithms)

This work is licensed under a Creative Commons Attribution 4.0 International License.