Using Knowledge Graphs and Reinforcement Learning for Malware Analysis

Submitted by grigby1 on Tue, 04/27/2021 - 2:40pm

Title	Using Knowledge Graphs and Reinforcement Learning for Malware Analysis
Publication Type	Conference Paper
Year of Publication	2020
Authors	Piplai, A., Ranade, P., Kotal, A., Mittal, S., Narayanan, S. N., Joshi, A.
Conference Name	2020 IEEE International Conference on Big Data (Big Data)
Date Published	Dec. 2020
Publisher	IEEE
ISBN Number	978-1-7281-6251-5
Keywords	artificial intelligence, Big Data, computer security, cybersecurity, graph theory, Knowledge graphs, machine learning algorithms, Malware, malware analysis, Measurement, Metrics, pubcrawl, reinforcement learning, resilience, Resiliency, Scalability, security, Semantics
Abstract	Machine learning algorithms used to detect attacks are limited by the fact that they cannot incorporate the back-ground knowledge that an analyst has. This limits their suitability in detecting new attacks. Reinforcement learning is different from traditional machine learning algorithms used in the cybersecurity domain. Compared to traditional ML algorithms, reinforcement learning does not need a mapping of the input-output space or a specific user-defined metric to compare data points. This is important for the cybersecurity domain, especially for malware detection and mitigation, as not all problems have a single, known, correct answer. Often, security researchers have to resort to guided trial and error to understand the presence of a malware and mitigate it.In this paper, we incorporate prior knowledge, represented as Cybersecurity Knowledge Graphs (CKGs), to guide the exploration of an RL algorithm to detect malware. CKGs capture semantic relationships between cyber-entities, including that mined from open source. Instead of trying out random guesses and observing the change in the environment, we aim to take the help of verified knowledge about cyber-attack to guide our reinforcement learning algorithm to effectively identify ways to detect the presence of malicious filenames so that they can be deleted to mitigate a cyber-attack. We show that such a guided system outperforms a base RL system in detecting malware.
URL	https://ieeexplore.ieee.org/document/9378491
DOI	10.1109/BigData50022.2020.9378491
Citation Key	piplai_using_2020

Groups:

Science of Security VO