Classification and traceability method and device for malicious software in smart power plant and storage medium
By combining BERT and CNN models, static and dynamic information of API call sequences are extracted, and the problem of low detection accuracy of malware in smart power plants is solved, higher detection accuracy and traceability are achieved, and the network security of the power plant is improved.
Patent Information
- Application Number
- CN202510112345.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
AI Technical Summary
Existing malware detection methods in smart power plants are difficult to effectively detect encrypted and shelled malware, and relying solely on static information is difficult to capture dynamic changes in process behavior.
The BERT model and CNN model in deep learning technology are adopted to comprehensively consider the static element information and context timing information of the API call sequence, and the semantic features of the text of the ultra-long API call sequence are extracted to improve the multi-classification detection accuracy of malware.
It improves the multi-class detection accuracy of malware, enhances traceability, strong adaptability and scalability, and improves the level of network security protection for power plants.
Smart Images

Figure CN120105415A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of malware classification in smart power plants, and in particular to a method, device and storage medium for classifying and tracing malware in smart power plants. Background Art
[0002] Smart power plants are digitalized and intelligentized traditional power plants based on digitalization and informatization, using advanced information technology to monitor, control and manage the entire process of power production. Smart power plants combine advanced information technologies such as sensor measurement, information communication, automatic control, artificial intelligence, cloud computing, big data, and three-dimensional visualization with industrialized technologies and power plant management technologies in the power generation process, and are highly integrated with the power plant's infrastructure. Data in smart power plants runs through the entire process of enterprise production management. This requires smart power plants to have high security protection measures in the face of network threats, to be able to recover quickly when attacked by malware, and to effectively avoid accidents. The main methods of static malware detection are mostly to extract static information from non-encrypted, non-packed (including successfully unpacked) executable binary files, and then use machine learning algorithms for classification, which makes it difficult to detect encrypted and packed malware. The main method of dynamic detection is to extract the API call information of malware through the API extraction algorithm, extract the feature vectors in the API call information using the feature extraction algorithm, and finally use machine learning methods to classify the feature vectors, thereby achieving malware classification detection.
[0003] Most of the current API feature extraction algorithms use statistical methods of static information (such as the frequency of occurrence of different types of sequence elements in a process). Through this type of method, some process execution features can be easily obtained. However, in addition to the static information of the process, the dynamic behavior of the process is more critical for malicious detection, because many malicious processes are likely to imitate normal process behavior most of the time, and only attack the system a few times. Therefore, it is difficult to capture the changes in process behavior over time using only static information. Summary of the invention
[0004] The main technical problem solved by this application is to provide a method for classifying and tracing malware in a smart power plant to solve the above-mentioned problems.
[0005] To solve the above technical problems, a technical solution adopted in this application is to provide a method for classifying and tracing malware in a smart power plant, comprising the following steps: using the BERT model and CNN model in deep learning technology, comprehensively considering the static element information and contextual timing information of the API call sequence, and extracting the semantic features of the ultra-long API call sequence text to improve the multi-classification detection accuracy of malware.
[0006] The present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method when executing the computer program.
[0007] The present application also provides a computer-readable storage medium having a computer program stored thereon, and the steps of the method described above are implemented when the computer program is executed by a processor.
[0008] The beneficial effects of the present application are as follows: In the present application, the BERT model and the CNN model in the deep learning technology are adopted, and the static element information and the contextual timing information of the API call sequence are comprehensively considered, which has the advantages of improving the accuracy of multi-classification detection, enhancing the traceability capability, strong adaptability and scalability, and improving the level of network security protection of power plants. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a flow chart of intelligent classification of malware based on deep learning according to an embodiment of the present application;
[0010] Figure 2 It is a schematic diagram of a malware intelligent classification detection model based on deep learning and API behavior according to an embodiment of the present application;
[0011] Figure 3 It is a schematic diagram of an advanced network threat source tracing and evidence collection method based on behavior and context information according to an embodiment of the present application;
[0012] Figure 4 It is a schematic diagram of an attack reconstruction example according to an embodiment of the present application. DETAILED DESCRIPTION
[0013] In order to facilitate the understanding of the present application, the present application is described in more detail below in conjunction with the accompanying drawings and specific embodiments. The preferred embodiments of the present application are provided in the accompanying drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive.
[0014] It should be noted that, unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art of the present application. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more related listed items.
[0015] Figure 1An embodiment of the method for classifying and tracing malware in a smart power plant of the present application is shown, including: using the BERT model and CNN model in deep learning technology, comprehensively considering the static element information and contextual timing information of the API call sequence, and extracting the semantic features of the ultra-long API call sequence text to improve the multi-classification detection accuracy of malware.
[0016] This application uses deep learning methods to make full use of the semantic information in API sequences to improve the accuracy and reliability of malware classification. By building and training a suitable deep learning model, it can better capture the complex features in API sequences and achieve more accurate malware detection and classification.
[0017] like Figure 1 As shown in the figure, this application adopts the BERT model and CNN model in deep learning technology, comprehensively considers the static element information and contextual timing information of the API call sequence, and extracts the semantic features of the ultra-long API call sequence text to improve the multi-classification detection accuracy of malware, providing an effective way to intelligently solve the detection problem of cyberspace malware. The model framework is shown in the figure. Figure 2 As shown:
[0018] 1) Input and preprocessing layers
[0019] In this application, malware is used as input to the input layer, runs in a sandbox environment, and its dynamic behavior logs are collected by the data acquisition and fusion processing model. This method can capture the detailed behavior of malware during actual operation, including all API calls, file operations, network communications, etc. However, if all API call information of malware is sent to the system for analysis without processing, it will bring huge computing overhead. This will not only consume a lot of computing resources, but also may lead to a decrease in the efficiency of model training and analysis. Therefore, this application proposes an API sequence redundancy removal algorithm to reduce the computational complexity of subsequent models.
[0020] In the API call sequence, some API calls appear frequently, especially when malware performs repetitive operations. Although these repeated calls reflect the behavior patterns of malware to a certain extent, they do not provide additional useful information, but increase the redundancy of data and the computational complexity. To solve this problem, the proposed API sequence de-redundancy algorithm aims to reduce data redundancy by deleting repeated API calls or short API sequences and simplifying them into one.
[0021] Specifically, the algorithm includes the following steps:
[0022] 1. Preprocessing of API call sequence: First, preprocess the collected API call sequence to identify each API call and its parameters. Arrange these API calls in chronological order to form a complete call sequence.
[0023] 2. Identification of repeated calls: In the pre-processed API call sequence, identify repeated API calls or short sequences. These repeated calls are usually generated when malware performs looping operations or repetitive tasks.
[0024] 3. Deletion of redundant information: For identified repeated calls, they are deleted to one. For example, if an API call appears multiple times in a sequence, only the first occurrence is retained and the subsequent repeated parts are deleted. Similarly, for repeated short API sequences, only one occurrence is retained and the other repeated parts are deleted.
[0025] 4. Sequence reconstruction: Reassemble the processed API calls to form a de-redundant API call sequence. This new sequence retains the key behavioral features of the malware but removes redundant repetitive information.
[0026] Through the above steps, this application achieves effective de-redundancy of API call sequences, significantly reducing data redundancy and computational complexity. The de-redundant API sequence not only reduces the amount of data, but also retains the core behavioral characteristics of the malware, making subsequent model analysis more efficient and accurate.
[0027] The advantage of this de-redundancy algorithm is that it can significantly reduce the amount of data without losing key behavioral information, thereby reducing the consumption of computing resources and improving the efficiency of model training and analysis. This method is particularly effective for systems that need to process a large number of malware behavior logs. By reducing the redundancy of data, data processing and analysis can be performed faster, thereby improving the real-time and accuracy of malware detection and classification.
[0028] 2) BERT Layer
[0029] Most existing API feature extraction algorithms use statistical methods of static information to obtain some process execution features. However, it is difficult to capture the changes in process behavior over time using only static information. Therefore, it is necessary to conduct in-depth analysis of the order in which the API call sequence appears in order to better extract the feature vector. The BERT model structure consists of the encoder stack of the Transformer model. The BERT model extracts the semantics of the language based on comprehensive contextual information and has significant results in document vectorization.
[0030] This application uses the BERT model in deep learning technology to extract the contextual timing information of the API call sequence to provide high-quality data guarantee for subsequent detection and classification models.
[0031] like Figure 2 As shown, the BERT layer calculation flow chart is as follows:
[0032] ①Separate operation
[0033] Divide each original position vector into several position vectors, such as dividing each 512-dimensional position vector into 8 parts, and the dimension of each part becomes 64.
[0034] ②Self-attention mechanism
[0035] The attention mechanism is used to analyze contextual information and perform multiple understandings of the text, increasing the weight of important information and reducing the weight of unimportant information. The attention mechanism means using multiple attention mechanisms at the same time to enhance the understanding of the text. There are many common APIs (such as Getfilesize, Getusername, etc.) in the malware API call sequence. These common APIs have little effect on malware classification and are unimportant information. However, some malware-specific API call information plays a significant role in classifying and identifying the malware, and these API call information is important information. Therefore, the use of the attention mechanism can help identify important and unimportant information in the API call sequence to improve the accuracy of malware classification and recognition.
[0036] ③Head connection and Dropout
[0037] The values output by the self-attention mechanism are concatenated to obtain multi-dimensional information. Dropout prevents overfitting of the neural network by probabilistically deleting certain neurons.
[0038] ④Residual & Normalization
[0039] The residual method is used to add the input and output of the attention mechanism. Then, the sum is normalized to remove the correlation of the data so that it meets the conditions of independent distribution. The residual & normalization method can effectively avoid the gradient vanishing of the neural network and accelerate the learning convergence speed.
[0040] ⑤ Feedforward Neural Network
[0041] The feedforward neural network consists of a fully connected layer and an activation function, which is responsible for transforming the dimension of the vector and removing noise, making it easier for the multi-head attention mechanism to extract important information.
[0042] 3) CNN Layer
[0043] The current malware feature vector classification model has problems such as low classification accuracy, poor anti-disturbance, slow detection speed, and inability to achieve multi-classification. First, the malware feature vector extracted by the BERT model is converted into an image, and the time series data classification problem with complex coupling relationships is converted into an image classification problem to simplify the problem. Then, a CNN-based classification algorithm is used to classify the malware feature vector graph to achieve high-precision, high-speed, and high-reliability real-time classification detection of malware.
[0044] Figure 2 What is shown is the classic meta-structure of the CNN network. The real CNN network consists of multiple meta-structures and pooling layers. Figure 2 The calculation process of CNN is as follows:
[0045] ①Convolution and activation function
[0046] Convolution is an effective method for extracting image features. Generally, a square convolution kernel is used to traverse every pixel in the image. Each pixel value corresponding to the overlap area between the image and the convolution kernel is multiplied by the weight of the corresponding point in the convolution kernel, and then summed and added with the bias to finally obtain a pixel value in the output image. Generally, the ReLU activation function is used to introduce nonlinear factors, improve the expression ability of the neural network to the model, increase the sparsity of the network, and alleviate the problem of overfitting.
[0047] ②Residual
[0048] The residual method is used to add the input and output of the neural network as the input of the next layer of the neural network, so as to effectively avoid the gradient disappearance of the neural network and accelerate the learning convergence speed.
[0049] ③Pooling
[0050] The pooling operation reduces the spatial size of the data. The reduction in the number of parameters and the amount of computation enables the model to extract a wider range of features while preventing overfitting.
[0051] ④Fully connected
[0052] Convert three-dimensional data vectors into one-dimensional data, which is beneficial for subsequent classification operations.
[0053] 4) Output layer
[0054] The output layer uses Softmax (normalized exponential function) to present the multi-classification results in the form of probability. The sum of the probability prediction results is 1, and the category with the largest probability is the classification result. Finally, the malware is classified into specific malware families, such as mining programs, Trojan horses, backdoor programs, etc.
[0055] Experimental part:
[0056] ① Dataset introduction
[0057] In this application, in order to conduct malware classification analysis, we first found some public datasets with labels on the Internet and downloaded Windows malware samples. These datasets provide rich sample data for research, and these samples have been classified into different malware categories, which is helpful for subsequent classification and analysis. However, malware samples alone are not enough. In order to build a balanced dataset, additional normal samples need to be collected.
[0058] The balance of the dataset is crucial when training deep learning models. If the ratio of malware samples to normal samples in the dataset is unbalanced, the model may tend to predict the class with a larger number, resulting in poor classification performance. Therefore, in order to balance the dataset, an additional number of normal samples equal to the number of malware samples is collected. In this way, the ratio of malware and normal samples in the dataset is balanced, which can effectively prevent the model from being biased towards any one category and improve the classification accuracy and stability of the model.
[0059] After collecting a balanced dataset, an important issue is how to avoid overfitting of deep learning classification models. Overfitting means that the model performs well on the training data, but performs poorly on unseen data. This is usually due to the model remembering the noise and details in the training data during training, rather than learning the general features of the data. In order to avoid overfitting, the dataset is reasonably divided into training set, validation set and test set.
[0060] Specifically, the steps for data set division are as follows:
[0061] 1. Training set: The training set contains most of the samples in the dataset, usually accounting for 70% of the entire dataset. The training set is used to train the deep learning model, learning the characteristics and classification rules of malware and normal samples through a large amount of sample data. During the training process, the model will continuously adjust parameters to optimize the classification performance.
[0062] 2. Validation set: The validation set usually accounts for 20% of the dataset. During the model training process, the validation set is used to evaluate the performance of the model in real time and help adjust hyperparameters. By regularly verifying the performance of the model on the validation set during the training process, overfitting can be discovered and prevented in a timely manner. For example, if the accuracy of the model on the training set continues to improve, but the accuracy on the validation set does not increase synchronously, or even decreases, it indicates that overfitting may have occurred.
[0063] 3. Test set: The test set usually accounts for 10% of the data set. The test set is used to perform a final evaluation of the model after the model training is completed. The data in the test set has never been seen by the model during the training process, so it can provide an objective indicator to measure the actual classification performance of the model. The evaluation results of the test set can reflect the performance of the model in the real world.
[0064] The above dataset division method ensures that the model can perform stably during training, validation, and testing. The training set provides sufficient data for the model to learn, the validation set helps the model to adjust and optimize during the training process, and the test set is used to finally evaluate the generalization ability of the model.
[0065] The details of the final collected data set are shown in Table 1 below:
[0066] Table 1 Malware datasets
[0067]
[0068]
[0069] ②API redundancy and deduplication
[0070] The length of API call sequences of different malwares ranges widely, from a dozen to tens of thousands. In the API sequence, many APIs or API sequence fragments are called repeatedly. Therefore, this application proposes a de-redundancy algorithm that reduces APIs that are repeated three or more times to one time; and reduces API sequence fragments that are repeated three or more times to one time. In this way, the length of the API sequence can be greatly reduced and the computational efficiency can be improved. Experiments have shown that tens of thousands of word vectors can be reduced to hundreds of word vectors, and the dimensionality reduction effect is very obvious.
[0071] ④Deep learning classification algorithm
[0072] For the API sequences after de-redundancy, BERT and CNN are used to extract features and classify them;
[0073] Table 2 Hyperparameter settings and experimental environment of the BERT model
[0074] batch 256 Iterations 500 Optimizer Adam Learning Rate 0.001 Number of encoders 16 Hidden Units 768 Number of longs 12 System version Ubuntu 19.04 GPU Nvidia 4090 CPU Intel(R)Core(TM)i9-9820XCPU@3.30GHz Deep Learning Frameworks pytorch1.5 Memory 64G
[0075] The parameters of CNN are:
[0076] layer=12, input_shape=(10,100), output_shape=110
[0077] The training set and test set are divided into 4:1 ratio; the Word2vec model and deep learning classification model are trained using the training set; the classification effect is tested using the test set, and the results are shown in Table 3 below.
[0078] Table 3 Malware multi-classification results
[0079]
[0080] The analysis results show that the classification accuracy of the validation set and the test set is basically the same, proving that the deep learning classification model proposed in this application performs well in effectively preventing overfitting. This result shows that the model has stable performance when processing unknown data and has high generalization ability. In the test set, for 13 different types of malware including Trojans, the classification accuracy reached about 99%, exceeding the contract indicator requirement of 98%, fully demonstrating the practicality and reliability of the model.
[0081] The success of this model is largely attributed to the multi-head attention mechanism and hybrid neural network structure it adopts. Through the multi-head attention mechanism, the BERT model can comprehensively analyze the API call sequence from a global perspective and extract the deep semantic information hidden in the call sequence. This mechanism allows the model to focus on different parts of the sequence and capture long-range dependencies, thereby more accurately understanding and classifying the behavior patterns of malware.
[0082] In addition, the introduction of CNN adds the ability to extract information from a local perspective to the model. The sliding window feature of the convolution kernel enables it to capture local patterns in the API call sequence, thereby performing fine-grained analysis of the data. By combining the global perspective of BERT and the local perspective of CNN, the deep learning classification model proposed in this application can comprehensively analyze the API call sequence from different levels and angles, ensuring accurate identification of malware families.
[0083] This study further verifies the potential of multimodal neural networks in the field of malware detection. Traditional malware detection methods often rely on manual feature extraction and rule matching, which are difficult to cope with the rapid evolution and diverse behaviors of malware. Deep learning models, especially hybrid models based on multi-head attention mechanisms and convolutional neural networks, can automatically learn and extract features, significantly improving the accuracy and efficiency of detection.
[0084] From the experimental results, the model not only achieved the expected classification accuracy, but also demonstrated its potential in practical applications. When faced with various complex and changing malware samples, the model can still maintain a high level of accuracy, indicating that it has strong robustness and adaptability.
[0085] Based on dynamic behavior network threat source tracing and forensics, malware is often used as a carrier and is widely used in advanced network threats. Usually, the system will issue an alarm message after identifying the malicious behavior of the malware, and then the security manager will use manual or automated methods to obtain attack-related information for source tracing and forensics. However, the existing source tracing and forensics methods require a lot of computing and storage costs, and there are major problems in performance, accuracy, integrity and granularity of attack chain reconstruction. To this end, this application proposes a network threat source tracing and forensics method based on dynamic behavior and contextual information to achieve source tracing and forensics of advanced network threats. The overall design of the model is as follows: Figure 3 shown.
[0086] Depend on Figure 3 It can be seen that the advanced network threat source tracing and forensics method based on dynamic behavior and context information consists of a data collection and fusion model, a malware fine-grained semantic behavior recognition model, a labeling algorithm, an offline backup database, and an attack chain reconstruction model. The data collection and fusion model is responsible for collecting malware operation dynamic logs; the malware fine-grained semantic behavior recognition algorithm is responsible for processing dynamic behavior logs. If malicious behavior is found, the labeling algorithm is triggered. The labeling algorithm extracts features from the dynamic behavior logs to form a memory-based detection data structure, which is uploaded to the offline backup database. Then, the attack chain reconstruction model is used to analyze the data to complete the source tracing and forensics of advanced network threats.
[0087] Tag algorithm, the attack duration of advanced network threats is long and the attacker can lurk for a long time before achieving their ultimate attack goal, that is, persistence. On the one hand, there are 59 known persistence technologies, and it is costly to detect these persistence methods one by one; on the other hand, even if a process is detected as a persistent process, it cannot prove its maliciousness. At the same time, APT attacks can also be non-persistent.
[0088] Therefore, this application does not attempt to detect APT attacks by detecting persistence technology. Because attackers can lurk for a long time, or malicious programs that are downloaded by mistake can be started several days after being downloaded. These have caused the use of context information in detecting APT attacks to become difficult. In order to detect such attacks, system events should be stored for a long time, and storing system events for each terminal every day will consume GB-level hard disk space. This makes it costly for enterprise-level users to use context information to detect APT attacks, and at the same time, it also makes such systems unavailable for individual users. Although some studies have attempted to reduce system log storage as a goal, such methods can only alleviate the problem, because real APT attacks may last for several years, and in order to detect these attacks, data should also be stored for such a long time. In addition, it also takes a lot of time and computing resources to associate attack-related data in massive data. Therefore, a provenance graph (also known as a dependency graph, a dependency graph, or an information graph) is used to speed up the traversal process of the log. Real-time detection systems based on provenance graphs usually store provenance graphs in memory in order to achieve better performance when constructing provenance graphs and performing detection. However, at the same time, the graph in memory becomes larger and larger over time, which means that this type of method cannot run stably for a long time. One optimization method is to remove or integrate "expired" data over time. This can indeed solve the problem of memory explosion, but the premise of these methods is that the weight of historical data becomes lighter over time, and this assumption is not applicable in the case of APT attacks characterized by long latent time. Therefore, existing methods based on provenance graphs cannot solve the problem of increasingly large memory consumption and the problem of fast matching in graph structures.
[0089] This application proposes a label-based APT detection framework. In this framework, each process and file object is in the form of a label, which contains all the context information, feature information and behavior information required for detection. Therefore, all the information required for detection can be obtained on each object, and detection can be performed directly to achieve the purpose of rapid detection. At the same time, the detection system itself does not need to store any historical events, thereby achieving the purpose of significantly saving memory consumption. The following will introduce the definition of detection labels, data structure, label generation rules and suspicious process determination methods in turn.
[0090] It should be noted that since the core of this framework is to track the impact of information flow on each entity in the system, and the impact is generated in an orderly manner, the input data stream should be arranged in time sequence; the data collector in this system has ensured the timing of the data, but when other data streams are used as input, the event timing should be ensured.
[0091] In the traditional manual forensic analysis process, by analyzing various events related to the object and combining context information to identify its high-level semantics, the same event is given different semantics and traced. This application proposes an automated identification method for data flow, control flow, and process behavior, and pre-places the expert manual analysis method in the system in the form of rules, so that the system can imitate experts to perform automated semantic analysis. Semantics is the basis of context-based intrusion detection methods, and labels carry semantic information. Labels mainly contain the following types of information: attack behavior, suspicious code, network connection, and suspicious control.
[0092] Some definitions of tags are shown in Table 4, and only some tags are selected here. The first column is the label number, where P represents process label and F represents file label. Tags are divided into the following categories: code source (CS), behavior (Beh.), feature (Fea.) and network connection (Net.).
[0093] Table 4 Partial label list
[0094]
[0095]
[0096] Each tag can be described as a triple: 〈No, Ty, De> Each tag has a unique number, No represents the position of its bitmap in the process object or file object. Ty is the category to which the tag belongs. De is a human-readable description to describe the semantics of the tag.
[0097] In order to achieve rapid tracing and evidence collection of advanced network threats, this application proposes a memory-based automaton-like data structure to store status information of processes and files that may be involved in the attack.
[0098] Each process and file object only stores its basic information and label information, that is, status information. Each process object can be described as a five-tuple 〈Na, Pi, Cl, Ui, St>, where Na is the process name, Pi is the process ID, and Cl is its startup command line. Ui is the globally unique identifier of each process object. St represents the existing label set of the process object, that is, the current status of the process object.
[0099] Each file object can be represented as a triple 〈Na, Ui, St>, where Na is the complete path of the file. Ui is a globally unique identifier for each file object. St is the existing tag set of the file object, that is, the current status of the file.
[0100] Label generation and state transfer, when analyzing, different semantics will be given to the same event according to its context information; for example, when process A reads file B, if file B is a user's personal file, it can be inferred that the event may be an access to the user's personal information, which may be related to an attack, and ultimately lead to user information leakage; and when file B is a common system environment configuration file, this read event may only be a necessary file read when the process starts, which has nothing to do with the attack; or when file B is a downloaded file, process A accesses untrusted information from the outside, which may be related to an attack, and ultimately leads to the execution of unknown code or behavior. Therefore, the method described in this application is based on the above observations and is summarized as follows: events of the same type may have different semantics when the objects associated with them are different. The module described in this application aims to automatically track untrusted data flows, untrusted control flows, and high-value data flows. One implementation method can give objects corresponding scores or labels representing different levels. However, in order to make the final detection results interpretable and the feasibility of traceability analysis, this application will adopt a more fine-grained and semantic way to identify event semantics.
[0101] The present application realizes automatic semantic extraction of events by means of a predefined rule base, and Table 5 shows some of the rules. Each rule can be represented as a six-tuple 〈No, Ss, Ev, So, Di, De>, where No represents the number of the rule; Ev represents the type of event; Ss and So are the corresponding labels of the subject and object, one of which is an existing label and the other is a label to be generated, and De is a description of the rule. [Beh] and [CS] shown in Table 5 represent all malicious behavior labels and all suspicious code labels, respectively. In the advanced network threat tracing and forensics method designed by the present application, the definition of labels and rules can be expanded.
[0102] Table 5 Some examples of rules
[0103]
[0104]
[0105] This application uses a labeling algorithm to extract dynamic behaviors and contextual information in logs, uses a memory-based data structure to unify the data format, and analyzes it to complete the source tracing and evidence collection of advanced network threats. The method flow is as follows:
[0106] ① The attack graph to be drawn is G, which is initially empty
[0107] ②The set of tags that have been traversed is T;
[0108] ③ Traverse all tags on the object O1 and put them into queue Q;
[0109] ④When queue Q is not empty, take out a label L1 from queue Q
[0110] ⑤ Find the event corresponding to the label L1 (i.e., the event that generates the label (backward traversal), or the event that generates other labels because of the label (forward traversal)); if the label is a directly marked label (not involving two events), discard the label; if the label L1 involves another object O2, add the object O2 and the event to G, and add the corresponding label L2 in the object O2 to the queue Q;
[0111] ⑥ Add (label L1 of object O1) to the traversed label set T; (In order to prevent the traversed labels from being traversed again, each label should have a globally unique number)
[0112] ⑦ Repeat 4-6 until there are no labels in Q
[0113] ⑧Draw G
[0114] In the process of drawing the attack chain, this method only needs to traverse the relevant events without complex calculations. Generally, the restored attack chain can be drawn within 1 minute to complete the attack tracing and evidence collection.
[0115] In order to complete the tracing and evidence collection of advanced network threats more quickly, this application adopts a method of matching hot and cold databases. Among them, the hot database uses the NoSQL graph database Neo4J to store short-term data; the cold database uses the Cassandra database, thereby solving the problem of excessive computing and space consumption of hot database storage data.
[0116] 5) Experimental part: Attack process: "Watering hole attack" is one of the common methods of hacker attack. As the name suggests, it is to set up a "watering hole (trap)" on the path that the victim must pass. The most common practice is that hackers analyze the Internet activity patterns of the target, find the weaknesses of the websites that the target frequently visits, first "break into" this website and implant the attack code, and once the target visits the website, it will be "hit".
[0117] like Figure 4As shown in (a), the attacker first sets up attack server A and attack server B; then, the attacker tampers with the DNS server so that the victim is redirected to attack server A when using the firefox browser to visit a legitimate website; then, the attacker triggers the firefox vulnerability through the attack code pre-deployed on server A, allowing the attacker to remotely execute malicious code a directly in the victim's firefox process, so that a thread for executing malicious code a is created in the process space of the legitimate program firefox. Malicious code a establishes a reverse connection with attack server A; the attacker executes the hostname command in the name of the firefox process through malicious code a, and determines whether this terminal is an attack target based on the host name of the terminal; then, the tasklist command is executed to try to understand the running status of the process on the victim's terminal. This command is often used by attackers to detect information, especially to find other processes that are easily exploited and to find and shut down terminal protection software. At the same time, the attacker also accesses the Default.rdp file, which records the default configuration when the user uses rdp, especially the IP address of the last login; this information helps the attack to achieve the Pass The Hash domain penetration attack.
[0118] After determining the operating environment of the victim's terminal, the attacker uses malicious code a to download and execute an independent, customized malicious program cloud.exe. Among them, malicious code a is often called a small horse, which is usually less than a few KB and implements simple reconnaissance and attack functions to prepare for subsequent tools; the malicious program cloud.exe is called a big horse, which has complete attack functions to help attackers achieve purposes such as lurking and eavesdropping. The cloud process automatically captures information on the victim's screen to steal possible sensitive information; executes the whoami command to determine the information of the currently logged-in user. Afterwards, the screenshot and user information are transmitted to the attacker's server B, and the attacker successfully completes the attack.
[0119] Detection process: The specific process of the detection is described as follows:
[0120] 1. When a user tries to access a legitimate IP address but is directed to the attack server A set up by the attacker, the process object is marked with P1 because the process firefox accesses the external network.
[0121] 2. The attacker successfully executed remote code by exploiting the vulnerability of Firefox. When the legitimate process executes unknown code, the data collection module of this system detects this information and sends it to the detection server. Therefore, the Firefox process is split into two different objects according to its dynamic call stack: the legitimate part and the unknown part; the object composed of the thread executing the unknown code is marked with P12;
[0122] 3. When firefox is successfully hacked, the attacker first executes the hostname and tasklist commands to collect information. At this time, since hostname and tasklist are cmd commands commonly used by attackers, the firefox object that executes unknown code is marked with P5; at the same time, since the process accesses the Default.rdp file that is previously identified as containing sensitive information, the firefox object is marked with P2;
[0123] 4. At this point, the firefox object has entered a malicious state, so the system alarm is triggered. The object has features such as user interaction and visual windows, which indicate that users can perceive the operation of the malicious program, and the execution of malicious behavior may come from user operations, Trojans disguised as legitimate programs, or memory-based attacks. Furthermore, the process object has network activities and executes unknown code, which can be inferred that the unknown code may be downloaded from the network. Therefore, this attack can be identified as an "attack that infects legitimate programs"
[0124] 5. Then, the binary file is downloaded by Firefox. Since Firefox already has the label P1, it is believed that Firefox has had network activity; when it downloads the file, the operating system will generate an event that the process reads and writes the cloud.exe file. When this event is received, the detection system matches the label generation rule based on the existing label P1 of Firefox and the event type "file write", and then assigns the label F1 to the file object according to the rule, indicating that the file may contain data from the network. Obviously, if it can be distinguished whether the file is downloaded from the network, it will be more helpful to detect such attacks. However, since the Windows system uses the same system call NtCreateFile to create and open files, and system events are records of key system calls, it is impossible to distinguish whether the file is created or opened through the system log. It is worth mentioning that judging whether the file is downloaded from the network by the creation time of the file seems to be a feasible method; however, in Windows, the file creation time can be forged, and the attacker may use the method of overwriting existing files to download data. Therefore, this application uses the description of "contains data from the network" to record the inference of this rule;
[0125] 6. The firefox process starts cloud.exe to create a new process cloud. After the process instance of cloud is created, cloud.exe will be loaded; at this time, when the loading event is parsed by the data collector, the data collector will find the file according to the path and determine that it does not have a legal code signature; this information and the image loading event are transmitted to the detection server at the same time. Therefore, the process cloud will be marked as being started by a suspicious process P13 and loaded with untrusted code from the network P10. At this step, the process cloud will still not trigger a detection alarm, because the process is only started and executed suspiciously, and there is no network activity or damage to the hacked system;
[0126] 7. After that, the program took a screenshot, which was identified by the data collector through the API sequence; the process also executed the whoami sensitive cmd command to collect information. The detection results of these two malicious behaviors made the process look more like a malicious program;
[0127] 8. Finally, the program transmits the acquired private data back through the network. At this point, the process has gathered the unknown code execution, malicious behavior, and network activity required to generate an alert. Due to its various existing characteristics, including the absence of a visible window, the type of malicious behavior, and the source of the unknown code, it can be determined that this attack is a common "download and execute attack."
[0128] Attack restoration: Taking the restoration from cloud.exe as an example, the specific process is as follows:
[0129] 1. Operation: The process cloud has P1, P4, P6, P7, P8;
[0130] After this operation, the queue of pending labels S is: cloudP1, cloudP4, cloudP6, cloudP7, cloudP8;
[0131] 2. Traversed tag queue: process cloud
[0132] After this step, the attack part to be drawn: None
[0133] 3. Operation: Tag cloudP1, corresponding event is network connection E7, related tag is IPS2
[0134] After this operation, the queue of pending tags S is: cloudP4, cloudP6, cloudP7, cloudP8
[0135] Traversed tag queue: cloudP1
[0136] After this step, the attack parts that should be drawn are: IPS2, network connection E7
[0137] 4. Operation: Tag cloudP4, corresponding event is process start E8, related tag is whoamiP3
[0138] After this operation, the queue of pending tags S is: cloudP6, cloudP7, cloudP8, whoamiP3
[0139] Traversed tag queues: cloudP1, cloudP4
[0140] After this step, the attack part should be drawn: process whoami, process start E8
[0141] 5. Operation: Tag cloudP6, corresponding event is process execution E5, related tag is firefoxP2
[0142] After this operation, the queue of pending tags S is: cloudP7, cloudP8, whoamiP3, firefoxP2
[0143] Traversed tag queues: cloudP1, cloudP4, cloudP6
[0144] After this step, the attack part should be drawn: process firefox, process start E5
[0145] 6. Operation: Tag cloudP7, corresponding event is loading E6, related tags are cloud.exeF2, cloud.exeF3
[0146] After this operation, the queue of pending tags S is: cloudP8, whoamiP3, firefoxP2, cloud.exeF2, cloud.exeF3
[0147] Traversed tag queue: cloudP1, cloudP4, cloudP6, cloudP7
[0148] After this step, the attack part should be drawn: file cloud.exe, load E6
[0149] 7. Operation: Tag cloudP8, no corresponding event
[0150] After this step, the queue of pending tags S is: whoamiP3, firefoxP2, cloud.exeF2, cloud.exeF3
[0151] The tag queue has been traversed: cloudP1, cloudP4, cloudP6, cloudP7, cloudP8. After this operation, the attack part to be drawn: None
[0152] 8. Operation: Tag whoamiP3, no corresponding event
[0153] After this step, the queue of pending tags S is: firefoxP2, cloud.exeF2, cloud.exeF3
[0154] Traversed tag queues: cloudP1, cloudP4, cloudP6, cloudP7, cloudP8, whoamiP3
[0155] After this step, the attack part should be drawn:
[0156] 9. Operation: label firefoxP2, no corresponding event
[0157] After this operation, the queue of pending tags S is: cloud.exeF2, cloud.exeF3
[0158] The tag queue has been traversed: cloudP1, cloudP4, cloudP6, cloudP7, cloudP8, whoamiP3, firefoxP2. After this step, the attack part to be drawn: None
[0159] 10. Operation: Tag cloud.exeF2, corresponding event is write file E4, related tag is firefoxP1
[0160] After this operation, the queue of pending tags S: cloud.exeF3, firefoxP1
[0161] Traversed tag queue: cloudP1, cloudP4, cloudP6, cloudP7, cloudP8, whoamiP3, firefoxP2, cloud.exeF2
[0162] After this step, the attack part to be drawn: None
[0163] 11. Operation: Tag cloud.exeF3, no corresponding event
[0164] After this step, the queue of pending tags S is: firefoxP1
[0165] Traversed tag queue: cloudP1, cloudP4, cloudP6, cloudP7, cloudP8, whoamiP3, firefoxP2, cloud.exeF2, cloud.exeF3
[0166] After this step, the attack part to be drawn: None
[0167] 12. Operation: label firefoxP1, corresponding event is network connection E1, related label is IPS1
[0168] After this step, the queue of pending tags S is empty.
[0169] Traversed tag queue: cloudP1, cloudP4, cloudP6, cloudP7, cloudP8, whoamiP3, firefoxP2, cloud.exeF2, cloud.exeF3, firefoxP1
[0170] After this step, the attack parts that should be drawn are: IPS1, network connection E1
[0171] Process Figure 4 As shown in (b) and (c), the direction of the arrow in the figure is the tracking direction during the attack graph reconstruction process, not the direction of the event itself.
[0172] The above are merely embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structural transformations made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for classifying and tracing malware in a smart power plant, characterized in that: This includes using the BERT model and CNN model in deep learning technology, considering the static element information and contextual timing information of the API call sequence, and extracting the semantic features of the ultra-long API call sequence text to improve the multi-classification detection accuracy of malware.
2. The method for classifying and tracing malware in a smart power plant according to claim 1, characterized in that: The BERT model and CNN model in deep learning technology are used to comprehensively consider the static element information and contextual timing information of the API call sequence to extract the semantic features of the ultra-long API call sequence text, including: input and preprocessing layers, with malware as the input of the input layer; The API sequence de-redundancy algorithm reduces data redundancy by deleting repeated API calls or short API sequences and simplifying them into one; It includes the following steps: ① Preprocessing of API call sequence: First, preprocess the collected API call sequence to identify each API call and its parameters; arrange these API calls in chronological order to form a complete call sequence; ② Identification of repeated calls: In the pre-processed API call sequence, identify repeated API calls or short sequences; these repeated calls are usually generated when the malware performs looping operations or repetitive tasks; ③ Deletion of redundant information: For identified repeated calls, delete them to one; if an API call appears multiple times in a sequence, only the first occurrence is retained and the subsequent repeated parts are deleted; for repeated short API sequences, only one occurrence is retained and the other repeated parts are deleted; ④Sequence reconstruction: Recombine the processed API calls to form a de-redundant API call sequence.
3. The method for classifying and tracing malware in a smart power plant according to claim 2, characterized in that: The BERT model structure consists of the encoder stack of the Transformer model; the BERT layer calculation flow chart is as follows: ①Separate operation Divide each original position vector into several position vectors, and divide each 512-dimensional position vector into 8 parts, and the dimension of each part becomes 64; ②Self-attention mechanism The attention mechanism is used to analyze contextual information and perform multiple understandings of the text, increasing the weight of important information and reducing the weight of unimportant information. The attention mechanism means using multiple attention mechanisms at the same time to enhance the understanding of the text. ③Head connection and Dropout Concatenate the values output by the self-attention mechanism to obtain multi-dimensional information; ④Residual & Normalization The residual method is used to add the input and output of the attention mechanism. Then, the sum is normalized to remove the correlation of the data so that it meets the conditions of independent distribution. ⑤ Feedforward Neural Network The feedforward neural network consists of a fully connected layer and an activation function, which is responsible for transforming the dimension of the vector and removing noise.
4. The method for classifying and tracing malware in a smart power plant according to claim 3 is characterized in that: The CNN layer first converts the malware feature vector extracted by the BERT model into an image, converting the time series data classification problem with complex coupling relationships into an image classification problem, thus simplifying the problem. Then, a CNN-based classification algorithm is used to classify the malware feature vector graph to achieve real-time classification and detection of malware. The calculation process of CNN is as follows: ①Convolution and activation function Convolution uses a square convolution kernel to traverse every pixel in the image; each corresponding pixel value in the overlapping area of the image and the convolution kernel is multiplied by the weight of the corresponding point in the convolution kernel, and then summed and added with the bias to finally get a pixel value in the output image; ②Residual Using the residual method, the input and output of the neural network are added as the input of the next layer of neural network; ③Pooling The pooling operation reduces the spatial size of the data, and the number of parameters and the amount of computation decrease, allowing the model to extract a wider range of features; ④Fully connected Convert three-dimensional data vectors into one-dimensional data, which is beneficial for subsequent classification operations.
5. The method for classifying and tracing malware in a smart power plant according to claim 4, characterized in that: The data set is reasonably divided into training set, validation set and test set; the steps of data set division are as follows: Training set: The training set contains most of the samples in the dataset, accounting for 70% of the entire dataset. The training set is used to train the deep learning model, and learn the characteristics and classification rules of malware and normal samples through a large amount of sample data. Validation set: The validation set usually accounts for 20% of the dataset. During the model training process, the validation set is used to evaluate the performance of the model in real time and help adjust hyperparameters. By regularly verifying the performance of the model on the validation set during the training process, overfitting can be discovered and prevented in a timely manner; Test set: The test set usually accounts for 10% of the data set. The test set is used to perform a final evaluation of the model after the model training is completed. The data in the test set has never been seen by the model during the training process, so it can provide an objective indicator to measure the actual classification performance of the model. The evaluation results of the test set reflect the performance of the model in the real world.
6. The method for classifying and tracing malware in a smart power plant according to claim 5, characterized in that: The advanced network threat source tracing and forensics method based on dynamic behavior and context information consists of a data collection and fusion model, a malware fine-grained semantic behavior recognition model, a labeling algorithm, an offline backup database, and an attack chain reconstruction model. The data collection and fusion model is responsible for collecting malware operation dynamic logs. The malware fine-grained semantic behavior recognition algorithm is responsible for processing dynamic behavior logs. If malicious behavior is found, the labeling algorithm is triggered. The labeling algorithm extracts features from dynamic behavior logs to form a memory-based detection data structure, which is uploaded to the offline backup database. The data is then analyzed through the attack chain reconstruction model to complete the source tracing and evidence collection of advanced network threats.
7. The method for classifying and tracing malware in a smart power plant according to claim 6, characterized in that: A memory-based automaton-like data structure is used to store the status information of processes and files that may be involved in the attack; each process and file object only stores its basic information and label information, that is, status information; each process object is described as a five-tuple 〈Na, Pi, Cl, Ui, St>, where Na is the process name, Pi is the process ID, and Cl is its startup command line; Ui is the globally unique identifier of each process object; St represents the existing tag set of the process object, that is, the current state of the process object; Each file object is represented as a triple 〈Na, Ui, St>, where Na is the complete path of the file; Ui is a globally unique identifier for each file object; and St is the existing tag set of the file object, that is, the current status of the file.
8. The method for classifying and tracing malware in a smart power plant according to claim 7, characterized in that: This application uses a labeling algorithm to extract dynamic behaviors and contextual information in logs, uses a memory-based data structure to unify the data format, and analyzes it to complete the source tracing and evidence collection of advanced network threats. The method is as follows: ① The attack graph to be drawn is G, which is initially empty ②The set of tags that have been traversed is T; ③ Traverse all tags on the object O1 and put them into queue Q; ④When queue Q is not empty, take out a label L1 from queue Q ⑤ Find the event corresponding to the label L1, or the event with other labels generated because of the label; if the label is a directly marked label, discard the label; if the label L1 involves another object O2, add the object O2 and the event to G, and add the corresponding label L2 in the object O2 to the queue Q; ⑥Add the traversed tag set T; ⑦ Repeat 4-6 until there are no labels in Q; ⑧Draw G.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.