APT detection method and device in smart power plant and storage medium
By adopting a combination of program behavior model and dynamic programming algorithms in smart power plants, integrating system call API and call stack API, building feature generation modules, solving the problem of malware identification and classification, and achieving efficient APT detection and network security protection.
Patent Information
- Application Number
- CN202510106306.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-06
AI Technical Summary
Existing malware identification and classification technologies are difficult to detect encrypted shelling and complex and changeable types of malware and their variants, resulting in great difficulties and challenges in advanced intelligent analysis of cyber threats.
The program behavior model is adopted, and the low-level system call API and the high-level call stack API are integrated to build a feature generation module, and combined with the optimal local graph matching algorithm based on dynamic programming, accurately identify the behavior of malware.
Through multi-level and multi-dimensional feature description methods, the accuracy and reliability of APT detection are improved, noise is effectively filtered out, malware behavior is accurately identified, and the network security protection capabilities of power plants are improved.
Smart Images

Figure CN120105409A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of APT detection in smart power plants, and in particular to an APT detection method, device and storage medium in a smart power plant. Background Art
[0002] Smart power plants are digitalized and intelligentized by using advanced information technology on the basis of digitalization and informatization to realize the monitoring, control and management of the whole process of power production. Smart power plants combine advanced information technologies such as sensor measurement, information communication, automatic control, artificial intelligence, cloud computing, big data, and three-dimensional visualization with industrial technology and power plant management technology of power generation production process, and are highly integrated with the infrastructure of power plants. Data in smart power plants runs through the full range and process of enterprise production management. This requires smart power plants to have high security protection measures in the face of network threats, and to be able to recover quickly when attacked by malware and effectively avoid accidents. Malware is usually used as a carrier by attackers to launch APT attacks (Advanced Persistent Threat), but existing malware identification and classification technologies are difficult to detect encrypted and complex malware types and their variants, which brings great difficulties and challenges to the intelligent analysis of advanced network threats. Therefore, it is urgent to solve the problem of malware semantic behavior identification and classification, which is an important prerequisite for realizing advanced network threat detection. However, existing malware semantic behavior identification and classification technologies have problems such as coarse granularity, high false alarm rate and poor real-time performance. At the same time, existing advanced network threat tracing and forensics methods have problems such as slow response speed, low accuracy, incomplete tracing information, and incomplete attack chain reconstruction. Summary of the invention
[0003] The main technical problem solved by this application is to provide an APT detection method in a smart power plant to solve the above-mentioned problems.
[0004] To solve the above technical problems, a technical solution adopted in this application is to provide an APT detection method in a smart power plant, comprising the following steps: adopting a program behavior model, integrating low-level system call APIs and high-level call stack APIs to construct a feature generation module to accurately describe software behavior; using an optimal local graph matching algorithm based on dynamic programming to solve the noise problem, thereby accurately identifying the behavior of malware.
[0005] The present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method when executing the computer program.
[0006] The present application also provides a computer-readable storage medium having a computer program stored thereon, and the steps of the method described above are implemented when the computer program is executed by a processor.
[0007] The beneficial effects of the present application are: In the present application, the use of program behavior models can fully and deeply understand the various behavioral characteristics of the software during operation. By integrating the low-level system call API and the high-level call stack API, a feature generation module is constructed to accurately describe the behavior pattern of the software, including normal behavior and abnormal behavior. This multi-level, multi-dimensional feature description method helps to improve the accuracy and reliability of APT detection. The optimal local graph matching algorithm based on dynamic programming is used to solve the noise problem. In the APT detection process, noise often interferes with the detection results and reduces the accuracy of the detection. The optimal local graph matching algorithm can effectively filter out noise, thereby accurately identifying the behavior of malware. Improve the network security protection capabilities of power plants as well as good adaptability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a flow chart of an APT intelligent detection method based on deep learning according to an embodiment of the present application;
[0009] Figure 2 is a flow chart of a fine-grained semantic behavior identification method for malware according to an embodiment of the present application;
[0010] Figure 3 This is a schematic diagram of a case of dynamic execution log partitioning of Key Logging according to an embodiment of the present application;
[0011] Figure 4 is a schematic diagram of a semantic case of system call data according to an embodiment of the present application;
[0012] Figure 5 This is a schematic diagram of a case of matching AATR Graph and ETW log according to an embodiment of the present application;
[0013] Figure 6 is a schematic diagram of a case of converting an AATR Graph into a Transition Graph according to an embodiment of the present application;
[0014] Figure 7 is a schematic diagram of a case of obtaining an optimally matched dynamic planning transfer path according to an embodiment of the present application;
[0015] Figure 8 It is an attack cause-effect diagram restored by RATScope for a simulated RAT attack according to an embodiment of the present application;
[0016] Fig. 9is a schematic diagram of experimental results of a RAT semantic behavior recognition accuracy test based on time distribution according to an embodiment of the present application;
[0017] Fig.10 It is a schematic diagram of experimental results of a runtime load test of a log collector based on ETW according to an embodiment of the present application. DETAILED DESCRIPTION
[0018] In order to facilitate the understanding of the present application, the present application is described in more detail below in conjunction with the accompanying drawings and specific embodiments. The preferred embodiments of the present application are provided in the accompanying drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive.
[0019] It should be noted that, unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art of the present application. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more related listed items.
[0020] Figure 1 An embodiment of the APT detection method in the smart power plant of the present application is shown, including: adopting a program behavior model, integrating low-level system call APIs and high-level call stack APIs to construct a feature generation module to accurately describe software behavior; using an optimal local graph matching algorithm based on dynamic programming to solve the noise problem, thereby accurately identifying the behavior of malware.
[0021] In this application, the program behavior model is used to fully and deeply understand the various behavioral characteristics of the software during operation. By integrating the low-level system call API and the high-level call stack API, a feature generation module is constructed to accurately describe the behavior pattern of the software, including normal behavior and abnormal behavior. This multi-level and multi-dimensional feature description method helps to improve the accuracy and reliability of APT detection. The optimal local graph matching algorithm based on dynamic programming is used to solve the noise problem. In the APT detection process, noise often interferes with the detection results and reduces the accuracy of the detection. The optimal local graph matching algorithm can effectively filter out noise, thereby accurately identifying the behavior of malware. Improve the network security protection capabilities of power plants as well as good adaptability and scalability.
[0022] The APT intelligent detection method of this application prototype clustering considers multiple APIs and uses the optimal local graph matching algorithm based on dynamic programming to accurately identify the semantic behavior of malware; in order to effectively improve the computer's intelligent detection capabilities for advanced network threats, such as Figure 1 shown.
[0023] Realizing fine-grained semantic behavior recognition of malware is an important foundation for malware classification and advanced network threat tracing and forensics. Existing malware behavior recognition technologies have problems of coarse granularity, low accuracy, and poor real-time performance. Therefore, this application proposes a method for fine-grained semantic behavior recognition of malware: using a program behavior model (AATR), integrating low-level system call APIs and high-level call stack APIs to construct a feature generation module to accurately describe software behavior; using an optimal local graph matching algorithm based on dynamic programming to solve the noise problem, thereby accurately identifying the behavior of malware. The algorithm framework diagram is shown below. Figure 2 shown.
[0024] From above Figure 2 It can be seen that the fine-grained semantic behavior recognition method of malware consists of three stages: feature training stage, online log collection stage and detection stage. In the feature training stage, malware and legitimate application software are used to generate dynamic execution logs, and the dynamic execution logs are de-redundant using semantic redundancy elimination technology, and then the feature generation module is used to generate a feature database; the online log collection stage will transmit the system data collected on each host to the detection server; in the detection stage, the feature matching module will read the features from the feature database and match them with the collected system data to achieve accurate and rapid detection of fine-grained semantic behavior of malware.
[0025] Semantic redundancy elimination technology: Common malware behavior identification methods require the target program to run in a specific environment for a long time to ensure the integrity of the log data. However, logs closely related to behavior account for a small proportion of the complete log data and have a high duplication rate. For example, Figure 3 This article provides an example of dynamic execution data of Key Logging. Dynamic data can be divided into three parts: Initialization, Loop body, and Ending. The main semantic behavior of Key Logging is reflected in the repeated Loop body. If the execution time of this behavior is too long, it will aggravate the problem of semantic redundancy.
[0026] To this end, the semantic redundancy elimination technology proposed in this application starts from the top-level loop structure in the log data and recursively performs semantic redundancy removal. The premise for triggering the semantic redundancy elimination technology is if and only if the call stacks of the two system calls are exactly the same. After all the loop bodies of the top-level loop structure are identified, the same semantic redundancy elimination is performed on the embedded next-level loop structure, and finally all loop structures are traversed to obtain the de-redundant data. The semantic elimination technology can be used to convert a section of ETW log data containing long-term software running into a concise and expressive log event sequence.
[0027] The semantic redundancy elimination algorithm process is as follows:
[0028] Algorithm 1 Semantic redundancy elimination algorithm
[0029] Input: An expanded AATR data φ input ={θ j =(system call, software library call stack, application call stack)|j=1…m}
[0030] Output: AATR sequence without causality
[0031] initialization:
[0032] I:function REDUNDANCY-REDUCTION(φ input )
[0033] 2: scan from θ m toθ 1 , find the last θ lbs ∈φ input where θ lbs =θ k ,lbs+1≤k≤m
[0034] 3:φ int ←θ 1 …θ lbs-1
[0035] 4get all loop body separatorα lbs ←{j k ∣θ jk =θ lbs}
[0036] 5:llbs←max(α lbs )
[0037] 6: scan from θ 1toθ m , find the last θ be ∈φ input where θ lbe =θ k ,lbs≤k≤llbs
[0038] 7:φ end =θ lbc+1 …θ m
[0039] 8:α lbs ←α lbs ∪{lbe+1}
[0040] 9: get the maximum gap index
[0041] 10:get selected loop body
[0042] 11:get nested loop recursively reduced resultφ neserec ←Redundancy-Reduction(φ slb )
[0043] 12:Φreduced←Concatenate(φ init ,φ nestrec ,φ end )
[0044] 13:retumφ redeueed
[0045] Feature generation module design,This application integrates the system call API and the call stack API to build a feature generation module to accurately describe the semantic behavior of malware and generate the corresponding AATR model. In order to more conveniently discuss the feature generation module, the basic concepts are first formally defined as follows:
[0046] Call Stack: The call stack CSs of a system call s is composed of multiple functions. Each layer of the call stack is the caller function, and the lower layer is its callee function. A call stack is generally in the order of functions defined in the application, functions in the system software library, and low-level system calls from high to low levels.
[0047] Top-layer API: The top-layer API of a system call TAs is the first system software library function called by the application in the call stack.
[0048] Call Stack Tree: CSTreea represents the call stack tree of function a, and the root node of this tree is function a. The call stack starting from function a refers to every path starting from the root node of the tree and ending at the leaf node.
[0049] From the perspective of the system call layer, the behaviors of two different programs can be exactly the same. However, from the perspective of the call stack, they are completely different. However, most existing malware semantic behavior recognition methods only consider the system call API without considering other factors, resulting in low behavior recognition accuracy and high false positive rate. To solve this problem, this application integrates the system call API and the API in the call stack to build a feature generation module to accurately describe malware behavior.
[0050] The feature generation algorithm process is as follows:
[0051] Algorithm 2AATR Graph Generation Algorithm
[0052] Input: (1) n expanded AATR data of a specific PHF behavior: φ ph f ={φ i ∣i=1…n}, where φ i ={θ j (system call, software library call stack, application call stack)|j=1…m}; (2) a causal analysis engine: CausalityEngine, which is used to convert the AATR sequence into an AATR Graph; (3) AATR data of other PHFs: Φ ooher ; (4) AATR data of legitimate application behavior: Φ benign ;
[0053] Output: n AATR Graphs, each of which is a directed acyclic graph with AATR as a node, used to describe a specific PHF behavior: Ψ = {ψ}, where ψ = (UAATR, Ucausatity) Initialization:
[0054] 1:procedure AGGREGATED API TREE RECORD GRAPH GENERATION
[0055] 2: Preprocess each traceφ i ∈φ phf .
[0056] 3: for each traceφ i ∈Φ phf do:
[0057] 4:ψ i ←REDUNDANCY-REDUCTION(φ i )
[0058] 5:Eliminate application call-stack inψ i and organizeψ i to be an AATRbehavior sequence
[0059] 6:ψ i ←CausalityEngine(ψ i )
[0060] 7:eligible←true
[0061] 8: for each traceφ i2 ∈(Φ other ∪φ benign )do
[0062] 9:ifψ i matchedφ i2 then
[0063] 10:eligible←false
[0064] 11:if eligible then
[0065] 12:Ψ←Ψ∪{ψ i}
[0066] The character sequence represents the original ETW log event sequence, and each character represents a specific system call and its call stack. Two identical characters mean that the two system calls and their corresponding call stacks are exactly the same.
[0067] Eliminating semantic redundancy from a given ETW log event sequence can be achieved through the following three steps:
[0068] In step 1, we first identify the initialization, ending, and loop bodies in the original ETW sequence by locating the loop body delimiters.
[0069] Then, repeat the above steps in step 2 to further remove semantic redundancy from the embedded loop body (surrounded by a red dashed box in the figure), and finally obtain a log sequence without semantic redundancy (lines 11-13 of Algorithm 1).
[0070] Finally, in step 3, the application layer call stack of each log event is first eliminated (replacing characters with characters ', such as replacing A with A' to represent the eliminated log event), and the log sequence is converted into an AATR sequence (line 5 of Algorithm 2). Finally, the causal relationship judgment engine is used to replace the original AATR temporal relationship with the causal relationship of AATR (marked by a black solid line in the figure) (line 6 of Algorithm 2). Taking file download as an example, the AATR with the root node Create Fil must be called before the AATR with the root node WriteFile, so the AATR Graph generation module will create a causal relationship from the former to the latter. The AATR Graph generation module performs the above three steps for each dynamic execution data in the training set.
[0071] The design of the feature matching module is to match the trained feature library with the log data collected on the monitored host. If the match is successful, it means that the dynamic data contains a malicious behavior. However, due to the different runtime contexts of the software, the dynamic data of the same program when running at different time points has slight differences, that is, noise. The noise is irrelevant to the behavior itself, so the feature matching module must tolerate this noise when performing log matching. To this end, this application proposes a feature matching module based on the optimal local graph matching algorithm of dynamic programming to solve the noise problem, thereby realizing the precise identification of fine-grained semantic behavior of malware. The optimal local graph algorithm based on dynamic programming is divided into two parts: dynamic programming algorithm and local graph matching algorithm. The local graph matching algorithm uses a matching score to calculate the similarity between the behavior of the monitored host log and the software behavior in the feature library, and judges whether the behavior is malicious based on the similarity; the dynamic programming method is used to solve the noise problem, so as to quickly and accurately perform local graph matching, thereby realizing the fine-grained semantic behavior identification of malware.
[0072] 1) Optimal local graph matching problem
[0073] Given a and a copy of ETW log data φ={θ j ∣1≤j≤n}, the goal is to find a bijection in and At the same time, the bijection needs to satisfy the following conditions: ①f can maximize the matching density function ②For So that θ p =f(v x) and θ q =f(v y ) holds. If there exists a path from v in G x to v y If the path is p<q, it must also satisfy. Among them, It refers to the power set of V×φ, and N represents the set of natural numbers. Function t represents the matching similarity function, which is used to reflect the matching similarity of a graph.
[0074] 2) Matching similarity function
[0075] Given a and a sequence φ = {θ j ∣1≤j≤n} and a bijection in and Then the matching similarity function t:P(V×φ)→N is defined as: where b onus is a matching score function, b onus The definition is
[0076]
[0077] Among them, sysc v Represents v's system call, lib_stack v Represents the software library call stack corresponding to the system call. Δ(lib_stack v ,lib_stack θ ) is the longest common part of the two software library call stacks of v and θ from the bottom (system calls) to the top (top-level API).
[0078] The core idea of the matching function is to first calculate the similarity of each matching pair (v, θ), add up all the similarities to get a sum score, and then calculate the sum score and the optimal matching score. The higher the matching rate, the higher the matching rate between AATR Graph and ETW log event sequence.
[0079] 3) Optimal local graph matching algorithm based on dynamic programming
[0080] According to the constraint ② of the optimal local graph matching problem, when a graph node v i Match the previous event θ j When i All ancestor nodes of can only match those that appear in θ j The previous event, so v i The best match f for all ancestor nodes of jmust be certain. This discovery provides new inspiration: the optimal local graph matching problem has an optimal substructure property, so this problem can be solved by dynamic programming. For example, in the most extreme case, the graph structure can be a sequence, then the matching algorithm of this application becomes the maximum weighted common sequence problem of matching two sequences (a generalized version of the longest common subsequence problem). In general, the matching algorithm of this application solves the matching problem of a sequence and a directed acyclic graph.
[0081] In order to perform graph matching based on dynamic programming on AATR Graph, this application first defines a concept graph state (Graph State) to represent the matching sub-problem.
[0082] Graph state: Given a directed acyclic graph G = (V, E), a graph state gs is defined as a mapping gs:V→{0,1}, where gs(v) = 0 means that node v is not matched, and gs(v) = 1 means that node v is matched.
[0083] To promote a dynamic planning transition on a graph state gs, it is necessary to select the next graph node to match. This represents a graph state gs that is a direct descendant of the graph state gs. new It can only be constructed by adding a matching node to gs (Lines 5-13 of Algorithm 3). It should be noted that a node v can be added to gs only if all direct or indirect ancestors v' of a node v satisfy gs(v') = 1, so as not to violate constraint ② in the optimal local graph matching problem.
[0084] The entire matching algorithm consists of two algorithms. Algorithm 3 is responsible for the important initialization part, which converts a directed acyclic graph into a graph state transition graph by topologically sorting the AATR Graph. Algorithm 4 describes the core logic of AATR Graph matching. In short, it finds the optimal solution of the objective function through dynamic programming.
[0085] Feature matching module workflow: This application will describe a specific case to demonstrate the basic workflow of the AATR Graph matching module. Specifically, Figure 5 The output and ingress of the matching module are shown from a black-box perspective. Figure 6 and Figure 7 The example in describes in detail the flow of the matching module from input to output.
[0086]
[0087] Figure 8 Medium θ nRepresents the nth event in the ETW log event sequence; v m Represents an AATR in the AATR Graph; a capital letter surrounded by a circle represents a specific system call and its call stack. For ease of understanding, it is assumed that each AATRv m There is only one node, namely v m It also represents a specific system call and its call stack. Here, two identical letters indicate that they have exactly the same system call and call stack. That is, if θ n and v m have the same letters, then their matching score bonus (v m ,θ n ) (as defined in Definition 2) will be 1 otherwise it is 0.
[0088]
[0089]
[0090] like Figure 8 As shown in Figure 1, the matching module accepts an ETW log event sequence and an AATR Graph as input and outputs a matching similarity between them. Specifically, v 1 、v 2 and v 4 They are matched with θ in the ETW log event sequence respectively 1 ,θ 2 and θ 4 , which means bonus(v 1 ,θ 1 )、bonus(v 2 ,θ 3 ) and bonus(v 4 ,θ 4 ) are all 1. Therefore, the optimal matching score is 3, and the optimal matching rate of the AATR graph (4 nodes) is 3 / 4=0.75.
[0091] Figure 6 and Figure 7 Demonstrates the dynamic programming algorithm to obtain Figure 8 The optimal matching method in . Algorithm 3 first constructs a graph state transition graph ( Figure 6 The right half of the AATR Graph ( Figure 6 Specifically, initialize the graph state gs init is an empty set, which means that no node is matched. 1Must be matched before all nodes, the next graph state gs 1 = {v 1} is generated, indicating that v 1 In this state it has been matched.
[0092] In gs 1 Going a step further, once v 1 is matched, v 2 and v 3 are likely to be matched, so the two new graph states gs 2 = {v 1 ,v 3} and gs 3 = {v 1 ,v 2} is added to the transition graph. The above process is repeated until no new graph state is generated.
[0093] Finally, Algorithm 4 finally finds the state transition diagram based on the dynamic programming method ( Figure 6 ) is the optimal path in . Figure 7 Shows one of the optimal paths gs init →gs 1 →gs 3 →gs 4 →gsf in For each transfer step, the matching score will be updated and changed accordingly on the graph. Specifically, gs init to gs 1 As an example, the algorithm finally decides to transfer v 1 With θ 1 Matching, rather than selecting other events, is because this local matching decision can help the global matching get the highest score. Similarly, the algorithm then 1 to gs 3 In the transfer, v 2 and θ 3 Match, the match score is now 2. In gs 3 to gs 4 In the transfer, v 3 The match with all ETW log events is 0, so the algorithm chooses to skip v 3 , the score is still 2. Finally, the algorithm will v 4 and θ 4 Match, the final score is 3. Through the above steps, the algorithm can finally obtain the optimal matching score and Figure 7 The optimal matching rate is shown.
[0094] Experimental part: This application will evaluate the performance and effectiveness of RATScope from the following three aspects:
[0095] Q1: The effectiveness of RATScope in dealing with real-world RAT attacks.
[0096] Q2: Robustness of RATScope when dealing with unknown RAT families.
[0097] Q3: Runtime and storage load introduced by RATScope.
[0098] When deploying RATScope in a smart power plant, each host needs to enable its own ETW function and use the collector proposed in this article to collect logs; in addition, a dedicated server is required to receive the log data collected by the host and perform AATR Graph matching. In the experiment of this application, the host is configured with an i5-4590 processor and 8GB of memory, and the server is configured with an Intel Xeon E5-2650 processor and 252GB of memory.
[0099] 1) Effectiveness of RATScope in Simulated Attacks
[0100] The experiment in this application simulates a real-world RAT attack. The attack details and strategies (including attack surface, RAT family used, anti-detection techniques, etc.) are recorded in detail on blogs, so this paper can accurately simulate the RAT attack based on these records. It is worth noting that these blogs cannot provide fine-grained semantics at the PHF level, so this part is simulated by this paper based on the attacker's purpose.
[0101] Real-world RAT attack: The purpose of this RAT attack is to steal the intelligence of the activist (such as account passwords, chat texts and voice records, etc.). Specifically, the attacker uses complex social engineering techniques (such as spear phishing attacks to steal Skype accounts and send messages containing malicious code to the target) and visual deception techniques (such as extension deception techniques) to induce victims to execute malicious code (here refers to RATs, such as DarkComet, NjRAT and Xtreme), and uses anti-detection techniques (such as obfuscation techniques and process injection techniques) to protect malicious processes. After the RAT is successfully executed, the attacker remotely triggers PHF, such as keylogging, microphone monitoring and camera monitoring, to obtain intelligence from the victim.
[0102] Simulated RAT Attacks in This Paper: The simulated attacks in this experiment used the same attack strategy as the real RAT attacks and attempted to obtain the same victim information. Table 1 provides the details of the simulated attacks, including the victim, the attack surface used, the RAT family used, the anti-detection techniques, and the PHF performed. The table also shows the attack duration and the accuracy and false positive rate of RATScope for the attack.
[0103] Specifically, Alice and Bob are two employees of a company. Bob has information that the attacker is interested in, including emails and chat records. However, Bob has a strong sense of security, and the attacker cannot directly attack Bob's machine. So the attacker uses two steps to indirectly attack Bob's host. In the first step, the attacker first attacks Alice's host and uses it as a springboard. The attacker first sends a phishing email to Alice's computer. The email contains a malicious executable attachment urgentgpj.scr, which uses the file extension spoofing technology Right to Left Override to disguise itself as a JPG image file urgentrcs.jpg. Alice lacks security awareness and directly double-clicks the executable attachment. After the malicious attachment is executed, a JPG file is first opened in the foreground to confuse Alice, and at the same time, an obfuscated DarkComet RAT is started in the background. The RAT injects itself into the Internet Explorer process to bypass the firewall and connects back to the attacker's server xxxx:80. After the connection is established, the attacker can remotely trigger the RAT's keylogging function and finally obtain Alice's Skype account and password. In the second step, the attacker also used the file extension spoofing technique to generate a fake PDF file proposalrcs.pdf and sent it to Bob via Skype. Since Alice and Bob often transfer files via Skype, Bob did not realize the danger and double-clicked to open the file. This led to the successful execution of NjRAT on Bob's machine and connected back to the server XXXX:443. The attacker then triggered the RAT's camera monitoring, microphone monitoring, and keylogging functions to obtain intelligence on Bob. During this attack, Alice and Bob used office software normally for their daily work.
[0104] Table 1: Experimental details of simulated RAT attacks
[0105]
[0106]
[0107] Using RATScope for attack auditing: This experiment installed an ETW-based log collection module on Alice and Bob's machines. Now assume that a third-party threat intelligence system finds XXXX in the log as a malicious IP and generates a warning. Taking this warning as a starting point, the existing audit system can only find processes and files related to the attacker. It is difficult for attack auditors to fully understand the purpose and behavior of this RAT attacker based only on these process and file information. In contrast, RATScope successfully captured the PHF executed by the RAT. Figure 8 A simplified causal graph generated by RATScope is provided. As can be seen from the figure, in addition to files and processes, RATScope also provides fine-grained semantic information at the PHF level. This semantic information can help auditors better understand the attacker's strategy and intentions and make corresponding emergency responses. For example, based on the execution time of the microphone monitoring behavior, auditors can determine which important information was stolen by the attacker.
[0108] 2) Behavior recognition accuracy experiment across RAT families
[0109] This experiment will evaluate the PHF recognition ability of RATScope for unknown RAT families (not included in the training set).
[0110] Experimental setup:
[0111] Table 2: Programming languages used by 53 RAT families and their distribution
[0112]
[0113]
[0114] Table 3: Common Potentially Harmful Functions (PHFs) Equipped with RATs
[0115] Potentially malicious functionality (PHF) describe Frequency KeyLogging Logs the victim's keystrokes 81.13% RemoteShell Open a shell remotely and allow the attacker to execute 81.13% DownloadandExecute Open a shell remotely and allow the attacker to execute 84.90% RemoteCamera Remotely access the victim's webcam 66.03% AudioRecord Remotely record the victim's microphone 49.05%
[0116] Family-split PHF dataset: This experiment first randomly selects a RAT sample for each of the 53 RAT families in Table 2. Then, for the 53 selected samples, 5 commonly used PHFs (as shown in Table 3) were executed respectively, including keyboard logging, remote shell, download and automatic execution, camera monitoring and microphone monitoring, and the log collection module was used to collect the dynamic execution data of each PHF execution. Then the dynamic execution data of the same PHF are divided into the same group, and each group is labeled with the corresponding PHF. Through the above steps, 5 Family-split PHF sub-datasets were finally collected. It is worth noting that in the same PHF sub-dataset, any two dynamic execution data cannot be collected from the same RAT family.
[0117] Benign Dataset: This experiment selected 90 legitimate applications that are widely present in smart power plants. According to the category, the selected programs include editing applications (such as Notepad++, GIMP and Word, etc.), communication software (such as Skype, Foxmail and Outlook, etc.), browsers (such as Chrome and IE), file transfer tools (such as WinSCP, FileZilla and FreeFileSync, etc.) and a large number of long-running system processes (such as explorer and dllhost, etc.). In addition, this experiment also collected legitimate applications with similar functions to PHF, such as audio-related applications (QuicktimePlayer and Audiorecorder), Shell-related programs (CMD) and download-related programs (FreeDownloadManager). This experiment installed these programs on common user machines and manipulated them according to normal user habits (such as browsing the web with a browser), and then used the collected log data as the Benign dataset.
[0118] Experimental method: This experiment uses the fold cross-validation method to evaluate the capabilities of RATScope on the Family-split PHF dataset and the Benign dataset. Specifically, for each tested PHF, this experiment first divides the Benign dataset and the Family-split PHF sub-dataset corresponding to the PHF into 10 parts respectively, and randomly selects 1 part as the test set and the remaining 9 parts as the training set. Repeat the above steps 10 times, and take the average of the results as the final result. It is worth noting that in this experiment, the dynamic execution data of a RAT family cannot appear in both the training set and the test set, which avoids the bias problem introduced by the random cross-validation method, that is, using the data of a RAT family for training and using the data of the family as the test set. This bias will significantly improve the evaluation results.
[0119] Baseline Methods: API and system calls are the two most common types of log data used for program behavior modeling. The AATR model proposed in this paper combines API and system calls to more accurately identify fine-grained program semantic behaviors and solves the semantic conflict problem mentioned earlier due to the lack of parameters in ETW. In order to make a more comprehensive comparison, this experiment replaced the AATR nodes in the existing AATR Graph with the corresponding APIs and system calls, and generated two baseline models: API-only model and syscall-only model to approximate the previous work. Specifically, in the AATR model proposed in this paper, each node represents a system library API or system call, the root node represents a top-level API, and the leaf node represents a system call. By replacing the AATR in the generated AATR Graph with the root node top-level API or the leaf node system call, two baseline models were generated respectively. It is worth noting that, except for the lack of parameters, these models are similar to the models proposed in existing work.
[0120] Analysis of experimental results: Table 4 shows the experimental results in detail. In this article, True Positive (TP) represents the number of RAT families whose PHF behaviors are correctly identified by RATScope, and False Positive (FP) represents the number of legitimate applications whose PHF behaviors are incorrectly identified by RATScope. RATScope can accurately identify RAT behaviors. As can be seen from Table 4, for each HF, the AATR model can identify most of the RAT families in the test set (with an accuracy of approximately 85% to 93%), which verifies the key finding in the RAT survey, that is, different RAT families use similar implementations to implement the same PHF. For example, the remote shell AATR model generated by SpyNet RAT can match the same PHF behavior of 28 other RAT families.
[0121] Compared with the API-only and syscall-only models, the AATR model has a lower false alarm rate. As shown in Table 4, the AATR model has a lower false alarm rate (0% to 2.2%) than the other two models. The reason is that the AATR model can solve the semantic conflict problem and combine the advantages of API and system calls at the same time, so it can achieve higher accuracy. The other two models cannot handle the semantic conflict problem due to the lack of parameters, resulting in a decrease in accuracy. In addition, this experiment found that compared with system calls, the top-level API directly called by the program contains more semantics, which also explains the phenomenon that the false alarm rate of the APIonly model is lower than that of the syscall-only model. This experiment also found that when identifying the PHF behavior of camera monitoring, the API-only model can obtain the same accuracy as the AATR model because the API without parameters already contains enough semantic information to identify this PHF behavior.
[0122] Table 4 Comparison of behavior recognition accuracy of different models
[0123]
[0124]
[0125] The AATR model has a small number of false positives. The semantics of some legitimate program behaviors are similar to PHF behaviors, such as the audio recording tool AudioRecorder that comes with Windows. Experiments have found that the AATR model will mistakenly identify the "PHF-like behavior" of legitimate programs as PHF behaviors. Further analysis of these false positive data shows that it is difficult to distinguish legitimate application developers and RATs with "PHF-like behaviors" from the perspective of low-level logs, because their implementation methods are similar, and the implementation methods themselves are not malicious or legitimate. In addition, some malware will abuse the functions of legitimate programs to implement malicious behaviors, making these behaviors even more difficult to distinguish. This is why this article calls RAT behaviors potential harmful functions. However, it is still valuable to report these false positives at the first point in time. In the subsequent audit process, security personnel can use other contextual information to comprehensively determine whether the warning is a real malicious attack. In fact, both in industry and academia, security personnel have realized that relying on a single behavior to diagnose malicious attacks is insufficient. NoDoze uses the contextual information of the warning (such as the preceding log events that caused the warning and the subsequent damage caused by the warning) to automatically diagnose whether the warning is a malicious attack. Likewise, RATScope will be able to leverage contextual information to comprehensively diagnose these PHF warnings in the future. For example, Figure 8It is difficult to judge whether this is a malicious attack just by accessing the camera from the IE browser. However, by checking the context information, it is found that the IE browser is created by a suspicious process proposalrcs.pdf, and the executable file of this process is downloaded from Skype. Combining these context information, it can be judged that this is not a normal browser behavior.
[0126] 3) Experiment on the accuracy of behavior recognition based on RAT time distribution
[0127] This experiment takes into account that RATs will be updated over time, so we choose to use RAT samples that appeared earlier for training and use RAT samples that appeared later for testing.
[0128] Experimental setup: Temporally-sorted PHF dataset: First, five Family-split PHF sub-datasets were generated using the method of generating the Family-split PHF dataset. Then, the data in each sub-dataset were arranged from early to late according to the corresponding RAT appearance year. Finally, five Temporally-sorted PHFs were obtained.
[0129] Benign dataset: The same legitimate applications were used in this experiment. Then, the applications were randomly divided into two groups. For the applications in the first group, the release versions released between 2012 and 2013 were selected. For the applications in the second group, the release versions released between 2015 and 2016 were selected. Then, using the same data generation method as the cross-RAT family experiment, two Benign datasets were obtained: the OLD Benign dataset containing dynamic execution data of applications from 2012 to 2013 and the NEW Benign dataset containing dynamic execution data of applications from 2015 to 2016.
[0130] Experimental method: This experiment follows the best practices proposed by the TESSERACT work, dividing the training set and test set according to the time of RAT appearance to avoid introducing timing bias into the experiment. That is, this experiment will use RATs known in the past for training and RATs unknown in the future for testing. Specifically, for each PHF tested, this experiment uses RAT families that appeared between 2015 and 2016 and the NEW Benign dataset as fixed test sets. Next, observe the change in accuracy by continuously changing the training set (continuously adding RAT families from 1999 to 2014). The OLD Benign dataset is used in the training set. Fig. 9The test results are shown. A RAT family R on the X-axis indicates that the current training set contains all RAT families that appeared before R. The test set is fixed and contains RAT families that appeared from 2014 to 2015.
[0131] Experimental results analysis: Fig. 9 This shows that RATScope is capable of using past known RAT samples as training sets to detect new RAT samples in the future. For most PHFs, using 8 RAT family samples can detect 50% of the RAT families that appeared between 2015 and 2016, with a false positive rate of less than 3%. When the training set includes RAT families between 2012 and 2014, the accuracy can rise to about 80%. To further understand these experimental results, this experiment manually analyzed these RATs and found that: (a) RAT developers tend to reuse the code base of existing RAT families, which are often cracked or published on hacker forums. For example, NjRAT is a notorious RAT family whose source code was leaked and made public in 2013. In addition, the code architecture of Kiler RAT and NjRAT that appeared in 2015 is highly similar, and the configuration file format of Coringa RAT that appeared in 2015 is also highly similar to NjRAT. Other RATs including CyberGate and NanoCore were also found to share code bases. (b) RAT developers use public software libraries to implement PHF behaviors. For example, this experiment found that CtOS and Imminent Monitor used a well-known third-party software library AForge to implement remote camera functions, Crimson and jSpy used the software library JNativeHook to implement keylogging functions, and Quasar used the software library GlobalMouseKeyHook to implement keylogging functions. This finding is reasonable because it is not easy to develop a stable and complete PHF function from scratch.
[0132] Performance test experiment:
[0133] Runtime Load: In order to evaluate the runtime load of the ETW-based log collection module, this experiment was tested under different workloads. Specifically, this paper conducted two experiments to evaluate the performance of the log collection module.
[0134] First, this experiment builds a test program that can generate different types of ETW log events and control the number of logs generated per second. Then the program is used to test the log collection module. Fig.10It can be seen that when the number of logs generated per second reaches 20,000 / second, the load is about 7%. When the generation speed reaches the maximum value of 229,475 / second that the current machine can withstand, the load is 83.72%. This shows that the runtime load is closely related to the event generation speed.
[0135] Then, this experiment tested the load generated by the log collection module on a single application under high application load and real user environment. Specifically, this experiment selected 12 common applications, including browsers, communication software, editing software, multimedia software and development tools. Then a GUI automation tool was used to continuously trigger the functions of these applications, such as browsing web pages, sending and receiving emails, editing documents, playing music and compressing documents. At the same time, the dynamic execution data of these applications was collected and the log event generation speed was calculated. Fig.10 The experimental results shown in show that the average event generation rate of most applications is less than 20,000 events per second, that is, the runtime load is less than 7%. It is worth noting that the above GUI automation tools will operate the application continuously, while real users will stay for a few seconds to read the content on the GUI interface after performing an operation. Therefore, in order to simulate the real user environment, this experiment sets the GUI automation tool to pause for 5 seconds each time it clicks the mouse. The simulation process lasts for 30 minutes, and the final average runtime load is 3.7%.
[0136] ETW parsing performance: To evaluate the effectiveness of the parsing technology proposed in this paper, this experiment conducted a comparative test on the parser proposed in this paper and the parser that comes with the system. Specifically, this experiment used different parsers on the same machine to parse the same segment of raw ETW data. The experimental results are shown in Figure 5. RATScope's parsing speed is much faster than other parsers, nearly 6 times the speed of TDH and nearly 2 times the speed of TraceEvent. Faster parsing speed means more savings in system resources and faster response to attacks.
[0137] Table 5 Comparison of parsing performance tests of different parsers
[0138]
[0139] Storage load: In this experiment, the log collection module was deployed on two machines and collected nearly 1 day of data. Table 6 shows the comparison results before and after using the data reduction technology proposed in this paper, and the results are compressed sizes. The results show that the data reduction technology proposed in this paper can delete 97.61% of the data, which confirms the previous observation that there is a lot of semantic redundancy in the log data. In short, the log collection module proposed in this paper collects an average of 0.2GB of data per day. Considering a smart power plant with 30 machines, the collection module in this paper will generate 2TB of data per year. The current market price of a 2TB hard drive is about US$60, so this storage consumption is affordable for enterprises.
[0140] Table 6 Experimental results of storage load test based on ETW log collector
[0141] Alice Bob Average BeforeReduction 7.5G 9.4G 8.4G AfterReduction 0.18G 0.23G 0.2G Reduction Ratio 2.4% 2.44% 2.38%
[0142] The above are merely embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structural transformations made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for detecting APT in a smart power plant, characterized in that: This includes adopting a program behavior model, integrating low-level system call APIs and high-level call stack APIs to build a feature generation module to accurately describe software behavior; and using an optimal local graph matching algorithm based on dynamic programming to solve the noise problem, thereby accurately identifying the behavior of malware.
2. The APT detection method in a smart power plant according to claim 1, characterized in that: The fine-grained semantic behavior recognition method of malware consists of three stages: feature training stage, online log collection stage and detection stage; In the feature training phase, malware and legitimate application software are used to generate dynamic execution logs, semantic redundancy elimination technology is used to remove redundancy from the dynamic execution logs, and then the feature generation module is used to generate a feature database; During the online log collection phase, the system data collected on each host will be transmitted to the detection server; during the detection phase, the feature matching module will read features from the feature database and match them with the collected system data to achieve accurate and rapid detection of fine-grained semantic behaviors of malware.
3. The APT detection method in a smart power plant according to claim 2, characterized in that: The semantic redundancy elimination technology starts from the top-level loop structure in the log data and recursively removes semantic redundancy. The semantic redundancy elimination technology is triggered only when the call stacks of two system calls are exactly the same. After all loop bodies of the top-level loop structure are identified, the same semantic redundancy elimination is performed on the embedded next-level loop structure, and finally all loop structures are traversed to obtain the de-redundant data.
4. The APT detection method in a smart power plant according to claim 3, characterized in that: Eliminating semantic redundancy from a given ETW log event sequence can be achieved through the following three steps: In step 1, we first identify the initialization, ending, and loop body in the original ETW sequence by locating the loop body delimiter; In step 2, repeat the above steps to further remove semantic redundancy from the embedded loop body, and finally obtain a log sequence without semantic redundancy. In step 3, the application layer call stack of each log event is first eliminated, and the log sequence is converted into an AATR sequence; finally, the causal relationship judgment engine is used to replace the original AATR time series relationship with the AATR causal relationship.
5. The APT detection method in a smart power plant according to claim 4, characterized in that: Given a and a copy of ETW log data φ={θ j ∣1≤j≤n}, the goal is to find a bijection in and At the same time, the bijection needs to satisfy the following conditions: ①f can maximize the matching density function ②For So that θ p =f(v x ) and θ q =f(v y ) holds; if there exists a path from v in G x to v y If the path is p<q, it must also satisfy; It refers to the power set of V×φ, N represents the set of natural numbers; the function t represents the matching similarity function, which is used to reflect the matching similarity of a graph.
6. The APT detection method in a smart power plant according to claim 5, characterized in that: Given a and a sequence φ={θ j ∣1≤j≤n} and a bijection in and Then the matching similarity function t:P(V×φ)→N is defined as: Where bonus is a matching score function, and bonus is defined as Among them, sysc v Represents v's system call, lib_stack v Represents the software library call stack corresponding to the system call; Δ(lib_stack v ,lib_stack θ ) is the longest common part of the two software library call stacks of v and θ from bottom to top.
7. The APT detection method in a smart power plant according to claim 6, characterized in that: According to the constraint ② of the optimal local graph matching problem, when a graph node v i Match the previous event θ j When i All ancestor nodes of can only match those that appear in θ j Previous events.
8. The APT detection method in a smart power plant according to claim 7, characterized in that: First, we define a conceptual graph state to represent the matching subproblem; graph state: given a directed acyclic graph G = (V, E), a graph state gs is defined as a mapping gs:V→{0,1}, where gs(v) = 0 represents that node v is not matched, and gs(v) = 1 represents that node v is matched.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.