A Method for Mining SCADA Software Vulnerabilities Based on Reinforcement Learning
By generating test cases based on reinforcement learning, the problem of insufficient self-learning models in SCADA software vulnerability detection is solved, efficient and automated vulnerability detection is achieved, and the diversity and coverage of detection is improved.
Patent Information
- Application Number
- CN202211332080.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-10-28
AI Technical Summary
The lack of self-learning models of existing SCADA software vulnerability detection methods has led to limited diversity and coverage of test detection samples and limited rules formulated by experts.
Using a reinforcement learning-based method, test cases are generated through state analysis network and information evaluation network, combined with the software running state for iterative optimization, and using Python system interface and AFL proxy module for automated test cases generation and adjustment.
It realizes vulnerability detection with high coverage of self-learning, improves the diversity and coverage of test cases, and improves the efficiency and accuracy of vulnerability detection.
Smart Images

Figure CN115687114B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method for mining SCADA software vulnerabilities based on reinforcement learning. Background Art
[0002] With the vigorous development of the industrial Internet, traditional industries have begun to closely integrate with advanced Internet technologies. Compared with the relatively closed and secure environment of traditional industrial systems, the industrial Internet is exposed to an open Internet environment, which means that the risk of industrial systems being attacked has increased significantly. In recent years, there have been several security incidents in which SCADA systems have been attacked, resulting in serious economic losses and security risks. Discovering vulnerabilities in advance through vulnerability mining is an effective way to avoid security attack incidents.
[0003] Traditional means of mining SCADA software vulnerabilities include two major categories: static analysis methods and dynamic analysis methods. The difference between the two lies in whether the program needs to run. Static methods do not require the program to run and can directly discover vulnerabilities by detecting the source code. Mainstream static analysis methods include those based on abstract syntax trees and control flow graphs, such as Flawfinder, Fortify, and Coverity. Dynamic methods require the program to run for vulnerability detection, so they can detect both source code and binary code. Mainstream dynamic analysis mainly includes three methods: symbolic execution, dynamic taint analysis, and fuzz testing. Compared with static methods, dynamic methods have good scalability and low false positive rates and are more popular.
[0004] Currently, mainstream vulnerability detection methods still require experts to formulate rules and lack self-learning models, resulting in limited diversity and coverage of test detection examples.
[0005] 4 Inventive Objectives
[0006] As described above, in order to establish a self-learning vulnerability mining model and improve the diversity and coverage of test cases, we model the vulnerability detection process as a Markov decision process and propose a vulnerability mining method based on reinforcement learning. The reinforcement learning agent receives the software running state and formulates test cases according to its policy network, and the SCADA software gives incentive information according to the test cases. Our method has the following advantages:
[0007] 1. The model has self-learning ability. The reinforcement learning method improves the strategy for generating test cases through the incentives obtained by interacting with the environment, without the need for experts to supervise the formulation of test cases or test case generation rules.
[0008] 2. High coverage. The method adaptively adjusts the generated test cases according to the feedback of the test environment and can have a high coverage rate in different test environments.
[0009] 3. The algorithm has multiple adjustable parameters, so it can be adjusted and set according to specific tasks and problems, and the algorithm has good portability. Summary of the Invention
[0010] To this end, the present invention first proposes a method for mining SCADA software vulnerabilities based on reinforcement learning, which mainly uses the method of reinforcement learning to iteratively generate targeted test cases, and tests the SCADA system software in a software test environment, and dynamically adjusts the test case generation strategy according to the collected test feedback information, so as to efficiently mine software vulnerabilities.
[0011] The test case generation part based on reinforcement learning includes two modules: a state analysis network and an information evaluation network;
[0012] The software test environment part includes a Python system interface and an AFL proxy module;
[0013] The test case generation part of the reinforcement learning receives the state information during software operation through the system interface of Python. The state analysis network outputs all possible test case rules through the analysis and encoding of the state information. After receiving the test case rules generated by the state analysis network, the information evaluation network combines the current software state information, evaluates and selects each rule, and finally obtains a series of case rules that are most likely to cause the software to make mistakes, and inputs them into the software test environment part through the Python system interface. After receiving the test case rules output by the reinforcement learning network, the Python system interface generates test cases for the AFL proxy module. The AFL proxy module adjusts the test cases by statistically analyzing the relationship between code coverage and test cases, and at the same time returns the test results of the cases and the running state of the software to the reinforcement learning network through the Python system interface. The reinforcement learning network also learns and iterates through this feedback information to adjust the generated case rules.
[0014] The specific implementation method of the state analysis network is as follows: The state analysis network receives the software state S output by the test software environment, including the test structure information T of the software, which is represented by a graph, T = <G, E>, where G = {i1, i2, i3... i n} represents the structure information output by the software, i n represents an input node, E = {(i2, i3), (i m , i n ),...}, (i m , i n ) represents the input nodes i n , i mThere is a relationship. The software state analysis network outputs the above-mentioned software state S and software input structure information T, encodes them, and then predicts the information of the edges of the graph formed by the software input structure. The result is the software test rule A = {i1→i4, i1→i7, i7→i 1, ...}. Its calculation process is shown in Formula 1:
[0015]
[0016] The specific implementation method of the information evaluation network is as follows: For the rule A generated by the state analysis network, analyze the rules in combination with the current state of the software, and then calculate and output a score score. This score represents the probability that the current rule may cause the software to have an error. In order to improve the quality of test cases, select some software test rules with a higher probability for output. At the same time, the incentive information r of the software test structure will be used to update the weights of the evaluation network itself, that is, adjust the evaluation of the generated rules. The calculation method of the score score is shown in Formula 2:
[0017] score = W*(concat[A, S]) + b#(2)
[0018] where W and b are learnable parameters.
[0019] The method for generating case rules of the Python system interface is as follows: The python system call interface in the system test environment is mainly responsible for the information transfer between the reinforcement learning network and the software to be tested. It receives the test case rules output by the reinforcement learning network and then batch-generates test cases according to the given case template. The case template includes input boundaries, type conditions, etc.; the system call interface is also responsible for receiving the running state of the software as the current running state of the software, and calculating the incentive information r generated by the current test case and passing it to the reinforcement learning network for the policy update of the reinforcement learning network. First, calculate the loss function loss according to Formula (3), and then use the gradient descent method to update the parameters of the network.
[0020] loss = (Q(A, S) - r) 2 #(3)
[0021] where Q is the parameter of the reinforcement learning network.
[0022] The implementation method of the AFL proxy module is as follows: For each test case E for the input software, count the running time and the number of code lines of the software under this case, and record the running state of the software under this test case, and pass this running state to the reinforcement learning network through the python system interface.
[0023] The technical effects to be achieved by the present invention are as follows:
[0024] 1. The model has self - learning ability. The reinforcement learning method improves the strategy for case generation through the incentives obtained by interacting with the environment, without the need for experts to formulate cases or case - generation rules under supervision.
[0025] 2. High coverage rate. The method adaptively adjusts the generation of cases according to the feedback of the test environment and can have a high coverage rate in different test environments.
[0026] 3. Multiple parameters of the algorithm are adjustable, so it can be adjusted and set according to specific tasks and problems, and the algorithm has good portability. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Vulnerability mining framework based on reinforcement learning;
[0028] Figure 2 Software test rule generation;
[0029] Figure 3 Software test rule evaluation DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The following are the preferred embodiments of the present invention in combination with the accompanying drawings. The technical solutions of the present invention will be further described, but the present invention is not limited to this embodiment.
[0031] The present invention proposes a method for mining vulnerabilities in SCADA software based on reinforcement learning, which can effectively test the security and robustness of existing industrial SCADA software.
[0032] In this embodiment, for the scenario of mining security vulnerabilities in industrial SCADA software, using the method of reinforcement learning, in - depth analysis and calculation can be performed on any industrial SCADA software to be detected. Specifically, in this embodiment, a method for mining vulnerabilities in SCADA software based on reinforcement learning is used. By inputting and analyzing the real - time status information of the software, a large number of targeted, highly available, and comprehensive - coverage test cases are automatically generated in an iterative manner. These test cases are used to comprehensively test the security and robustness of the SCADA software, and at the same time, the generation strategy of the test cases is dynamically adjusted according to the test effect. The above - mentioned test cases are imported into the SCADA software for execution, and it is observed whether the execution result of the software is consistent with the expected result. If not, it means that there are vulnerabilities in the SCADA software, and the corresponding vulnerability types and trigger conditions can be given.
[0033] The method combines software structure status information, generates different test cases through the analysis of the software input structure, thereby testing the security of the SCADA software, and at the same time, the test results are fed back to the learning algorithm for the generation of the next test case, and the testing is carried out iteratively.
[0034] This method mainly includes two parts - test case generation based on reinforcement learning and software test environment.
[0035] The test case generation part based on reinforcement learning includes two modules: state analysis network and information evaluation network.
[0036] State analysis network: The state analysis network receives the output of the system interface of the software test environment, which contains the status information of the software to be tested. This module analyzes and encodes the received software status information to generate a series of test case rules.
[0037] Information evaluation network: The information evaluation network receives the test case information generated by the state analysis network, combines the current status information of the software, and selects the most effective test case rules for the current software for output.
[0038] The software test environment part includes a Python system interface and an AFL agent.
[0039] Python system interface: This module is mainly responsible for the batch generation of software test cases and the interaction of software status. It receives the test case rules of the reinforcement learning model, and then generates test cases batch by batch according to the rules; at the same time, it obtains the software status output by the AFL module and uniformly transmits it to the reinforcement learning network for training.
[0040] AFL agent: By comparing the input test cases with the code coverage rate, it can automatically adjust the combination of test cases and improve the probability of vulnerability detection.
[0041] Through this software testing method based on reinforcement learning, the security of the software can be comprehensively and efficiently tested, software vulnerabilities can be automatically detected, and the efficiency of software testers can be improved.
[0042] The technical framework adopted by the present invention is as Figure 1As shown in the figure, first, the reinforcement learning network receives some state information during software operation through the system interface of Python. The state analysis network in it analyzes and encodes the state information, and outputs all possible test case rules. After receiving the test case rules generated by the state analysis network, the information evaluation network combines the current software state information, evaluates and selects each rule, and finally obtains a series of test case rules that are most likely to cause the software to malfunction, and inputs them into the test environment through the Python system interface. After receiving the test case rules output by the reinforcement learning network, the Python system interface generates batch test cases for the AFL proxy module. The AFL proxy module automatically adjusts the test cases by statistically analyzing the relationship between code coverage and test cases, and at the same time returns the test results of the cases and the software operation state to the reinforcement learning network through the Python system interface. The reinforcement learning network also learns and iterates through this feedback information to adjust the generated test case rules.
[0043] State analysis network
[0044] The test case generation method in this invention uses the reinforcement learning method for generation. As Figure 2 shown, first, the state analysis network receives the software state S output by the test software environment, including the test structure information T of the software, which is represented by a graph, T = <G, E>, where G = {i1, i2, i3... i n} represents the structure information output by the software, and i n represents the input node, E = {(i2, i3), (i m , i n ),...}, (i m , i n ) represents that there is a relationship between the input nodes i n , i m . The software state analysis network outputs the above software state S and software input structure information T, encodes them, and then predicts the information of the edges of the graph formed by the software input structure. The result is the software test rule A = {i1→i4, i1→i7, i7→i1,...}.
[0045] Its calculation process is shown in Formula 1:
[0046]
[0047] Information evaluation network
[0048] For rule A generated by the state analysis network, the goal of the rule evaluation network is to analyze the rules in combination with the current state of the software, and then calculate and output a score. This score represents the probability that the current rule (a certain edge in the figure) may cause errors in the software. To improve the quality of test cases, we select some software testing rules with relatively high probabilities for output (greater than 0.5). On the other hand, the incentive information r of the software testing structure will be used to evaluate the update of the network's own weights, that is, to adjust the evaluation of the generated rules.
[0049] The calculation method of the score is shown in Formula 2:
[0050] score = W * (concat[A, S]) + b #(2)
[0051] Generation of batch test cases
[0052] The python system call interface in the system test environment is mainly responsible for facilitating the information transfer between the reinforcement learning network and the software to be tested. It receives the test case rules output by the reinforcement learning network, and then generates test cases in batches according to the given case template. The case template includes input boundaries, type conditions, etc. The system call interface is also responsible for receiving the running state of the software as the current running state of the software, and thereby calculating the incentive information r generated by the current test case and passing it to the reinforcement learning network for the policy update of the reinforcement learning network.
[0053] First, calculate the loss function loss according to Formula (3), and then use the gradient descent method to update the parameters of the network.
[0054] loss = (Q(A, S) - r) 2 #(3)
[0055] AFL agent
[0056] The AFL agent module is mainly responsible for statistics through test cases and code execution situations, so as to make fine adjustments to the inputs of all test cases. Specifically, for each test case E for testing the input software, statistics are made on the running time of the software and the number of code lines under this case, and at the same time, the running state of the software under this test case is recorded, and this running state is passed to the reinforcement learning network through the system interface of python.
[0057] In order to make full use of the given test cases, this module will randomly combine according to the code running situations between different test cases, so as to maximize the use of the information of the current test cases to mine software vulnerabilities.
Claims
1. A SCADA software vulnerability mining method based on reinforcement learning, characterized by: The method architecture consists of two parts: reinforcement learning-based test case generation and software testing environment. This method ultimately provides the corresponding vulnerability types and triggering conditions. The test case generation part based on reinforcement learning includes two modules: state analysis network and information evaluation network. The test case generation part based on reinforcement learning includes two modules: state analysis network and information evaluation network; The software testing environment includes a Python system interface and an AFL agent module; The test case generation part of the reinforcement learning receives the state information of the software during runtime through the Python system interface. The state analysis network outputs all possible test case rules by analyzing and encoding the state information. After receiving the test case rules generated by the state analysis network, the information evaluation network evaluates and selects each rule in combination with the current software state information, and finally obtains a series of test case rules that are most likely to cause software errors. The rules are input into the software testing environment part through the Python system interface. After receiving the test case rules output by the reinforcement learning network, the Python system interface generates test cases for the AFL agent module. The AFL agent module adjusts the test cases by counting the relationship between the code coverage and the test cases. At the same time, the test results of the test cases and the running status of the software are returned to the reinforcement learning network through the Python system interface. The reinforcement learning network also performs learning iteration based on the feedback information and adjusts the generated use case rules. The specific implementation of the state analysis network is as follows: the state analysis network receives the software state S output by the test software environment, including the test structure information T of the software, and represents it with a graph, T=<G,E> , where G={i1,i2,i3...i n } represents the output structure information of the software, i n Represents the input node, E={(i2,i3),(i m ,i n ),...},(i m ,i n ) represents the input node i n ,i m There is a relationship between them. The software state analysis network outputs the above-mentioned software state S and software input structure information T, encodes them, and then predicts the information of the edges of the graph composed of the software input structure. The result is the software testing rule A = {i1→i4,i1→i7,i7→i1,...}. The calculation process is shown in the formula: The specific implementation of the information evaluation network is as follows: for the rule A generated by the state analysis network, the rule is analyzed in combination with the current state of the software, and then a score is calculated and output. This score represents the probability that the current rule will cause the software to have an error. In order to improve the quality of the test case, some software test rules with a higher probability are selected for output. At the same time, the stimulus information r of the software test structure will be used to update the evaluation network's own weight, that is, to adjust the evaluation of the generated rules. The score is calculated as shown in the formula: score=W*(concat[A,S])+b Where W,b are learnable parameters.
2. A SCADA software vulnerability mining method based on reinforcement learning as claimed in claim 1, characterized in that: The use case rule generation method of the Python system interface is as follows: the Python system call interface in the system test environment is mainly responsible for the information transmission between the reinforcement learning network and the software to be tested. It receives the test case rules output by the reinforcement learning network, and then generates test cases in batches according to the given use case template, wherein the use case template contains the input boundary and type conditions; the system call interface is also responsible for receiving the running state of the software as the current software running state, and thereby calculating the incentive information r generated by the current test case, and passing it to the reinforcement learning network for the policy update of the reinforcement learning network. First, the loss function loss is calculated according to the following formula, and then the gradient descent method is used to update the network parameters. loss=(Q(A,S)-r) 2 where Q is the parameter of the reinforcement learning network.
3. A SCADA software vulnerability mining method based on reinforcement learning as claimed in claim 2, characterized in that: The AFL agent module is implemented as follows: for each test case E of the input software, the running time and number of code lines of the software under the test case are counted, and the running status of the software under the test case is recorded. The running status is then transmitted to the reinforcement learning network through the Python system interface.
Citation Information
Patent Citations
Program vulnerability mining method, device, terminal, and storage medium
CN109086606A
Vulnerability mining method based on machine learning and deep learning
CN110110525A