Intelligent security auditing method, system and equipment based on hierarchical reinforcement learning

By adopting a hierarchical reinforcement learning method in security audit, the problems of inefficiency of traditional security audits in complex network environments and instability of policy convergence in complex network environments are solved, and more efficient and accurate security threat identification is achieved.

CN120185872AActive Publication Date: 2025-06-20JINAN TIMES CONVINCED INFORMATION SECURITY EVALUATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510281211.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-20
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Traditional security audit methods have problems such as low efficiency, narrow coverage and unstable policy convergence when dealing with complex network environments and diverse attack scenarios, which are difficult to meet the actual network security audit needs.

Method used

The intelligent security audit method based on hierarchical reinforcement learning is adopted to model the security audit process as a hierarchical Markov decision-making process. Through the hierarchical design of macro strategies and micro strategies, global planning and risk control and single-step attack actions are achieved.

Benefits of technology

It significantly improves the detection ability of complex attack scenarios, reduces manual intervention, improves the efficiency and accuracy of security threat identification, and adapts to the needs of multi-step and multi-stage complex network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185872A_ABST
    Figure CN120185872A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security auditing and automatic penetration testing, in particular to an intelligent security auditing method, system and equipment based on hierarchical reinforcement learning, and the method comprises the following steps: modeling a security auditing process into a hierarchical Markov decision process, generating an attack sequence by a server, and sending the attack sequence to a client, the attack sequence is generated based on hierarchical reinforcement learning, then the client side executes an attack action according to the received attack sequence, collects feedback and sends the feedback to the server side, the server side updates the hierarchical reinforcement learning according to the feedback and carries out optimization and iteration, and when a set iteration condition is met, the security audit process is ended. Otherwise, continuing to carry out the safety audit operation, and ending the safety audit process until an iteration stop condition is met. According to the method, the detection capability of coping with a complex attack scene is improved through automatic hierarchical learning and an adaptation mechanism, and manual intervention is reduced, so that more efficient and accurate security threat identification is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network security auditing and automated penetration testing, and particularly to an intelligent security auditing method, system and device based on hierarchical reinforcement learning. Background Art

[0002] With the continuous improvement of the complexity of the network environment and application systems, traditional security auditing methods have deficiencies in terms of response efficiency, coverage, and diverse attack scenarios. Manual auditing is not only time-consuming and relies on expert experience, making it difficult to adapt to the evolving threat landscape. Therefore, in recent years, reinforcement learning methods have been studied and applied in the fields of automated security auditing and penetration testing, with the expectation of enhancing the detection ability for complex attack means through automated learning and adaptation mechanisms and reducing the degree of manual participation. However, traditional reinforcement learning often faces problems such as an overly large search space, low training efficiency, and unstable policy convergence, making it difficult to fully meet the actual network security auditing requirements.

[0003] Therefore, the present invention proposes an intelligent security auditing method, system and device based on hierarchical reinforcement learning to solve the above problems. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention develops an intelligent security auditing method, system and device based on hierarchical reinforcement learning. The present invention improves the detection ability for complex attack scenarios through an automated hierarchical learning and adaptation mechanism, reduces manual intervention, and thus obtains more efficient and accurate security threat identification.

[0005] The technical solution for the present invention to solve the technical problem is an intelligent security auditing method based on hierarchical reinforcement learning, which is specifically as follows: Model the security auditing process as a hierarchical Markov decision process, and conduct security auditing operations by the server generating an attack sequence and sending it to the client. The attack sequence is generated based on hierarchical reinforcement learning, which includes a macro policy and a micro policy. Then the client executes attack actions according to the received attack sequence and collects feedback, and then sends the feedback to the server. The server updates the macro policy and the micro policy according to the feedback, iterates and optimizes the macro policy and the micro policy, sets an iteration stop condition. If the iteration condition is met, the security auditing process ends; otherwise, continue to conduct security auditing operations according to the updated macro policy and micro policy until the iteration stop condition is met and the security auditing process ends.

[0006] In the specific implementation manner, the hierarchical reinforcement learning includes a macro policy and a micro policy , and the macro policy is the high level in the hierarchical reinforcement learning, and the micro policy It is the lower layer in hierarchical reinforcement learning and defines subtasks by the upper layer. and the upper-layer goal , , defines executable micro-actions by the lower layer , , represents an attack action; The macro policy selects a subtask under the state of the system , denoted as , the micro policy selects a micro-action under the subtask and the state of the system , denoted as , the state , denoted as , the state , represents the network environment, system configuration information, and current attack progress.

[0007] In the specific implementation manner, the process of the macro policy selecting a subtask is as follows: The macro policy analyzes the current system state through threat intelligence information and the system's historical audit records , dynamically determines the upper-layer goal to be executed , and then based on the macro policy under the state of the system selects a suitable subtask from the upper-layer goal ; ; The threat intelligence information includes data sources from network detection systems, intelligence sharing platforms, and security vulnerability databases; The upper-layer goal consists of a series of subtasks , , represents the th subtask in the upper-layer goal represents the number of subtasks in the upper-layer goal , ; The macro policy calculates the macro value generated by the immediate reward through the macro value function, and arranges the execution order and timing of subtasks. The specific process is as follows: The calculation formula of the macro value is as follows: , where represents the discount factor, represents the expected value, represents the preset number of planning time steps, represents the state of the system The macro value of the subtask , represents the state of the system at time when the subtask is executed and the immediate reward generated thereafter; Then, based on the macro value, priority determination operations are performed to arrange the execution order and timing of the subtasks. Specifically, by setting a priority threshold , subtasks that meet are filtered out as candidate subtasks, and then the candidate subtasks are sorted in descending order according to the value to obtain the priority sequence of the subtasks. Among them, the higher the value, the higher the execution priority; Finally, based on the priority sequence of the subtasks and the current resource status of the system, on the premise of ensuring that the state of the system allows the subtasks to be executed, each candidate subtask is started in turn, and the is updated in real time according to the task feedback during the execution process.

[0008] In the specific implementation manner, the micro strategy selects the micro action as follows: Based on the micro strategy and the state of the system in the subtask select the micro action , , and the attack action includes instruction execution, port and traffic detection, vulnerability exploitation actions, privilege escalation actions, and resource usage and concealment-related actions; The micro strategy evaluates the execution value of different micro actions in each subtask by defining the micro value , and continuously updates the micro strategy in combination with the immediate feedback. The specific process is as follows: In the system initialization stage, for each subtask and all possible micro actions , the micro value is assigned a random initial value; In the state of the system , when the subtask is executed, the micro strategy selects a micro action for execution according to the value of the current micro value , and then adopts the -greedy strategy to balance exploration and exploitation, according to the probability Select the action that maximizes the value. Randomly select other actions according to the probability . After the micro-action is executed, the state of the system changes from to , and an immediate feedback reward is obtained. After multiple iterations and continuous updates of the immediate feedback reward, the micro-policy tends to be optimal for the state of the system and the sub-task selected according to the micro-value .

[0009] In the specific implementation manner, the execution and feedback of the client: (1) At time, the client receives the micro-action sent by the server, and based on the system state at time and the corresponding sub-task , the client performs specific attack operations. During the execution of the specific operations, the client executes attack actions on the audit object according to the instructions of the micro-action ; (2) After the client finishes the execution, the state of the system changes from to , indicating the latest state after executing the action . The state includes network topology and port changes, system configuration and resource usage, current attack progress and sub-task completion indication information, and whether security exceptions or defense mechanism trigger exception events occur; (3) The client calculates the immediate feedback reward of the micro-policy according to the update of the system state. The calculation formula is as follows: wherein, it sequentially includes five reward items, namely , , , and . represents the success degree of quantifying an attack operation, represents the time overhead of the attack operation, represents the security risk brought by the attack operation, represents the system resources consumed by the attack operation, represents the state-action exploration reward, , , , and respectively represent in the state the next reward item , , , and dynamic weight coefficients; State-action exploration reward The calculation formula is as follows: , where, represents the exploration reward base coefficient, represents the total number of attack actions executed in history, represents the index of represents the th attack action executed in history, represents the current attack action and the state-action space distance between represents the distance sensitivity adjustment factor; is determined by the multi-level security priority evaluation model and the security state entropy weight method , , represents the index of the number of reward items, , The calculation formula is as follows: , where, also represents the index of the number of reward items, represents the th preset base weight of the reward item, represents the state th dynamic adjustment factor of the th reward item of the state represents the state entropy of the state represents the th preset base weight of the reward item, represents the state th dynamic adjustment factor of the State The calculation formula of the state entropy is as follows: , where, represents the number of system state classifications, represents the state belongs to the th class of states; State Dynamic adjustment factor The calculation formula is: , where, represents the important index of the th reward item in the state, represents the normalization constant, and represent two different adjustment coefficients, represents the adjustment coefficient of the th reward item , represents the adjustment coefficient of the th reward item , represents the current attack progress index, represents the th historical success rate of the reward item; The calculation formula of the historical success rate is as follows: , where, represents the total number of attack actions executed up to the moment, represents the index of, represents the set of relevant actions of the th reward item, represents the indicator function. If , then , otherwise , represents judging whether the th attack action is successful. If successful, , if failed, ; (4) The client sends the latest state , immediate reward , sub-task completion indication information, and auxiliary diagnosis information back to the server in the form of a data packet.

[0010] In the specific implementation manner, the iteration and optimization of the macro strategy and the micro strategy: (1) Update the micro strategy: The server performs strategy iteration according to the data packet fed back by the client, and updates the micro value using a value-based method. The calculation process is as follows: , Among them, represents the learning rate used to control the fusion ratio of new and old values, represents the discount factor, represents the candidate action that obtains the maximum value among all microscopic actions, represents the subtask selected at the moment; represents the value brought by the attack action of the optimal microscopic policy; (2) Update the macro policy: When the subtask is completed or switched in the service segment, policy iteration is performed according to the revenue and risk information of the macro policy. The immediate reward of the macro policy , and the calculation formula is as follows: , Among them, represents the degree of coverage improvement for unknown network nodes, ports, and potential vulnerability areas, represents the success degree of quantifying an attack operation, represents the measurement of the risk of the attack action at the macro level, represents the cost of executing the attack action; When the subtask is completed or switched, the immediate reward of the macro policy is used to update the macro value , and the calculation formula is as follows: , Among them, represents the macro learning rate, represents the discount factor, represents the subsequent value obtained by executing the subtask of the optimal macro policy layer, represents the update operation; (3) Set the iteration stop condition: Set the change threshold of the value function. When the in the microscopic policy and the update amplitude in the macro policy are both lower than the given threshold, it is judged that the microscopic policy layer is stable, specifically as follows: Define the update amplitude thresholds for the microscopic policy and the macro policy respectively. For the microscopic policy layer, set the microscopic threshold , and for the macro policy, set the macro threshold ; In each iteration process, for all system states , subtasks and microscopic actions , calculate the maximum absolute change before and after the update of the microscopic value, and the calculation formula is as follows: , wherein, represents calculating the maximum value based on and represents the updated microscopic value, represents the microscopic value before update; Meanwhile, for all system states , calculate the maximum absolute change before and after the update of the macroscopic value , and the calculation formula is as follows: , wherein, represents calculating the maximum value based on and represents the updated macroscopic value, represents the macroscopic value before update; If in consecutive iterations, both and are satisfied, it is considered that the policy update has tended to be stable and the iteration can be terminated; When the set main security audit objective has been completed, or when the available resources are not enough to continue the large-scale audit, the training process can also be aborted; If the audit policy performance still cannot be significantly improved after exceeding the maximum number of iteration rounds, one can choose to exit or switch to the manual review mode according to specific requirements.

[0011] The present invention also provides an intelligent security audit method based on hierarchical reinforcement learning, including a module for executing each step processing instruction in the intelligent security audit method based on hierarchical reinforcement learning.

[0012] The present invention also provides a device for executing an intelligent security audit method based on hierarchical reinforcement learning, and the device includes: a memory, a processor, and a program of the intelligent security audit method based on hierarchical reinforcement learning stored on the memory and executable on the processor.

[0013] The present invention also provides a memory, on which a program of the intelligent security audit method is stored, and when the program of the intelligent security audit method is executed by a processor, an intelligent security audit method based on hierarchical reinforcement learning is implemented.

[0014] The effects provided in the summary of the invention are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: The security audit task is split at different levels through a hierarchical reinforcement learning framework, thereby narrowing the search space and improving learning efficiency. Among them, the macro policy is responsible for global planning and risk control, and the micro policy further refines the execution and effect evaluation of single-step attack actions, which can better meet the requirements of multi-step and multi-stage complex network environments; By building a hierarchical reinforcement learning system between the server and the client, the combination of macro and micro audit tasks is realized. It can not only more effectively discover system vulnerabilities, but also adapt to various scenario requirements and dynamically control risks, significantly enhancing the scalability and practical value of network security audit and automated penetration testing; Through the hierarchical design of the macro policy and the micro policy, the problem of state-action space explosion of traditional single-layer reinforcement learning in complex security audit scenarios is effectively solved, and intelligent security audit in complex network environments is realized; The macro layer of the present invention is responsible for subtask planning and high-level goal decomposition, and the micro layer focuses on the execution and optimization of specific attack actions. The two cooperate to significantly narrow the search space, improve the learning convergence speed and policy quality, and further introduce a multi-dimensional adaptive reward mechanism based on state entropy and dynamic adjustment factors. Through state-action exploration rewards, the common local optimum problem in reinforcement learning can be effectively avoided. A dynamic mapping relationship between the system state and the reward weight is established through the security state entropy weight method, thereby realizing intelligent risk control in the security audit process.

[0015] In summary, the present invention can automatically learn and optimize the attack sequence, greatly improve the vulnerability discovery efficiency and coverage rate, reduce the dependence on professional security personnel, lower the enterprise security audit cost, and has the ability to adapt to network environments of different complexities. It has strong discovery ability for new threats and unknown vulnerabilities, provides more proactive and forward-looking technical support for network security protection, and has important theoretical value and practical significance for improving the security protection ability of critical information infrastructure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention.

[0017] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to clearly illustrate the technical features of the present solution, the present invention will be described in detail below through specific embodiments and in conjunction with its drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention.

[0019] Example 1 An intelligent security audit method based on hierarchical reinforcement learning is as follows: Model the security audit process as a hierarchical Markov decision process. Conduct security audit operations by the server generating an attack sequence and sending it to the client. The attack sequence is generated based on hierarchical reinforcement learning, which includes a macro policy and a micro policy. Then the client executes attack actions according to the received attack sequence and collects feedback, and then sends the feedback to the server. The server updates the macro policy and the micro policy based on the feedback, iterates and optimizes the macro policy and the micro policy, sets an iteration stop condition. If the iteration condition is met, the security audit process ends; otherwise, continue to conduct security audit operations according to the updated macro policy and micro policy until the iteration stop condition is met and the security audit process ends.

[0020] In the specific implementation manner, the hierarchical reinforcement learning includes a macro policy and a micro policy . The macro policy is the high level in the hierarchical reinforcement learning, and the micro policy is the low level in the hierarchical reinforcement learning. The high level defines subtasks and high-level goals . . The low level defines executable micro actions . . denotes attack actions; The macro policy selects a subtask under the state of the system, expressed as . The micro policy selects a micro action under the subtask and the state of the system, expressed as . The state . denotes the network environment, system configuration information, and current attack progress.

[0021] In the specific implementation manner, the process of the macro policy selecting a subtask is as follows: The macro policy analyzes the current system state through threat intelligence information and the historical audit records of the system , dynamically determines the high-level goal that needs to be executed, and then based on the macro policy selects a suitable subtask from the high-level goals under the state Threat intelligence information includes data sources from network detection systems, intelligence sharing platforms, and security vulnerability databases; High-level goal Consists of a series of subtasks and , represents the th subtask in the high-level goal where represents the number of subtasks in the high-level goal ; Each subtask can be regarded as a high-level audit goal, for example: information collection subtask, vulnerability detection subtask, privilege escalation subtask, persistence and lateral movement subtask, and risk assessment subtask; The macro strategy calculates the macro value generated by immediate rewards through the macro value function to measure the advantages and disadvantages of different subtasks in terms of long-term benefits and risk control, and arranges the execution order and timing of subtasks. The specific process is as follows: The calculation formula for the macro value is as follows: , where represents the discount factor represents the expected value represents the preset number of planning time steps represents the state of the system at which the subtask has a macro value and represents the state of the system at the moment when the subtask is executed, and the immediate reward generated reflects the comprehensive index of benefits and risk control; Then, based on the macro value, priority determination operations are performed to arrange the execution order and timing of subtasks. Specifically, by setting a priority threshold , subtasks that meet are screened out as candidate subtasks, and then the candidate subtasks are sorted in descending order according to values to obtain the priority sequence of subtasks. Among them, the higher the value, the higher the execution priority; Finally, based on the priority sequence of subtasks and the current resource status of the system, on the premise that the state of the system allows the execution of subtasks, each candidate subtask is started in turn, and is updated in real time according to the task feedback during the execution process, and then the priorities and execution timing of subsequent tasks are adjusted.

[0022] In the specific implementation, the micro-strategy selects micro-actions The process is as follows: In the sub-task and the state of the system Based on the micro-strategy Select micro-actions , , the attack actions include instruction execution, port and traffic detection, vulnerability exploitation actions, privilege escalation actions, and resource usage and concealment-related actions; The micro-strategy evaluates the execution value of different micro-actions in each sub-task by defining the micro-value , and continuously updates the micro-strategy in combination with the immediate feedback, making the action selection in the audit process gradually tend to be optimal. The specific process is as follows: In the system initialization stage, for each sub-task and all possible micro-actions , the micro-value is assigned a random initial value; In the state of the system , when executing the sub-task , the micro-strategy selects a micro-action for execution according to the value of the current micro-value , and then adopts -greedy strategy to balance exploration and exploitation, and selects the action that maximizes the value with probability , and randomly selects other actions with probability . After the micro-action is executed, the state of the system changes from to , and an immediate feedback reward is obtained. After multiple iterations and continuous updates of the immediate feedback reward, the micro-action selected by the micro-strategy in the state of the system and the sub-task tends to be optimal according to the value of the micro-value .

[0023] In the specific implementation, the execution and feedback of the client: (1) At moment, the client receives the micro-action sent by the server, and based on the system state at moment and the corresponding sub-task Perform specific attack operations. During the execution of specific operations, the client performs attack actions on the audit object according to the instructions of the micro-actions . (2) After the client finishes the execution, the state of the system changes from to , indicating the latest state after performing the action . The state includes network topology and port changes, system configuration and resource usage, current attack progress and subtask completion indication information. When the completion condition is met, the subtask completion flag is set to 1, otherwise it is set to 0, whether a security exception or a defense mechanism trigger exception event occurs; (3) The client calculates the immediate feedback reward of the micro-strategy according to the update of the system state . The calculation formula is as follows: , where, in sequence, it includes five reward items, which are , , , and . represents the success degree of quantifying an attack operation, represents the time overhead of the attack operation, represents the security risk brought by the attack operation, represents the system resources consumed by the attack operation, represents the state-action exploration reward, , , , and respectively represent the dynamic weight coefficients of the next reward item in the state , , , and ; The calculation formula of the state-action exploration reward is as follows: , where, represents the exploration reward base coefficient, represents the total number of attack actions executed in history, represents 's index, represents the th attack action executed in history, represents the current attack action and The state-action space distance between represents the distance sensitivity adjustment factor; Determined by the multi-level security priority evaluation model and the security state entropy weight method , , The index representing the number of reward items, , The calculation formula of is as follows: , Among them, also represents the index of the number of reward items, represents the th preset basic weight of the reward item, represents the state th dynamic adjustment factor of the reward item, represents the state entropy of the state , represents the th preset basic weight of the reward item, represents the state th dynamic adjustment factor of the reward item; The state entropy of the state The calculation formula is as follows: , Among them, represents the number of system state classifications, represents the state belongs to the th class of states; The dynamic adjustment factor of the state The calculation formula of is: Among them, , Among them, represents the important index of the th reward item under the state , represents the normalization constant, and represent two different adjustment coefficients, represents the th adjustment coefficient of the reward item , represents the th adjustment coefficient of the reward item , represents the current attack progress index, represents the Historical success rate of each reward item; Historical success rate The calculation formula is as follows: , wherein, represents the total number of attack actions executed as of time, represents index of represents the th relevant action set of the reward item, represents the indicator function. If , then , otherwise , represents judging whether the th attack action is successful. If successful, , if failed, ; (4)The client sends the latest status , immediate reward , sub-task completion degree indication information and auxiliary diagnosis information back to the server in the form of data packets.

[0024] In the specific implementation manner, the iteration and optimization of the macro strategy and the micro strategy: (1)Update the micro strategy: The server performs strategy iteration according to the data packet fed back by the client and updates the micro value using a value-based method. The calculation process is as follows: , wherein, represents the learning rate used to control the fusion ratio of the new and old values, represents the discount factor, represents the candidate action that obtains the maximum value among all micro actions, represents the selected sub-task at time, (2)Update the macro strategy: When the sub-task is completed or switched, the server performs strategy iteration according to the profit and risk information of the macro strategy. The immediate reward of the macro strategy, the calculation formula is as follows: , wherein, Indicates the degree of coverage improvement for unknown network nodes, ports, and potential vulnerability areas. Indicates the success degree of a single attack operation quantification. Indicates the measurement of the risk of attack actions at the macroscopic level. Indicates the cost of executing an attack action; When completing or switching subtasks, the immediate reward according to the macroscopic strategy Update the macroscopic value , and the calculation formula is as follows: , Among them, Indicates the macroscopic learning rate, Indicates the discount factor, Indicates the subsequent value obtained by executing the subtasks of the optimal macroscopic strategy layer, Indicates the update operation; (3) Set the iteration stop condition: Set the change threshold of the value function. When the in the microscopic strategy and the in the macroscopic strategy both have update amplitudes lower than the given threshold, it is judged that the microscopic strategy layer is stable, specifically as follows: Define the update amplitude thresholds for the microscopic strategy and the macroscopic strategy respectively. For the microscopic strategy layer, set the microscopic threshold , and for the macroscopic strategy, set the macroscopic threshold ; In each iteration process, for all system states , subtasks and microscopic actions , calculate the maximum absolute change before and after the microscopic value update, and the calculation formula is as follows: , Among them, Indicates the maximum value calculation based on , Indicates the updated microscopic value, Indicates the microscopic value before update; At the same time, for all system states , calculate the maximum absolute change before and after the macroscopic value update, and the calculation formula is as follows: , Among them, Indicates the maximum value calculation based on , Indicates the updated macroscopic value, Indicates the macroscopic value before update; If within consecutive iterations, both and are satisfied, set . Then it is considered that the policy update has tended to be stable and the iteration can be terminated; When the set main security audit objective has been completed, or when the available resources are insufficient to continue the large-scale audit, the training process can also be aborted; If the performance of the audit policy still fails to improve significantly after exceeding the maximum number of iteration rounds, one can choose to exit or switch to the manual review mode according to specific requirements.

[0025] Embodiment 2 An intelligent security audit method based on hierarchical reinforcement learning, including a module for executing processing instructions for each step in an intelligent security audit method based on hierarchical reinforcement learning.

[0026] Embodiment 3 A device that executes an intelligent security audit method based on hierarchical reinforcement learning, the device including: a memory, a processor, and a program for the intelligent security audit method based on hierarchical reinforcement learning that is stored on the memory and can run on the processor.

[0027] Embodiment 4 A memory, on which a program of an intelligent security audit method is stored, and when the program of the intelligent security audit method is executed by a processor, it implements an intelligent security audit method based on hierarchical reinforcement learning.

[0028] Embodiment 5 In order to verify that the method proposed by the present invention has better performance compared with traditional rule-based security audit methods and basic single-layer reinforcement learning methods, an enterprise internal network security simulation environment is selected for experimental verification. This simulation environment constructs a virtual enterprise internal network, which includes elements such as servers, workstations, network devices, firewalls, etc., and integrates network monitoring, data collection, and log recording modules to simulate the security status and operation of the enterprise internal network; The experimental objective is to deploy the intelligent security audit system based on hierarchical reinforcement learning proposed by the present invention to achieve automated attack sequence optimization and improve the vulnerability discovery efficiency and coverage; in addition, under the same background, verify the performance of traditional rule-based security audit methods and basic single-layer reinforcement learning methods, and compare them with the method proposed by the present invention in terms of vulnerability discovery efficiency, location threat detection rate, average audit cost, critical vulnerability coverage rate, and corresponding APT attack time.

[0029] The process of deploying the method of the present invention to this virtual enterprise internal network is as follows: (1) Construct the server and client: Server: Macro Strategy Engine: Generate high-level goals (such as "obtain database permissions") based on threat intelligence (such as the MITRE ATT&CK framework); Micro Strategy Optimizer: Dynamically adjust attack actions (such as SQL injection, port scanning); Feedback Analysis Module: Calculate the reward function and update the strategy; Client: Attack Execution Agent: Deployed in each subnet, execute actions and collect status feedback.

[0030] Security State Sensor: Real-time monitor network topology and the triggering situation of defense mechanisms.

[0031] The functions of relevant core components in the server and client are shown in Table 1 specifically; Table 1 Comparison Table of Core Component Functions (2)Implementation Process and Dynamic Optimization (2-1)Macro Strategy Generation: Input: The current state s_t (such as "Web service exposes ports 22 / 80 / 443, and the defense mechanism is not triggered").

[0032] Output: A sequence of subtasks G={ : Port scanning, : Vulnerability exploitation, : Privilege escalation}, sorted by (s, ).

[0033] (2-2)Micro Strategy Execution: For the subtask (port scanning), select an action (such as "perform SYN stealth scan using Nmap").

[0034] Exploration-Exploitation Balance: ϵ = 0.2, with an 80% probability of selecting the action with the highest value, and a 20% probability of randomly trying a new action (such as "ICMP stealth scan").

[0035] (2-3)Feedback and Reward Calculation: Example of the reward function: , : 2 open ports are found during scanning → +0.4; : Firewall alarm is triggered → -0.3; ; Dynamic weight adjustment: If the state entropy is high (the network state is complex), the weight of (exploration reward) is increased to 0.15.

[0036] (2 - 4) Policy iteration optimization: Micro - policy update: , Macro - policy update: If the subtask is completed and (s, ) is increased by 10%, then the priority of its subsequent tasks is increased.

[0037] (3) Similarly, deploy the traditional rule - based security audit method and the basic single - layer reinforcement learning method on the internal network of this virtual enterprise, and then obtain the comparison results of the three methods. The test period is one month, as shown in Table 2 specifically; Table 2 Evaluation and comparison results of each method From the data in Table 2, it can be seen that due to relying on fixed rules, the traditional rule - based security audit method has a low detection rate of unknown threats and a low coverage rate of key vulnerabilities in a dynamic environment. At the same time, the average audit cost and the response time to APT attacks are high; the basic single - layer reinforcement learning method has a certain degree of self - adaptability, but in a complex network environment, there are still deficiencies in the vulnerability discovery efficiency, the detection rate of unknown threats, and the coverage rate of key vulnerabilities; while the method based on hierarchical reinforcement learning proposed in this application has a significant improvement in the vulnerability discovery efficiency, the detection rate of unknown threats, and the coverage rate of key vulnerabilities, and the number of iterative convergence is also greatly reduced, indicating that this method can perform security audits on the target system faster and more accurately, thus comprehensively improving the effect and efficiency of security detection.

[0038] Although the specific implementation manners of the invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the invention. Based on the technical solutions of the invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the invention.

Claims

1. An intelligent security audit method based on hierarchical reinforcement learning, characterized by: The security audit process is modeled as a hierarchical Markov decision process. The server generates an attack sequence and sends it to the client to perform security audit operations. The attack sequence is generated based on hierarchical reinforcement learning. Hierarchical reinforcement learning includes macro-strategy and micro-strategy. Then the client executes attack actions and collects feedback according to the received attack sequence, and then sends the feedback to the server. The server updates the macro-strategy and micro-strategy according to the feedback, iterates and optimizes the macro-strategy and micro-strategy, and sets the iteration stop condition. If the iteration condition is met, the security audit process ends. Otherwise, the security audit operation continues according to the updated macro-strategy and micro-strategy until the security audit process ends when the iteration stop condition is met.

2. According to claim 1, an intelligent security audit method based on hierarchical reinforcement learning is characterized by: Hierarchical reinforcement learning including macro strategies and micro strategies , macro strategy For high-level, micro-strategies in hierarchical reinforcement learning It is the lower layer in hierarchical reinforcement learning, and the subtasks are defined by the higher layer. and high-level goals , , executable micro-actions defined by the lower layer , , Indicates an attack action; The status of macro strategy in the system Select Subtask , expressed as , micro-strategies in subtasks and the status of the system Select Micro Actions , expressed as ,state , Indicates the network environment, system configuration information, and current attack progress.

3. The intelligent security audit method based on hierarchical reinforcement learning according to claim 2 is characterized in that: Macro Strategy Selection Subtask The process is as follows: Macro strategies analyze the current system status through threat intelligence information and historical audit records of the system , dynamically determine the high-level goals that need to be executed , and then based on the macro strategy Status in the system From high-level goals Choose the appropriate subtask ; Threat intelligence information includes data sources from network detection systems, intelligence sharing platforms, and security vulnerability databases; High-level goals A series of subtasks composition, , Represents high-level goals Middle subtasks, Represents high-level goals The number of neutron tasks, ; The macro strategy calculates the macro value generated by the immediate reward through the macro value function and arranges the execution order and timing of the subtasks. The specific process is as follows: The calculation formula of macro value is as follows: , in, represents the discount factor, represents the expected value, represents the preset planning time steps, Indicates the status of the system Next subtask The macro value of Indicated in The status of the system at the moment Execute subtasks The immediate reward generated after Then, according to the macro value, priority is determined and the execution order and timing of subtasks are arranged. Specifically, priority thresholds are set. , filter out Subtasks As candidate subtasks, the candidate subtasks are then The values ​​are sorted in descending order to obtain the priority sequence of the subtasks, among which the top ranked A higher value has a higher execution priority; Finally, according to the priority sequence of subtasks and the current resource status of the system, the system status is ensured. Under the premise of allowing subtasks to be executed, each candidate subtask is started in sequence, and updated in real time according to task feedback during execution .

4. According to claim 3, the intelligent security audit method based on hierarchical reinforcement learning is characterized in that: Micro-strategy selection of micro-actions The specific process is as follows: In subtask and the status of the system Based on micro-strategy Select Micro Actions , , attack action Including command execution, port and traffic detection, vulnerability exploitation, privilege escalation, and resource usage and concealment related actions; Micro-strategy defines micro-value In each subtask Evaluate different micro actions The execution value of the micro-strategy is continuously updated with instant feedback. The specific process is as follows: During the system initialization phase, for each subtask And all possible micro-movements , micro value are assigned random initial values; Status in the system Next, execute the subtask When the micro-strategy is based on the current micro-value The value of selects a micro action Execute and then use - Greedy strategy to balance exploration and exploitation, based on probability Select The action with the largest value, according to the probability Randomly select other actions, micro actions After being executed, the system status is changed by Transformed into , and receive instant feedback rewards After multiple iterations and continuous updates of instant feedback rewards, the micro-strategy is in the state of the system and subtasks According to micro value The micro-actions selected by the value of tend to be optimal.

5. The intelligent security audit method based on hierarchical reinforcement learning according to claim 4 is characterized in that: Client execution and feedback: (1) At this moment, the client receives the micro-action sent by the server , and based on System status at the moment and the corresponding subtasks Carry out specific attack operations. During the execution of specific operations, the client performs micro-actions The instructions execute attack actions on the audit object; (2) After the client completes the execution, the system status is changed to Transformed into , Indicates execution of an action The latest status after the attack, including network topology and port changes, system configuration and resource usage, current attack progress and subtask completion indication information, whether there are security anomalies or defense mechanisms trigger abnormal events; (3) The client calculates the immediate feedback reward of the micro-strategy based on the update of the system status , the calculation formula is as follows: , Among them, there are five reward items, namely , , , and , It indicates the degree of success of a quantified attack operation. represents the time cost of the attack operation, Indicates the security risks brought by the attack operation. Indicates the system resources consumed by the attack operation. represents the state-action exploration reward, , , , and Respectively in the state Next reward , , , and Dynamic weight coefficient of Status-Action Exploration Reward The calculation formula is as follows: , in, represents the basic coefficient of exploration reward, Indicates the total number of attack actions executed in history. express The index of Indicates the number of historical executions An attack action. Indicates the current attack action and The state-action space distance between represents the distance sensitivity adjustment factor; Determined through a multi-level security priority assessment model and security status entropy weight method , , An index representing the number of reward items, , The calculation formula is as follows: , in, It also represents the index of the number of reward items. Indicates The default basic weights of each reward item are: Indicates status No. Dynamic adjustment factor for each reward item, Indicates status The state entropy of Indicates The default basic weights of the rewards are: Indicates status No. Dynamic adjustment factors for each reward item; state The calculation formula of the state entropy is as follows: , in, Indicates the number of system status categories, Indicates status Belong to Probability of class status; state Dynamic adjustment factor The calculation formula is: , in, Indicates status Next The important indicators of the award items are: represents the normalization constant, and represents two different adjustment coefficients, Indicates The adjustment factor of the reward item , Indicates Adjustment factor for each reward item , Indicates the current attack progress indicator, Indicates Historical success rate of each reward item; Historical success rate The calculation formula is as follows: , in, Indicates that as of The total number of attack actions executed at any moment. express The index of Indicates The set of actions related to the reward items, represents the indicator function, if ,but ,otherwise , Indicates the judgment Attack Action Is it successful? If successful, If it fails, ; (4) The client will update the latest status , Instant Rewards , subtask completion indication information and auxiliary diagnosis information are sent back to the server in the form of data packets.

6. The intelligent security audit method based on hierarchical reinforcement learning according to claim 5 is characterized in that: Iteration and optimization of macro and micro strategies: (1) Update micro-strategies: The server responds to the data packet sent by the client. Perform strategy iteration and use value-based methods to analyze micro-values To update, the calculation process is as follows: , in, Represents the learning rate used to control the ratio of new and old value fusion, represents the discount factor, Indicates the maximum gain among all micro actions The candidate actions for the value, express The subtask selected at the moment, The value of attack actions representing the optimal micro-strategy; (2) Update macro strategy: When the service segment completes or switches subtasks, it iterates the strategy based on the benefit and risk information of the macro strategy, and the immediate reward of the macro strategy , the calculation formula is as follows: , in, Indicates the degree of coverage improvement of unknown network nodes, ports, and potential vulnerability areas. It indicates the degree of success of a quantified attack operation. Indicates the risk of measuring attack actions at the macro level. Represents the cost of performing an attack action; When completing or switching sub-tasks, instant rewards based on macro strategies Update macro value , the calculation formula is as follows: , in, represents the macro learning rate, represents the discount factor, represents the subsequent value obtained by executing the subtasks of the optimal macro-strategy layer, Indicates an update operation; (3) Set the iteration stop condition: Set the value function change threshold when the micro-strategy and macro strategies When the update amplitude is lower than the given threshold, the micro-strategy layer is judged to be stable, as follows: Define update amplitude thresholds for micro strategies and macro strategies respectively. For the micro strategy layer, set the micro threshold , set macro thresholds for macro strategies ; In each iteration, the status of all systems , Subtask And micro-action , calculate the maximum absolute change before and after the micro value update , the calculation formula is as follows: , in, Indicates based on Calculate the maximum value. represents the updated micro value, Indicates the micro value before updating; At the same time, for all system states , calculate the maximum absolute change before and after the macro value is updated , the calculation formula is as follows: , in, Indicates based on Calculate the maximum value. represents the updated macro value, Indicates the macro value before updating; If in continuous In all iterations, and , it is considered that the strategy update has stabilized and the iteration can be terminated; The training process may also be terminated when the major security audit objectives set have been achieved or when the available resources are insufficient to continue with a large-scale audit; If the audit policy performance is still not significantly improved after exceeding the maximum number of iterations, you can choose to exit or switch to manual review mode based on specific needs.

7. An intelligent security audit method based on hierarchical reinforcement learning, characterized by: Including for executing claims A module for processing instructions at each step in an intelligent security audit method based on hierarchical reinforcement learning as described in any one of the claims.

8. A device for executing the claim An intelligent security audit method based on hierarchical reinforcement learning as described in any one of the claims, characterized in that: The device comprises: a memory, a processor, and a program stored in the memory and capable of running an intelligent security audit method based on hierarchical reinforcement learning on the processor.

9. A memory, characterized in that: The memory stores a program of an intelligent security audit method, and when the program of the intelligent security audit method is executed by the processor, the program of the intelligent security audit method is implemented as claimed in claim An intelligent security auditing method based on hierarchical reinforcement learning as described in any one of the above.

Citation Information

Patent Citations

  • Collaborative penetration testing method and system based on heterogeneous layered reinforcement learning, and medium

    CN117834283A

  • Intelligent security auditing method, system and equipment based on reinforcement learning, and memory

    CN118473767A

  • Oil depot tank field inspection robot task allocation method and system based on Internet of Things

    CN119417192A

  • Reinforcement-learning based queue-management method for undersea networks

    US20230319163A1