A dynamic web vulnerability detection method, device, equipment and storage medium

By using real-time system monitoring data, partial differential equation dynamic model and MDP decision model, dynamic detection of web vulnerabilities and determining defense strategies, the problem of low accuracy of web vulnerability detection in the existing technology is solved, and more efficient Web system security is achieved.

CN119628960BActive Publication Date: 2025-05-13JIANGSU QITONG HAOMIAO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510135929.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-13
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The existing web vulnerability detection methods cannot fully identify all web vulnerabilities, resulting in low detection accuracy.

Method used

By acquiring real-time system monitoring data, using the pre-constructed dynamic model of partial differential equations and the MDP decision model, the current system security state and target defense action set are output, and the target defense strategy is determined to improve detection accuracy.

Benefits of technology

Improve the accuracy of web vulnerability detection, enhance the security of the web system, and can promptly and accurately detect web system vulnerabilities and maximize defense effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119628960B_ABST
    Figure CN119628960B_ABST
Patent Text Reader

Abstract

The present invention provides a dynamic Web vulnerability detection method, device, electronic device and storage medium. The method includes using a pre-set partial differential equation dynamic model to output the current system security state and target defense action set based on real-time system monitoring data, and based on a pre-built MDP decision model, selecting the best defense action execution sequence based on the target value corresponding to each target defense action included in the target defense action set as the target defense strategy. By applying the embodiment of the present invention, by calculating the system security state based on real-time system monitoring data, the Web system vulnerability can be discovered timely and accurately through the system security state. At the same time, by using the partial differential equation dynamic model to output the target defense action set based on the system monitoring data, and using the MDP model to determine the current target defense action execution sequence, the defense effect can be maximized and the security of the Web system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a dynamic Web vulnerability detection method, device, equipment and storage medium. Background Art

[0002] With the rapid development of network technology, network security has become an issue that needs to be focused on in practical applications. In the field of network security, Web vulnerability detection is a means to effectively improve system security. Web vulnerability detection refers to the process of comprehensively scanning and analyzing Web applications, websites or servers through various technical means and tools to discover security vulnerabilities in them.

[0003] In the related technologies, Web vulnerability detection is usually performed through black box testing, white box testing, gray box testing, etc., or using existing Web vulnerability scanning tools. However, with the increasing number of Web attack methods, the current Web vulnerability detection methods cannot fully identify all Web vulnerabilities, resulting in low accuracy of Web vulnerability detection. Summary of the invention

[0004] In view of this, embodiments of the present invention provide a dynamic Web vulnerability detection method, apparatus, device and storage medium to improve the accuracy of Web vulnerability detection and thereby improve the security of a Web system.

[0005] According to one aspect of the present invention, a dynamic Web vulnerability detection method is provided, the method comprising:

[0006] Acquire real-time system monitoring data, wherein the real-time system monitoring data includes resource usage data and access flow data;

[0007] Based on the real-time system monitoring data and the pre-built partial differential equation dynamic model, output the current system security state and the target defense action set, wherein the partial differential equation dynamic model is used to define the association between the system security state and the attack data and the defense data, and the target defense action set includes at least one defense action;

[0008] Calculating the target value corresponding to each target defense action based on the current system security state and each target defense action included in the target defense action set using a pre-built MDP decision model; the MDP decision model is used to define the association between the value function and the defense action and the system security state;

[0009] Based on the target values ​​corresponding to the target defense actions, the target defense action execution order with the highest sum of the target values ​​is determined as the target defense strategy.

[0010] In a possible embodiment, the partial differential equation dynamic model is pre-constructed by the following steps:

[0011] Constructing a partial differential equation, wherein the partial differential equation defines a relationship between a system security state and attack data and defense data, wherein the attack data includes an attack target and an attack method obtained based on a preset attack behavior model, and the defense data includes a defense action taken;

[0012] Constructing a cost function, wherein the cost function includes the cost corresponding to the adopted defense action and the security assessment result of the system at a preset time point;

[0013] A Hamiltonian is constructed based on the partial differential equation and the cost function to obtain the partial differential equation dynamic model.

[0014] In a possible embodiment, the partial differential equation is expressed by the following formula:

[0015]

[0016] Wherein, u(x,t) represents the safety state of the system at time t and spatial position x, k is the preset conduction coefficient, represents an external control variable, wherein the external control variable at least includes an attack target, an attack method, and a defense action obtained based on a preset attack behavior model, Indicates the influence of the external control variable on the system safety;

[0017] The cost function is expressed by the following formula:

[0018]

[0019] in, is the cost corresponding to the adopted defensive action, T is the preset time point, as well as is the preset space type range, is the safety assessment result of the system at the preset time point T;

[0020] The Hamiltonian is expressed by the following formula:

[0021]

[0022] Wherein, H is the Hamiltonian, is the accompanying variable.

[0023] In a possible embodiment, outputting the current system security state and the target defense action set based on the real-time system monitoring data and the pre-built partial differential equation dynamic model includes:

[0024] Optimizing the Hamiltonian and determining a target external control variable corresponding to the optimal Hamiltonian;

[0025] Calculating partial derivatives of the target external control variable based on the defensive action types included in the external control variable, and determining based on each partial derivative result that the defensive actions corresponding to the extreme values ​​of the external control variable constitute a target defensive action set;

[0026] The Hamiltonian is solved based on the target defense action set and the accompanying variables to obtain the current system safety state.

[0027] In a possible embodiment, the method of obtaining partial derivatives of the target external control variables based on the defense action types included in the external control variables, and determining the defense actions corresponding to the extreme values ​​of the external control variables to form a target defense action set based on the partial derivative results includes:

[0028] Taking partial derivatives of the target external control variables based on the attack data types included in the target external control variables, and determining that the attack data corresponding to the external control variable mechanism constitutes an optimal attack strategy set based on each of the partial derivative results; the attack data types at least include an attack target and an attack method;

[0029] updating the target external control variable based on the optimal attack strategy set;

[0030] The partial derivatives of the target external control variables are calculated based on the defense action types included in the updated external control variables, and the defense actions corresponding to the extreme values ​​of the external control variables are determined to constitute a target defense action set based on the partial derivative results.

[0031] In a possible embodiment, the MDP decision model is pre-built through the following steps:

[0032] Based on the state rewards corresponding to each defense action, a profit function is constructed; the state rewards include resource consumption and / or system security state changes corresponding to the defense action;

[0033] Constructing a state transition probability function based on the state reward corresponding to each of the defensive actions and the current system security state;

[0034] A value function is constructed based on the benefit function and the state transition probability as the MDP decision model.

[0035] In a possible embodiment, the using of the pre-built MDP decision model based on the current system security state and each target defense action included in the target defense action set to calculate the target value corresponding to each target defense action includes:

[0036] Calculating the value function value corresponding to each first target defense action included in the target defense action set by using the MDP decision model;

[0037] Calculating the value function value of each remaining defense action corresponding to each first target defense action by using the MDP decision model, wherein the remaining defense actions are other defense actions in the target defense action set except the first target defense action;

[0038] Determine the defense action with the highest value function value among the remaining defense actions as the second target defense action corresponding to the first target defense action;

[0039] The second target defense action is used as the first target defense action, and the step of using the MDP decision model to calculate the value function value of each remaining defense action corresponding to each first target defense action is returned until the remaining defense action corresponding to each first target defense action is empty, thereby obtaining a plurality of defense action execution orders;

[0040] The step of determining the target defense action execution order with the highest sum of target values ​​based on the target values ​​corresponding to the target defense actions as the target defense strategy includes:

[0041] Based on the sum of the value function values ​​corresponding to the execution sequences of the defense actions, the execution sequence of the defense actions with the largest sum of the value function values ​​is determined as the target defense strategy.

[0042] According to another aspect of the present invention, a dynamic Web vulnerability detection device is provided, the device comprising:

[0043] An acquisition module, used to acquire real-time system monitoring data, wherein the real-time system monitoring data includes resource usage data and access flow data;

[0044] An output module, configured to output the current system security state and a target defense action set based on the real-time system monitoring data and a pre-built partial differential equation dynamic model, wherein the partial differential equation dynamic model is used to define the association between the system security state and the attack data and the defense data, and the target defense action set includes at least one defense action;

[0045] A calculation module, configured to calculate the target value corresponding to each target defense action based on the current system security state and each target defense action included in the target defense action set by using a pre-built MDP decision model; the MDP decision model is configured to define the association between the value function and the defense action and the system security state;

[0046] The determination module is used to determine the target defense action execution order with the highest sum of target values ​​as the target defense strategy based on the target values ​​corresponding to the target defense actions.

[0047] According to another aspect of the present invention, there is provided an electronic device, comprising:

[0048] Processor; and

[0049] Memory for storing programs,

[0050] The program includes instructions, and when the instructions are executed by the processor, the processor executes any of the above-mentioned dynamic Web vulnerability detection methods.

[0051] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute any of the above-mentioned dynamic Web vulnerability detection methods.

[0052] One or more technical solutions provided in the embodiments of the present invention obtain real-time system monitoring data of the Web system, and use a pre-set partial differential equation dynamic model to output the current system security state and target defense action set based on the real-time system monitoring data, and then select the best defense action execution order based on the target value corresponding to each target defense action included in the target defense action set based on the pre-built MDP decision model as the target defense strategy. By applying the embodiments of the present invention, by calculating the system security state based on real-time system monitoring data, the Web system vulnerabilities can be discovered timely and accurately through the system security state. At the same time, by using the partial differential equation dynamic model to output the target defense action set based on the system monitoring data, and using the MDP model to determine the current target defense action execution order, the defense effect can be maximized and the security of the Web system can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Further details, features and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0054] Figure 1 A schematic diagram of a flow chart of a dynamic Web vulnerability detection method provided by an embodiment of the present invention;

[0055] Figure 2 A schematic diagram of a process for solving a target defense action set and a system security status in an embodiment of the present invention;

[0056] Figure 3 A schematic diagram of a process for solving a target defense strategy in an embodiment of the present invention;

[0057] Figure 4A logical structure diagram of a dynamic Web vulnerability detection device provided by an embodiment of the present invention

[0058] Figure 5 A block diagram of an exemplary electronic device that can be used to implement an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0059] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0060] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0061] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0062] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0063] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.

[0064] In order to improve the accuracy of Web vulnerability detection and thus improve the security of Web systems, the embodiments of the present invention provide a dynamic Web vulnerability detection method, device, electronic device and storage medium. The dynamic Web vulnerability detection method provided by the embodiments of the present invention can be applied to any electronic device with a dynamic Web vulnerability detection function, such as a server, a computer or a mobile terminal. The following describes the solution of the present invention with reference to the accompanying drawings:

[0065] Figure 1 A flow chart of a dynamic Web vulnerability detection method provided by an embodiment of the present invention may include the following steps:

[0066] S101, acquiring real-time system monitoring data, wherein the real-time system monitoring data includes resource usage data and access flow data;

[0067] S102, based on the real-time system monitoring data and the pre-built partial differential equation dynamic model, output the current system security state and the target defense action set, wherein the partial differential equation dynamic model is used to define the association relationship between the system security state and the attack data and the defense data, and the target defense action set includes at least one defense action;

[0068] S103, using a pre-built MDP (Markov Decision Process) decision model based on the current system security state and each target defense action included in the target defense action set, calculating the target value corresponding to each target defense action; the MDP decision model is used to define the association relationship between the value function and the defense action and the system security state;

[0069] S104. Based on the target values ​​corresponding to the target defense actions, determine the target defense action execution order with the highest sum of the target values ​​as the target defense strategy.

[0070] In an embodiment of the present invention, by acquiring real-time system monitoring data of a Web system, and using a pre-set partial differential equation dynamic model to output the current system security state and a target defense action set based on the real-time system monitoring data, a pre-built MDP decision model is then used to select the best defense action execution sequence based on the target value corresponding to each target defense action included in the target defense action set as a target defense strategy. By applying an embodiment of the present invention, by calculating the system security state based on real-time system monitoring data, Web system vulnerabilities can be discovered in a timely and accurate manner through the system security state. At the same time, by using a partial differential equation dynamic model to output a target defense action set based on system monitoring data, and using an MDP model to determine the current target defense action execution sequence, the defense effect can be maximized and the security of the Web system can be improved.

[0071] The above S101-S104 are exemplarily described below:

[0072] In S101, the request traffic data of the Web application and the resource usage data in the system can be monitored in real time. The above-mentioned request traffic data may include information such as the interface accessed by the traffic and the access time. The above-mentioned resource usage data may include CPU usage data, memory usage data, etc. The above-mentioned usage data may include resource utilization, resource usage, etc.

[0073] In a possible embodiment, the flow and resource usage in the system can be monitored by a system monitoring tool, which can be Prometheus, Nagios, etc. As a possible implementation, the request flow data and resource usage data can be obtained by the above monitoring tool at a preset time interval, and the time interval can be preset according to the actual application scenario, such as 1s, 0.5s, etc.

[0074] After obtaining the real-time system monitoring data, the system monitoring data can be input into a pre-established partial differential equation dynamic model, which is established based on a partial differential equation, wherein the partial differential equation is a partial derivative equation containing an unknown function. In a possible embodiment, the partial differential equation can be constructed based on the system security state and the external control variable, so that the partial differential equation dynamic model can define the relationship between the system security state and the external control variable, and then the partial differential equation dynamic model can be used to identify the system security state and output a set of defensive actions that can be taken.

[0075] The type of the above external control variable can be preset. As a possible implementation, the external control variable can include attack data and defense data, wherein the attack data can be obtained based on the attacker behavior model and real-time monitoring data, and can include attack targets, attack methods, attack timing, attack intensity, and resource allocation, etc. The attack intensity can be determined based on the number of abnormal machines and private data stolen caused by the attack. The above defense data can include current defense actions, defense intensity, defense timing, response speed, and resource allocation. The current defense action can be preset according to the actual application scenario, such as the default defense strategy including access rights, bandwidth restrictions, etc.

[0076] In a possible embodiment, an attacker behavior model can be obtained by training historical attack records, and the historical attack records can include historical attack targets, attack methods, attack intensity, and corresponding traffic data, resource usage data, etc. For example, traffic data and resource usage data can be used as training data, and attack targets, attack methods, and attack intensity can be used as training data labels to train the model, so that the model can predict information such as attack targets, attack methods, and attack intensity based on traffic data and resource usage data. In a possible embodiment, the attack target, attack method, and attack intensity can be obtained by setting analysis rules. The above attack data can also be obtained by combining existing vulnerability scanning tools, such as OWASP ZAP, Nmap (Network Mapper), and Goby.

[0077] In a possible embodiment, the partial differential equation dynamic model can be pre-constructed by the following steps:

[0078] S121. Construct a partial differential equation, in which the relationship between the system security state and the attack data and defense data is defined, wherein the attack data includes an attack target and an attack method obtained based on a preset attack behavior model, and the defense data includes a defense action taken.

[0079] In a possible embodiment, the system security status may be defined in association with time and space, that is, the relationship between the system security status of the space determined by time and the attack data and defense data may be determined. The space may refer to different types of system security. For example, a Web system may include storage space, memory space, network transmission space, etc. Based on the system security status corresponding to each space, the overall system security status may be obtained. The system security status corresponding to the above space may include whether the corresponding space is attacked, whether there is a vulnerability, etc.

[0080] As a possible implementation, the influence relationship of the external control variable on the system security state can be preset, and the external control variable can include the above-mentioned attack data, defense data and external influencing factors, and the external influencing factors can include machine hardware status, operating system vulnerabilities and industry specifications, etc. Exemplarily, the influence relationship between the external control variable and the system security state can be determined based on historical data, such as the influence relationship between the attack data and the system security state can be obtained based on the influence of historical attack data on the system security state. Specifically, the influence relationship may include the influence of the attack data on the system resource usage data, response time, response success rate, etc.

[0081] In a possible embodiment, when constructing partial differential equations, the conduction of system security impacts caused by attacks can also be considered. The system security impact conduction refers to the possible impact of the system security status at a certain moment or a certain space on the system security at the next moment or other space, so as to improve the comprehensiveness of the determination of the system security status.

[0082] Exemplarily, the above partial differential equation can be expressed by the following formula:

[0083]

[0084] Among them, u(x, t) represents the system security status of the Web system at time t and spatial position x, and k is a preset conduction coefficient, which can reflect the impact of u(x, t) on the system security status at the next moment or other spatial locations. The conduction coefficient can be set according to the actual application scenario. represents external control variables, which include the above-mentioned attack data, defense data, and external factors, etc. Indicates the impact of the external control variable on the system security, which may include the impact of attack data on the system security state, the impact of defense data on the system security state, and the impact of external factors on the system security state.

[0085] S122: construct a cost function, wherein the cost function includes the cost corresponding to the adopted defense action and the security evaluation result of the system at a preset time point.

[0086] In practical applications, various defense actions require consumption of resources or maintenance budgets, etc. For example, adjusting firewall rules will cause additional resource consumption, and increasing bandwidth will increase the bandwidth budget. In the present invention, the total amount of resources consumed during the execution of the defense action and the maintenance budget can be defined as the cost corresponding to the defense action. As a possible implementation method, a cost can be pre-set for each defense action based on the resources consumed by various defense actions and the maintenance budget, and the defense action and the cost can be stored in correspondence.

[0087] The above preset time point can be preset according to the actual application scenario. For example, the preset time point can be determined based on the preset maximum duration and the detection time. For example, if the preset duration is 10s and the current time is t, then the preset time point is t+10s. The preset maximum duration can be the maximum duration that the Web system can be attacked. The preset time point can also be the time point when the system security state reaches the preset security state.

[0088] The security assessment result of the system at the preset time point, that is, the security status of the system at the preset time point, may include whether the system has been attacked, whether the data has been leaked, etc.

[0089] When solving the optimal control rule, that is, when solving the optimal defense action set that can be taken, the minimum value of the cost function can be used as the optimization goal. Correspondingly, the above cost function can be expressed by the following formula:

[0090]

[0091] in, is the cost corresponding to the adopted defensive action, T is the preset time point, as well as is the preset space type range, It is the safety assessment result of the system at the preset time point T.

[0092] S123, constructing a Hamiltonian based on the partial differential equation and the cost function to obtain a dynamic model of the partial differential equation

[0093] The goal of determining the optimal control strategy in the present invention is to make the system safety state meet the preset safety state conditions while consuming the lowest cost, that is, it is necessary to solve the above partial differential equations and cost functions simultaneously. In a possible embodiment, the partial differential equations and cost functions can be solved simultaneously by Hamiltonian. Hamiltonian is an important physical concept in classical mechanics and quantum mechanics. It represents the total energy of the system. In the present invention, the total energy of the Web system can be represented by Hamiltonian, and the total energy can include the system safety state and the system cost.

[0094] In a possible embodiment, the above partial differential equation and cost function can be simply added to obtain the Hamiltonian, but since the external control variable will affect both the system safety state and the system consumption cost, this will cause the system safety state and the system consumption cost to influence each other to a certain extent, and it is easy to find that the system safety state and the system consumption cost are difficult to balance, thereby increasing the difficulty of solving the optimal control rule. Therefore, in a possible embodiment, a Hamiltonian can be constructed based on partial differential equations, cost functions and accompanying variables, wherein the accompanying variable is a variable that has no causal relationship with the dependent variable, but changes regularly with the change of the dependent variable. In an embodiment of the present invention, the accompanying variable may include system resource configuration, data protection measures, etc.

[0095] Exemplarily, the above Hamiltonian can be expressed by the following formula:

[0096]

[0097] Wherein, H is the Hamiltonian, is the accompanying variable.

[0098] In a possible embodiment, Figure 2As shown in the figure, the system security status and target defense action set can be solved based on the system monitoring data through the following steps:

[0099] S124, optimizing the Hamiltonian and determining a target external control variable corresponding to the optimal Hamiltonian;

[0100] In the embodiment of the present invention, the Hamiltonian can be optimized by any feasible method. For example, the Hamiltonian can be optimized by separation of variables method, solving functional extreme value method, etc. to obtain the external control variable corresponding to the optimal Hamiltonian. .

[0101] S125. Calculate partial derivatives of the target external control variables based on the defense action types included in the external control variables, and determine, based on the partial derivative results, that the defense actions corresponding to the extreme values ​​of the external control variables constitute a target defense action set.

[0102] As described above, the external control variables include attack data and defense data, that is, the variable types included in the external control variables include multiple types. Therefore, as a possible implementation method, the external control variables can be used to find partial derivatives of various types of data. According to the properties of the partial derivatives, the value corresponding to the partial derivative being 0 is the extreme point of the function, that is, the data with the greatest impact on the external control variables. Therefore, the data type corresponding to the partial derivative being 0 can be determined based on the partial derivative results as a defense action. The above data types involved in the partial derivative solution can be set according to the actual application scenario, such as various defense data and / or various attack data.

[0103] Exemplarily, the external control variables include attack variables and defense data, wherein the defense data may include various preset defense actions that may be taken, and the preset defense actions may be pre-set. Exemplarily, at the network level, defense may be performed by adjusting firewall rules or increasing bandwidth restrictions, and at the application level, defense may be performed by updating security patches, adjusting the sensitivity of intrusion detection systems, and the like. In the above external control variables, each defense data may be defined as an unknown quantity x1, x2…xn. Accordingly, the external control variables may be used to calculate partial derivatives of x1, x2…xn, respectively, to obtain partial derivative results corresponding to each type of defense data, and to determine each defense data whose partial derivative results are 0. Exemplarily, if the partial derivative results corresponding to the firewall rules and bandwidth restrictions are 0, these two defense actions may be used as target defense actions, and each target defense action constitutes a target defense action set.

[0104] In a possible embodiment, partial derivatives of various types of attack data can be obtained based on external control variables, and the optimal attack strategy can be obtained based on the partial derivative results, and the optimal attack strategy can be brought into the external control variables as a known quantity to update the target external control variables, and the target defense action can be solved based on the updated target external control variables. In this way, by predicting the optimal attack strategy based on the attack data, and then outputting the target defense action set based on the optimal attack strategy, the target defense action set with the highest defense efficiency can be obtained by playing a game between attack and defense.

[0105] S126. Solve the Hamiltonian based on the target defense action set and the accompanying variables to obtain a current system safety state.

[0106] In this step, the adjoint equation corresponding to the above partial differential equation can be solved to obtain the adjoint variable, wherein the adjoint equation is a differential equation that has a conjugate relationship with the given differential equation. In the present invention, the adjoint equation can be solved by any feasible method to obtain the adjoint variable, such as by separation of variables, canonical transformation, etc. The above target defense action set and the adjoint variable are brought into the Hamiltonian to obtain the current system security state.

[0107] In a possible embodiment, the theoretical safety state of the system after taking the target defense action set can be obtained based on a preset system state equation, and the target defense action set can be further judged based on the theoretical safety state. The above system state equation can define the correlation between the rate of change of the system safety state over time and the defense action taken and the current system state. For example, the system state equation can be:

[0108] du / dt=f(u,w,v)

[0109] Among them, u represents the system security state, w represents the defense action set, that is, the above-mentioned target defense action set, and v represents the attack strategy, which may include vulnerability scanning, vulnerability exploitation, etc. The initial value u(t0) of the system security state is usually known and can be determined based on the number of known vulnerabilities, attack intensity, etc. After obtaining the target defense action set w(t) at time t, the state equation can be solved by analytical solution or numerical method Euler method to obtain the evolution trajectory u(t) of the system security state u over time. If u(t) indicates that the system security state is a security state such as not being attacked or there is no data leakage, it can be determined that the target defense action set is reasonable.

[0110] In a possible embodiment, it may be impossible to execute all defense actions in the target defense action set at the same time. Therefore, a pre-built MDP decision model can be used to calculate the target value corresponding to each target defense action based on the current system security state and each target defense action included in the target defense action set; the MDP decision model is used to define the association relationship between the value function and the defense action and the system security state;

[0111] Based on the target values ​​corresponding to the target defense actions, the target defense action execution order with the highest sum of the target values ​​is determined as the target defense strategy.

[0112] The Markov decision process is used to describe how a system makes the best decision based on the current state when facing uncertainty. In a possible embodiment, the above MDP decision model can be defined by the Bellman equation. For example, the MDP decision model can be pre-built by the following steps:

[0113] S131. Construct a profit function based on the state reward corresponding to each defense action; the state reward includes resource consumption and / or system security state change corresponding to taking the defense action.

[0114] The benefit function is a function used to define the immediate return of each state-action pair in the MDP. It determines the "immediate reward" or "cost" that the system can obtain after taking a certain action in a certain state. The immediate reward reflects the effect of each decision or action. In the security defense scenario, the benefit function can represent the reward for successful defense, which can include avoiding data leakage, preventing attacks, etc. The benefit function can also represent the cost of taking a certain defense measure, such as resource consumption, time delay, etc. The calculation result of the benefit function can include positive rewards and negative rewards, among which positive rewards refer to the execution results that provide gains to the system, such as preventing an attack and protecting data security, and negative rewards refer to the execution results that cause consumption to the system, such as resource consumption, such as the cost of defense measures such as increasing computing power and increasing monitoring, etc.

[0115] For example, the profit function can be expressed as Indicates that this formula represents the state at time step t. Take action Instant rewards obtained when

[0116] The status rewards corresponding to the above-mentioned defense actions are preset and may include the impact of the defense action on the system security status, resource usage data, and the cost consumed.

[0117] S132: construct a state transition probability function based on the state rewards corresponding to each of the defense actions and the current system security state.

[0118] The state transition probability refers to the probability that the system will transition to another state after executing a certain defense action in a given state. In other words, the state transition probability can reflect the success rate of different defense measures. For example, firewall rule adjustments may reduce the probability of a successful attack. The state transition probability corresponding to each defense action can be obtained based on historical defense records. Exemplarily, the system security state corresponding to each defense action can be determined based on historical defense records, the candidate states corresponding to the defense action can be determined, and the transition probability of each candidate state corresponding to the defense action can be determined based on the proportion of each candidate state in the system security state corresponding to the defense action.

[0119] In a possible embodiment, the state transition probability may be obtained based on the state reward corresponding to the defense action and the current system security state. For example, the state transition probability may be determined based on the superposition result of the system security state corresponding to the defense action and the current system security state.

[0120] For example, in a defense system, assuming that the current state of the system is "undefended", if the defender performs a certain defense measure (such as enabling IDS (intrusion detection system)), the system may be transferred to the "partial defense" or "full defense" state. The probability of state transition may depend on the current network traffic, the attacker's strategy, etc. For example: if the current system state is "partial defense", after executing the "enhance firewall" action, the system will be transferred to the "full defense" state with an 80% probability and transferred to the "attack successful" state with a 20% probability.

[0121] Exemplarily, the above state transition probability can be expressed by the following formula:

[0122]

[0123] in, represents the system safety status at time t, represents the defensive action taken at time t, and u′ represents the state at the next time t+1. It represents the probability that the system transfers to state u′ after executing action a in state u.

[0124] S133. Construct a value function based on the benefit function and the state transition probability function as the MDP decision model.

[0125] In a possible embodiment, the benefit function and the state transition probability can be added to obtain the value function. In a possible embodiment, the long-term state impact caused by taking a defensive action can also be considered in the process of constructing the value function. For example, the long-term impact of the defensive action on the system can be represented by a discount factor. The discount factor determines the weight of the future reward relative to the current reward. The future reward refers to the long-term state impact caused by taking a defensive action, such as the impact of the system security state at the next moment of taking the defensive action, resource usage data, and the cost consumed. Through the discount factor, the construction of the cost function can take into account the long-term impact of the defensive action, which is conducive to selecting the target execution order of the target defensive action with better execution effect over a longer time period. The value range of the discount factor can be [0,1].

[0126] Exemplarily, the above value function can be expressed by the following Bellman formula:

[0127]

[0128] in, represents the optimal value function in state u, that is, the maximum expected return that can be obtained by taking the optimal strategy starting from state u. a is each defensive action in the corresponding defensive action set in state u. r(u,a) is the immediate reward obtained by executing action a in state u. is the probability of transitioning from state u to state u' after executing action a. γ represents the discount factor, which is used to control the weight of future rewards. It represents the weighted sum of all possible next states u' and their corresponding optimal values ​​after executing action a from state u.

[0129] Accordingly, the target execution order of the target defense actions can be solved by maximizing the value function, that is, the target execution order can be solved by the following formula:

[0130]

[0131] As a possible implementation method, the maximum value of the above value function can be solved by an iterative method, for example, Figure 3 As shown, the following steps may be included:

[0132] S1311. Initialize the value function corresponding to the current system security state. Specifically, the value function V(u) may be randomly initialized, for example, the value function may be initialized to 0 or any other arbitrary value.

[0133] S1312. Calculate the value function value corresponding to each first target defense action included in the target defense action set by using the MDP decision model.

[0134] In this step, a target defense action can be randomly selected from the target defense action set to calculate the value function until the value function corresponding to each target defense action is calculated. After the calculation is completed, the system security state at the next moment corresponding to the execution of the first target defense action can be obtained.

[0135] S1313. Calculate the value function value of each remaining defense action corresponding to each first target defense action using the MDP decision model.

[0136] The above-mentioned remaining target defense actions refer to all target defense actions in the target defense action set except the first target defense action. Exemplarily, the value function value can be calculated for each remaining defense action, and the remaining defense action with the largest value function value is selected as the second target defense action corresponding to the first target defense action.

[0137] S1314. Use the second target defense action as the first target defense action, and determine whether the remaining defense actions corresponding to the first target defense action are empty. If not, return to S1313. If yes, obtain multiple execution orders and execute S1315.

[0138] S1315. Based on the sum of the value function values ​​corresponding to the execution sequences of the defense actions, determine the defense action execution sequence with the largest sum of the value function values ​​as the target defense strategy.

[0139] In a possible embodiment, the weighted sum of the value functions corresponding to each execution order can be calculated, and the weight of each value function can be determined according to the position of the target defense action corresponding to the value function in the execution order. For example, it can be determined that the weight of the value function value corresponding to the target defense action with a higher execution position is higher. The specific weight of each value function value can be set according to the actual application scenario, and the present invention does not make specific limitations on this.

[0140] That is to say, in the embodiment of the present invention, the optimal execution strategy can be solved by the following formula: , that is, the execution order of the target defense actions:

[0141] Among them, A(u) is the target defense action set solved in state u, which means that in each state u, the defender should choose an action that maximizes the weighted sum of the immediate reward and the optimal value obtained from the new state.

[0142] By applying the above technical means, the resource consumption of each defense action and the system security status gain are reflected through the value function. Determining the optimal defense strategy based on the value of the value function can achieve the optimal allocation of defender resources, thereby ensuring the security of the system under different situations.

[0143] Furthermore, in the embodiment of the present invention, the optimal defense strategy is dynamically determined based on the current system security status and monitoring data, that is, the defense measures can be adjusted in time according to different states of the system, thereby improving the accuracy of the defense strategy.

[0144] In addition, in the embodiment of the present invention, when calculating the optimal defense strategy, not only the current instant reward is considered, but also the future security is comprehensively considered. In many security issues, the defender must not only prevent current attacks, but also ensure that the system can cope with potential threats in the future. Therefore, the optimal strategy combines the immediate reward and the future reward through the Bellman equation, ensuring the long-term stability and security of the system.

[0145] In a possible embodiment, after the target defense strategy begins to be executed, system monitoring data can be collected in real time, and the target defense strategy can be adjusted in time according to the feedback of the system security status until the system reaches a stable and safe state.

[0146] As a specific implementation, the dynamic Web vulnerability detection method provided by the present invention can be implemented by the following code:

[0147] import java.util.ArrayList;

[0148] import java.util.List;

[0149] import java.util.Optional;

[0150] / **

[0151] * Dynamic Web vulnerability detection method sample code.

[0152] * Combine differential game theory and partial differential equations (HJB equations) for vulnerability detection and defense optimization.

[0153] * /

[0154] public class DynamicWebVulnerabilityDetection {

[0155] public static void main(String[] args) {

[0156] / / A result object contains all code flow information

[0157] Result result = getResultFromSomewhere();

[0158] / / Save the parsed vulnerability call flow object

[0159] List <codesvulncallflowdo>codesVulnCallFlowDos = new ArrayList<>();

[0160] / / Traverse and parse all code flow information

[0161] result.getCodeFlows().stream()

[0162] .filter(Objects::nonNull) / / Filter empty CodeFlows

[0163] .flatMap(codeFlow ->codeFlow.getThreadFlows().stream()) / / Expand thread flow

[0164] .filter(Objects::nonNull) / / Filter empty ThreadFlows

[0165] .flatMap(threadFlow ->threadFlow.getLocations().stream()) / / Expand the location stream

[0166] .filter(Objects::nonNull) / / Filter empty ThreadFlowsLocations

[0167] .forEach(location ->{

[0168] / / Extract the message text of the location

[0169] String objectName = Optional.ofNullable(location.getLocation())

[0170] .map(Location::getMessage)

[0171] .map(Message::getText)

[0172] .orElse(null);

[0173] if (objectName != null) {

[0174] String[] objectNames = objectName.split(":");

[0175] / / Create and set the vulnerability call flow object

[0176] CodesVulnCallFlowDo callFlowDo = new CodesVulnCallFlowDo();

[0177] callFlowDo.setTaskId(getCurrentTaskId()); / / Get the current task ID from the context

[0178] callFlowDo.setCodesVulnUuId(getCurrentVulnUuid()); / / Get the current vulnerability UUID from the context

[0179] callFlowDo.setObjectName(objectNames[0]);

[0180] / / Add to the list

[0181] codesVulnCallFlowDos.add(callFlowDo);

[0182] }

[0183] });

[0184] / / Application of differential game and HJB equation optimization

[0185] applyDifferentialGameTheory();

[0186] applyHJBEquationOptimization();

[0187] }

[0188] / **

[0189] * Simulate obtaining the Result object.

[0190] * @return a Result object containing code flow information.

[0191] * /

[0192] private static Result getResultFromSomewhere() {

[0193] / / Code to get the result object

[0194] return new Result();

[0195] }

[0196] / **

[0197] * Apply differential game theory to optimize defense strategies.

[0198] * /

[0199] private static void applyDifferentialGameTheory() {

[0200] / / Specific implementation of optimized attack and defense strategies

[0201] System.out.println("Applying Differential Game Theory...");

[0202] / / Calculate the optimal control strategy placeholder logic

[0203] }

[0204] / **

[0205] * Apply the HJB equation to optimize the detection strategy.

[0206] * /

[0207] private static void applyHJBEquationOptimization() {

[0208] / / Optimize vulnerability detection strategy based on HJB equation

[0209] System.out.println("Applying HJB Equation Optimization...");

[0210] / / Placeholder logic for numerical solution of HJB equation

[0211] }

[0212] / **

[0213] * Get the current task ID (simulation context).

[0214] * @return The ID of the current task.

[0215] * /

[0216] private static int getCurrentTaskId() {

[0217] return 12345; / / Example task ID

[0218] }

[0219] / **

[0220] * Get the current vulnerability UUID (simulation context).

[0221] * @return The UUID of the current vulnerability.

[0222] * /

[0223] private static String getCurrentVulnUuid() {

[0224] return "vuln-uuid-example"; / / Example vulnerability UUID

[0225] }

[0226] }

[0227] / / The following is the definition of the auxiliary class:

[0228] class Result {

[0229] List <codeflows>getCodeFlows() {

[0230] / / Here is the code to get the CodeFlows list

[0231] return List.of(new CodeFlows());

[0232] }

[0233] }

[0234] class CodeFlows {

[0235] List <threadflows>getThreadFlows() {

[0236] / / Here is the code to get the list of ThreadFlows

[0237] return List.of(new ThreadFlows());

[0238] }

[0239] }

[0240] class ThreadFlows {

[0241] List <threadflowslocations>getLocations() {

[0242] / / Here is the code to get the Locations list

[0243] return List.of(new ThreadFlowsLocations());

[0244] }

[0245] }

[0246] class ThreadFlowsLocations {

[0247] Location getLocation() {

[0248] / / Here is the code to get the Location object

[0249] return new Location();

[0250] }

[0251] }

[0252] class Location {

[0253] Message getMessage() {

[0254] / / Here is the code to get the Message object

[0255] return new Message();

[0256] }

[0257] }

[0258] class Message {

[0259] String getText() {

[0260] / / Here is the code to get the text information

[0261] return "example:objectName";

[0262] }

[0263] }

[0264] class CodesVulnCallFlowDo {

[0265] void setTaskId(int id) {

[0266] / / Set the task ID

[0267] }

[0268] void setCodesVulnUuId(String uuid) {

[0269] / / Set the vulnerability UUID

[0270] }

[0271] void setObjectName(String objectName) {

[0272] / / Set the object name

[0273] }

[0274] }

[0275] For example, for the following Web application scenarios: Attacker strategy: Scan for vulnerabilities through OWASP ZAP and try to exploit SQL injection vulnerabilities. Defender strategy: Scan host vulnerabilities through Nmap, patch known vulnerabilities, and monitor abnormal traffic. State variables: system security status, attack strength, etc.

[0276] The optimal defense strategy can be solved by the following steps:

[0277] 1. Monitor request traffic and resource usage data in real time, analyze abnormal behavior based on traffic data, and detect abnormal resource consumption based on resource usage data.

[0278] 2. Use the partial differential equation dynamic model to solve the target defense action set and the current system security status based on abnormal resource consumption and abnormal behavior.

[0279] 3. Use the MDP decision model to output the optimal defense strategy based on the target defense action set and the current system security status.

[0280] By using the embodiments of the present invention, attack defense game and partial differential equation method can be used to predict attack behaviors, thereby more accurately and timely identifying potential vulnerabilities in Web applications. By using optimal control theory and Markov decision process, the target defense strategy is determined based on the value function value corresponding to the defense action, which can analyze complex attack methods and defense strategies and improve detection coverage.

[0281] Based on the same inventive concept, the embodiment of the present invention also provides a dynamic Web vulnerability detection device, such as Figure 4 As shown, the apparatus 400 may include:

[0282] An acquisition module 401 is used to acquire real-time system monitoring data, wherein the real-time system monitoring data includes resource usage data and access flow data;

[0283] An output module 402 is used to output the current system security state and a target defense action set based on the real-time system monitoring data and a pre-built partial differential equation dynamic model, wherein the partial differential equation dynamic model is used to define the association between the system security state and the attack data and the defense data, and the target defense action set includes at least one defense action;

[0284] A calculation module 403 is used to calculate the target value corresponding to each target defense action based on the current system security state and each target defense action included in the target defense action set by using a pre-built MDP decision model; the MDP decision model is used to define the association relationship between the value function and the defense action and the system security state;

[0285] The determination module 404 is used to determine, based on the target values ​​corresponding to the target defense actions, the target defense action execution order with the highest sum of target values ​​as the target defense strategy.

[0286] The exemplary embodiment of the present invention further provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to perform a method according to an embodiment of the present invention when executed by the at least one processor.

[0287] Exemplary embodiments of the present invention also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present invention.

[0288] An exemplary embodiment of the present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor of a computer, the computer is used to enable the computer to perform a method according to an embodiment of the present invention.

[0289] refer to Figure 5 , a block diagram of an electronic device 500 that can be used as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0290] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM 502 or a computer program loaded from a storage unit 508 into a RAM 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O interface 505 is also connected to the bus 504.

[0291] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 may be any type of device capable of inputting information to the electronic device 500, and the input unit 506 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 507 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 may include, but is not limited to, a disk, an optical disk. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0292] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, any of the above-described dynamic Web vulnerability detection methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 may be configured to perform any of the above-described dynamic Web vulnerability detection methods in any other appropriate manner (e.g., by means of firmware).

[0293] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0294] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0295] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0296] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0297] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0298] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.< / threadflowslocations> < / threadflows> < / codeflows> < / codesvulncallflowdo>

Claims

1. A dynamic Web vulnerability detection method, characterized in that: The method comprises: Acquire real-time system monitoring data, wherein the real-time system monitoring data includes resource usage data and access flow data; Based on the real-time system monitoring data and the pre-built partial differential equation dynamic model, output the current system security state and the target defense action set, wherein the partial differential equation dynamic model is used to define the association between the system security state and the attack data and the defense data, and the target defense action set includes at least one defense action; Calculating the target value corresponding to each target defense action based on the current system security state and each target defense action included in the target defense action set using a pre-built MDP decision model; the MDP decision model is used to define the association between the value function and the defense action and the system security state; Based on the target values ​​corresponding to the target defense actions, determining the target defense action execution order with the highest sum of the target values ​​as the target defense strategy; The partial differential equation dynamic model is pre-built by the following steps: Constructing a partial differential equation, wherein the partial differential equation defines a relationship between a system security state and attack data and defense data, wherein the attack data includes an attack target and an attack method obtained based on a preset attack behavior model, and the defense data includes a defense action taken; Constructing a cost function, wherein the cost function includes the cost corresponding to the adopted defense action and the security assessment result of the system at a preset time point; Constructing a Hamiltonian based on the partial differential equation and the cost function to obtain a dynamic model of the partial differential equation; The using of the pre-built MDP decision model based on the current system security state and each target defense action included in the target defense action set to calculate the target value corresponding to each target defense action includes: Utilizing the MDP decision model to calculate the value function value corresponding to each first target defense action included in the target defense action set; wherein the value function is constructed based on the state reward corresponding to each defense action; the state reward includes resource consumption corresponding to taking the defense action and / or system security state change; Calculating the value function value of each remaining defense action corresponding to each first target defense action by using the MDP decision model, wherein the remaining defense actions are other defense actions in the target defense action set except the first target defense action; Determine the defense action with the highest value function value among the remaining defense actions as the second target defense action corresponding to the first target defense action; The second target defense action is used as the first target defense action, and the step of using the MDP decision model to calculate the value function value of each remaining defense action corresponding to each first target defense action is returned until the remaining defense actions corresponding to each first target defense action are empty, thereby obtaining a plurality of defense action execution orders.

2. The method according to claim 1, characterized in that The partial differential equation is expressed by the following formula: Wherein, u(x,t) represents the safety state of the system at time t and spatial position x, k is the preset conduction coefficient, represents an external control variable, wherein the external control variable at least includes an attack target, an attack method, and a defense action, Indicates the influence of the external control variable on the system safety; The cost function is expressed by the following formula: in, is the cost corresponding to the adopted defensive action, T is the preset time point, as well as is the preset space type range, is the safety assessment result of the system at the preset time point T; The Hamiltonian is expressed by the following formula: Wherein, H is the Hamiltonian, is the accompanying variable.

3. The method according to claim 2, characterized in that The output of the current system security status and target defense action set based on the real-time system monitoring data and the pre-built partial differential equation dynamic model includes: Optimizing the Hamiltonian and determining a target external control variable corresponding to the optimal Hamiltonian; Calculating partial derivatives of the target external control variable based on the defensive action types included in the external control variable, and determining based on each partial derivative result that the defensive actions corresponding to the extreme values ​​of the external control variable constitute a target defensive action set; The Hamiltonian is solved based on the target defense action set and the accompanying variables to obtain the current system safety state.

4. The method according to claim 2, characterized in that: The method of obtaining partial derivatives of the target external control variables based on the defensive action types included in the external control variables, and determining the defensive actions corresponding to the extreme values ​​of the external control variables to form a target defensive action set based on the partial derivative results includes: Taking partial derivatives of the target external control variables based on the attack data types included in the target external control variables, and determining that the attack data corresponding to the external control variable mechanism constitutes an optimal attack strategy set based on each of the partial derivative results; the attack data types at least include an attack target and an attack method; updating the target external control variable based on the optimal attack strategy set; The partial derivatives of the target external control variables are calculated based on the defense action types included in the updated external control variables, and the defense actions corresponding to the extreme values ​​of the external control variables are determined to constitute a target defense action set based on the partial derivative results.

5. The method according to claim 1, characterized in that The MDP decision model is pre-built through the following steps: Based on the state rewards corresponding to each defense action, a profit function is constructed; the state rewards include resource consumption and / or system security state changes corresponding to the defense action; Constructing a state transition probability function based on the state reward corresponding to each of the defensive actions and the current system security state; A value function is constructed based on the benefit function and the state transition probability as the MDP decision model.

6. The method according to claim 5, characterized in that The step of determining the target defense action execution order with the highest sum of target values ​​based on the target values ​​corresponding to the target defense actions as the target defense strategy includes: Based on the sum of the value function values ​​corresponding to the execution sequences of the defense actions, the execution sequence of the defense actions with the largest sum of the value function values ​​is determined as the target defense strategy.

7. A dynamic Web vulnerability detection device, characterized in that: The device comprises: An acquisition module, used to acquire real-time system monitoring data, wherein the real-time system monitoring data includes resource usage data and access flow data; An output module is used to output the current system security state and a target defense action set based on the real-time system monitoring data and a pre-built partial differential equation dynamic model, wherein the partial differential equation dynamic model is used to define the association between the system security state and the attack data and the defense data, and the target defense action set includes at least one defense action; the partial differential equation dynamic model is pre-built by the following steps: Constructing a partial differential equation, wherein the partial differential equation defines a relationship between a system security state and attack data and defense data, wherein the attack data includes an attack target and an attack method obtained based on a preset attack behavior model, and the defense data includes a defense action taken; Constructing a cost function, wherein the cost function includes the cost corresponding to the adopted defense action and the security assessment result of the system at a preset time point; Constructing a Hamiltonian based on the partial differential equation and the cost function to obtain a dynamic model of the partial differential equation; A calculation module, used to calculate the target value corresponding to each target defense action based on the current system security state and each target defense action included in the target defense action set by using a pre-built MDP decision model; the MDP decision model is used to define the association between the value function and the defense action and the system security state; the use of the pre-built MDP decision model based on the current system security state and each target defense action included in the target defense action set to calculate the target value corresponding to each target defense action includes: Utilizing the MDP decision model to calculate the value function value corresponding to each first target defense action included in the target defense action set; wherein the value function is constructed based on the state reward corresponding to each defense action; the state reward includes resource consumption corresponding to taking the defense action and / or system security state change; Calculating the value function value of each remaining defense action corresponding to each first target defense action by using the MDP decision model, wherein the remaining defense actions are other defense actions in the target defense action set except the first target defense action; Determine the defense action with the highest value function value among the remaining defense actions as the second target defense action corresponding to the first target defense action; The second target defense action is used as the first target defense action, and the step of using the MDP decision model to calculate the value function value of each remaining defense action corresponding to each first target defense action is returned until the remaining defense action corresponding to each first target defense action is empty, thereby obtaining a plurality of defense action execution orders; The determination module is used to determine the target defense action execution order with the highest sum of target values ​​as the target defense strategy based on the target values ​​corresponding to the target defense actions.

8. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Markov signal game-based moving target defense strategy selection method and equipment

    CN110460572A

  • WEB dynamic adaptive defense system and defense method based on false response

    CN111917691A