Vehicle hybrid attack safety control method based on zero-sum game and reinforcement learning

By establishing a control system model for hybrid attacks and introducing a state observer, combined with zero-sum game and reinforcement learning methods, the problems of the single defense mechanism and insufficient stability of autonomous driving vehicles facing hybrid attacks are solved, real-time detection and dynamic adjustment are achieved, and the stability and performance of the vehicle are improved.

CN120669525APending Publication Date: 2025-09-19SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510701563.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies lack a comprehensive defense mechanism for hybrid attacks faced by autonomous vehicles. Traditional control algorithms are difficult to ensure stability in unknown or dynamically changing environments, and have not effectively combined game theory and learning technology to optimize attack and defense strategies.

Method used

A mathematical model of the control system including hybrid attacks is established, the state observer and augmented system are introduced, the cost function is defined, the Bellman equation is derived using zero-sum game and reinforcement learning methods, and the optimal control strategy and attack strategy are obtained through the critic and actor neural network.

Benefits of technology

It achieves real-time attack detection and dynamic adjustment of control strategies, improves the robustness and stability of the system, optimizes the vehicle's driving stability, path tracking accuracy and energy efficiency, and reduces the risk of system performance degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669525A_ABST
    Figure CN120669525A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle hybrid attack safety control method based on zero-sum game and reinforcement learning. The method comprises the following steps: establishing a control system mathematical model containing hybrid attacks; introducing a state vector and an external influence vector of a control system containing hybrid attacks, and constructing an augmented system; defining a cost function; deriving a Bellman equation of a zero-sum game problem and establishing an optimization problem of the zero-sum game by utilizing an optimality principle; and constructing an HJI equation based on the optimization problem of the zero-sum game, solving the HJI equation, and obtaining the optimal solutions of the control strategy and the attack strategy through the commentator neural network and the actor neural network. According to the method, the state observer, the zero-sum game framework and the reinforcement learning method are combined to form a mixed learning control strategy, the influence of attacks can be more effectively inhibited, the risk of system performance reduction is reduced, and a brand new thought is provided for solving the control problem when the automatic driving vehicle is subjected to cooperative attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security control technology, and in particular to a vehicle hybrid attack security control method based on zero-sum game and reinforcement learning. Background Art

[0002] AGVs may face two main types of cyberattacks during operation: false data injection (FDI) attacks and denial of service (DoS) attacks. FDI attacks disrupt the system's normal operation and decision-making process by injecting false data into the system, while DoS attacks prevent legitimate users from accessing or using the target system by exhausting its resources. These attacks can cause AGVs to misidentify road conditions, obstacles, and other vehicles, leading to dangerous driving behavior or accidents.

[0003] At present, the methods used for attack control include control methods based on event triggering mechanism, control methods based on adaptive dynamic programming (ADP), control methods based on zero-sum game, and control methods based on attack detection and defense mechanism.

[0004] However, in real-world scenarios, attackers may launch multiple types of attacks simultaneously (such as hybrid attacks). Existing methods have the following shortcomings: a lack of comprehensive defense mechanisms against hybrid attacks; traditional control algorithms are difficult to ensure stability in environments with unknown or dynamically changing models; and there is a failure to effectively combine game theory and learning techniques to optimize attack and defense strategies.

[0005] Therefore, there is an urgent need for a solution that can detect attacks in real time, dynamically adjust control strategies and ensure system robustness. Summary of the Invention

[0006] In response to the above-mentioned deficiencies in the prior art, the vehicle hybrid attack security control method based on zero-sum game and reinforcement learning provided by the present invention solves the problems of the single defense mechanism and insufficient stability in the prior art.

[0007] In order to achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a vehicle hybrid attack security control method based on zero-sum game and reinforcement learning, comprising:

[0008] Establish a mathematical model of the control system containing hybrid attacks; introduce the state vector and external influence vector of the control system containing hybrid attacks and construct an augmented system;

[0009] Define the cost function;

[0010] Based on the augmented system and cost function, using the optimality principle, the Bellman equation of the zero-sum game problem is derived and the optimization problem of the zero-sum game is established;

[0011] Based on the optimization problem of zero-sum game, the HJI equation is constructed and solved, and the optimal solutions of control strategy and attack strategy are obtained through critic neural network and actor neural network.

[0012] The beneficial effects of the present invention are:

[0013] 1. By combining a state observer, a zero-sum game framework, and reinforcement learning methods, a hybrid learning control strategy is developed, offering a novel approach to solving the control problem of autonomous vehicles under coordinated attacks. The state observer monitors the system state in real time, enabling timely detection of attacks and triggering appropriate defense mechanisms. The zero-sum game framework models the adversarial relationship between attackers and defenders, enabling the controller to optimize defense strategies in complex attack environments. Reinforcement learning, through continuous trial-and-error learning, enables the system to gradually approach the optimal control strategy in unknown environments, effectively addressing various uncertainties and dynamic changes.

[0014] 2. When subjected to a coordinated attack, this invention promptly detects attack signals and, by triggering a zero-sum game mechanism, rapidly forces the system into a defensive state. In this state, the controller treats the attacker as an adversary in the game and optimizes control strategies to minimize the damage inflicted by the attack on the system. Compared to traditional single-pronged defense methods, this game-theory-based active defense approach can more effectively suppress the impact of attacks and reduce the risk of system performance degradation.

[0015] 3. By using reinforcement learning methods to optimize the system online, the present invention enables the controller to gradually learn the optimal control strategy in a continuous trial and error process, thereby achieving continuous improvement in system performance. In actual operation, this is reflected in many aspects such as vehicle driving stability, path tracking accuracy, and energy efficiency. For example, in path tracking tasks, the optimized control strategy can enable the vehicle to travel along the preset path more accurately, reducing deviations and oscillations; in terms of energy consumption management, a reasonable control strategy helps to reduce the vehicle's energy consumption and increase cruising range. The optimization of these performances not only improves the vehicle's user experience, but also lays a solid foundation for the large-scale application of autonomous driving technology.

[0016] 4. This invention is applicable to a wide range of autonomous vehicles and operating conditions. Whether operating on urban roads, highways, or in challenging terrain, the control strategy proposed in this invention can rapidly adapt to new environments and vehicle characteristics through appropriate parameter adjustments and neural network training, providing stable and reliable control support. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the process of the present invention;

[0018] Figure 2 This is a schematic diagram of the effect of using this method to control the attack after being attacked;

[0019] Figure 3 This is a schematic diagram of the effect of using common methods to control the attack after being attacked. DETAILED DESCRIPTION

[0020] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0021] like Figure 1 As shown, in one embodiment of the present invention, a vehicle hybrid attack security control method based on zero-sum game and reinforcement learning includes the following steps:

[0022] S1. Establish a mathematical model of the control system that includes hybrid attacks. AGVs may face two main network attacks during operation: false data injection (FDI) attacks and denial of service (DoS) attacks. FDI attacks interfere with the normal operation and decision-making process of the system by injecting false data into the system, while DoS attacks prevent legitimate users from accessing or using the target system by exhausting its resources. These attacks may cause AGVs to misidentify road conditions, obstacles, and other vehicles, leading to dangerous driving behavior or accidents. In order to consider the combined interference of the two attacks, the specific attack model is designed in this embodiment as follows:

[0023] u(k)=α(k)u d (k)

[0024] y(k)=y c (k)+f a (k)

[0025] Among them, u(k) represents the control input under the DoS attack at time k, α(k) is the Bernoulli random variable, and u d (k) is the expected control input at time k; y(k) represents the control output under FDI attack, y c (k) represents the controlled output under normal network conditions, f a (k) is the FDI attack signal;

[0026] The expression of the control system containing hybrid attacks is:

[0027] x(k+1)=A xx(k)+α(k)B u u d (k)+E w w(k)

[0028] y(k)=y c (k)+f a (k)

[0029] Among them, x(k+1) represents the real state of the control system at time k+1, and x(k) represents the real state of the control system at time k; A x 、B u 、E w are all constant matrices; w(k) represents the external disturbance at time k.

[0030]

[0031] Before security control is implemented, the system detects attacks based on a mathematical model of the control system that includes hybrid attacks. When faced with coordinated attacks such as DoS and FDI attacks, the state observer promptly detects the presence of attacks and rapidly adjusts the control strategy to mitigate their impact on system performance. In the absence of attacks, the system uses a conventional output feedback controller to ensure stability. Upon detecting an attack, it quickly switches to a hybrid learning controller based on zero-sum game theory. This flexible control mode switching mechanism ensures stable operation under various attack scenarios.

[0032] The specific methods for detecting attacks are:

[0033] Design a state observer:

[0034]

[0035] in, is the state observation value at time k+1, is the state observation value at time k, L is the observer gain, is the state observation output at time k, and F is a constant matrix;

[0036]

[0037] Design residual evaluation function Γ resi (k):

[0038]

[0039] Among them, γ r (k) is the residual signal, Ω r is a positive definite weight matrix; the superscript T represents the transpose of the matrix;

[0040] Compare the relationship between the residual function value and the residual function threshold Γ, if Γ resii (k)≤Γ, then the detection result is no attack; if Γ resi If (k)>Γ, then there is an attack.

[0041] The state observer uses ordinary feedback control when observing. The feedback control is described as:

[0042] u d (k)=Ky(k)

[0043] Where K is measured by the quadratic performance function Determine the feedback control gain; R 11 and R 12 are all positive definite matrices.

[0044] S2, introduce the state vector and external influence vector of the control system containing hybrid attacks and construct an augmented system;

[0045] The expressions of the state vector and external influence vector of the control system are:

[0046]

[0047]

[0048] in, is the state vector, which represents the state of the control system at time k; ^ represents the estimated value of the parameter; Δx(k) is the state observation error at time k; is the external influence vector;

[0049] The expression of the augmented system is:

[0050]

[0051] in,

[0052]

[0053] S3. Define the cost function;

[0054] The cost function is Its expression is:

[0055]

[0056]

[0057] in, Indicates solving the expected value; Represents the discount factor, R1=diag[R 11, O], diag[·] is the diag function; O, R2 and R3 are all positive definite matrices;

[0058] S4. Using the optimality principle, derive the Bellman equation for the zero-sum game problem and establish the optimization problem of the zero-sum game;

[0059] The Bellman equation for the zero-sum game problem is:

[0060]

[0061] in, represents the value function at time k, represents the value function at time k+1;

[0062] The optimization problem of zero-sum game is expressed as:

[0063]

[0064] in, represents the optimal value of the control input expected at time k; represents the optimal value of external influence.

[0065] The HJI equation is established based on the optimization problem of zero-sum game, and its expression is:

[0066]

[0067] The expressions of the optimal control strategy and the worst attack strategy are:

[0068]

[0069] in, is the optimal control strategy, For the worst attack strategy,

[0070] S5. Based on the optimization problem of zero-sum game, the HJI equation is constructed and solved, and the optimal solutions of control strategy and attack strategy are obtained through critic neural network and actor neural network using reinforcement learning method.

[0071] The critic neural network is used to approximate the cost function, which is expressed as:

[0072]

[0073] The actor neural network is used to approximate the optimal control strategy and the worst attack strategy, which are expressed as:

[0074]

[0075]

[0076] in, express The optimal control strategy under express The worst attack strategy under are the ideal weights of the critic neural network, is the activation function vector, represents the approximation error of the critic neural network; and represents the weights of the actor's neural network, and is the activation function vector, and is the approximate error.

[0077] The specific method of using the critic neural network to approximate the cost function is:

[0078] Define the estimated cost function

[0079]

[0080] in, is the estimated weight of the critic neural network;

[0081] The TD error is introduced to update the estimated weights of the critic neural network to ensure the stability and robustness of the system. The TD error is expressed as

[0082]

[0083] in,

[0084] The estimated weights of the critic neural network after update are

[0085]

[0086] in, is a constant matrix; η v is the learning rate of the critic neural network;

[0087] The cost function is approximated using the critic neural network after updating the estimated weights.

[0088] The specific method of using actor neural network to approximate the optimal control strategy and the worst attack strategy is:

[0089] Define the estimated optimal control strategy and estimate the worst attack strategy They are:

[0090]

[0091] in, and are the estimated weights of the actor’s neural network;

[0092] Update the estimated weights of the actor's neural network; the estimated weights of the actor's neural network after update are:

[0093]

[0094] in, and is the estimated weight of the actor’s neural network after the update; are constant matrices, η u 、 denote the learning rates of the two actor neural networks respectively; All are optimization errors; K u represents the control gain; Both represent reconstruction errors;

[0095] The optimal control strategy and the worst attack strategy are approximated using the actor neural network with updated weights.

[0096] In order to verify the beneficial effects of the present invention, the method proposed by the present invention and the conventional method were used to perform attack control under the same experimental conditions. Figure 2 、 Figure 3 As shown in the figure, the attack control method proposed by the present invention can more effectively suppress the impact of attacks, reduce the risk of system performance degradation, and significantly improve the vehicle's survivability and mission completion capabilities in harsh network environments compared to traditional single defense methods.

[0097] In summary, the present invention optimizes the system through reinforcement learning, so that the controller gradually learns the optimal control strategy in the process of continuous trial and error, thereby achieving a stable improvement in system performance. In actual operation, this is reflected in many aspects such as vehicle driving stability, path tracking accuracy, and energy efficiency. For example, in path tracking tasks, the optimized control strategy can enable the vehicle to travel along the preset path more accurately, reducing deviations and oscillations; in terms of energy consumption management, a reasonable control strategy helps to reduce the vehicle's energy consumption and increase cruising range. The optimization of these performances not only improves the user experience of the vehicle, but also lays a solid foundation for the large-scale application of autonomous driving technology.

Claims

1. A vehicle hybrid attack security control method based on zero-sum game and reinforcement learning, characterized in that: include: Establish a mathematical model of control systems including hybrid attacks; The state vector and external influence vector of the control system containing hybrid attacks are introduced and an augmented system is constructed; Define the cost function; Based on the augmented system and cost function, using the optimality principle, the Bellman equation of the zero-sum game problem is derived and the optimization problem of the zero-sum game is established; Based on the optimization problem of zero-sum game, the HJI equation is constructed and solved, and the optimal solutions of control strategy and attack strategy are obtained through critic neural network and actor neural network.

2. The method according to claim 1, characterized in that Hybrid attacks include FDI attacks and DoS attacks. The specific attack model is as follows: u(k)=α(k)u d (k) y(k)=y c (k)+f a (k) Among them, u(k) represents the control input under the DoS attack at time k, α(k) is the Bernoulli random variable, and u d (k) is the expected control input at time k; y(k) represents the control output under FDI attack, y c (k) represents the controlled output under normal network conditions, f a (k) is the FDI attack signal; The expression of the control system containing hybrid attacks is: x(k+1)=A x x(k)+α(k)B u u d (k)+E w w(k) y(k)=y c (k)+f a (k) Among them, x(k+1) represents the real state of the control system at time k+1, and x(k) represents the real state of the control system at time k; A x 、B u 、E w are all constant matrices; w(k) represents the external disturbance at time k.

3. The method according to claim 2, characterized in that Before security control, the control system is detected based on the mathematical model of the control system including hybrid attacks to determine whether there is an attack. The specific method is as follows: Design a state observer: in, is the state observation value at time k+1, is the state observation value at time k, L is the observer gain, is the state observation output at time k, and F is a constant matrix; Design residual evaluation function Γ resi (k): Among them, γ r (k) is the residual signal, Ω r is a positive definite weight matrix; the superscript T represents the transpose of the matrix; Compare the relationship between the residual function value and the residual function threshold Γ, if Γ resii (k)≤Γ, then the detection result is no attack; if Γ resi If (k)>Γ, then there is an attack.

4. The method according to claim 3, characterized in that The expressions of the state vector and external influence vector of the control system are: in, is the state vector, which represents the state of the control system at time k; ^ represents the estimated value of the parameter; Δx(k) is the state observation error at time k; is the external influence vector; The expression of the augmented system is: in, 5. The method according to claim 4, characterized in that The cost function is Its expression is: in, Indicates solving the expected value; Represents the discount factor, R1=diag[R 11 , O], diag[·] is the diag function; O, R2 and R3 are all positive definite matrices; 6. The method according to claim 5, characterized in that The Bellman equation for the zero-sum game problem is: in, represents the value function at time k, represents the value function at time k+1; The optimization problem of zero-sum game is expressed as: in, represents the optimal value of the control input expected at time k; represents the optimal value of external influence.

7. The method according to claim 6, characterized in that The HJI equation is: The expressions of the optimal control strategy and the worst attack strategy are: in, is the optimal control strategy, For the worst attack strategy, 8. The method according to claim 7, characterized in that The critic neural network is used to approximate the cost function, which is expressed as: The actor neural network is used to approximate the optimal control strategy and the worst attack strategy, which are expressed as: in, express The optimal control strategy under express The worst attack strategy under are the ideal weights of the critic neural network, is the activation function vector, represents the approximation error of the critic neural network; and represents the weights of the actor's neural network, and is the activation function vector, and is the approximate error.

9. The method according to claim 8, characterized in that The specific method of using the critic neural network to approximate the cost function is: Define the estimated cost function in, is the estimated weight of the critic neural network; The TD error is introduced to update the estimated weights of the critic neural network. The TD error is expressed as in, The estimated weights of the critic neural network after update are in, is a constant matrix; η v is the learning rate of the critic neural network; The cost function is approximated using the critic neural network after updating the estimated weights.

10. The method according to claim 9, characterized in that The specific method of using actor neural network to approximate the optimal control strategy and the worst attack strategy is: Define the estimated optimal control strategy and estimate the worst attack strategy They are: in, and are the estimated weights of the actor’s neural network; Update the estimated weights of the actor's neural network; the estimated weights of the actor's neural network after update are: in, and is the estimated weight of the actor’s neural network after the update; are constant matrices, η u 、 denote the learning rates of the two actor neural networks respectively; All are optimization errors; K u represents the control gain; Both represent reconstruction errors; The optimal control strategy and the worst attack strategy are approximated using the actor neural network with updated weights.