Attack strategy determination method and device, equipment and storage medium

Through a deep reinforcement learning algorithm based on a near-end strategy optimization algorithm and a maximum entropy reinforcement learning framework, the target network model is trained, and the problem of hidden attack design of nonlinear autonomous driving vehicles is solved, and more accurate and concealed attack strategies are achieved, which improves vehicle safety.

CN120389878APending Publication Date: 2025-07-29SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340695.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The research on the design of covert attacks for nonlinear and safety-critical autonomous driving vehicles in the prior art is limited, especially the lack of effective methods for false data injection attacks against actuators.

Method used

A deep reinforcement learning algorithm based on a near-end strategy optimization algorithm and a maximum entropy reinforcement learning framework is adopted. By obtaining the current state of the vehicle's driving trajectory, the target network model is trained, the attack execution actions against autonomous driving vehicles are determined, and a hidden attack strategy is designed.

Benefits of technology

It improves the accuracy and concealment of attack strategies, enhances the safety and reliability of autonomous vehicles, and reduces the probability of being detected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120389878A_ABST
    Figure CN120389878A_ABST
Patent Text Reader

Abstract

The invention provides an attack strategy determination method and device, equipment and a storage medium. The method comprises the steps of obtaining a current state of a driving track of a target vehicle; based on the current state and a pre-established target network model, determining an execution action for attacking the target vehicle, the target network model being a network model based on a near-end strategy optimization algorithm or a network model based on a deep reinforcement learning algorithm of a maximum entropy reinforcement learning framework, attack design of the autonomous vehicle can be realized through a network model of a near-end strategy optimization algorithm or a network model of a deep reinforcement learning algorithm based on a maximum entropy reinforcement learning framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of autonomous driving technology, and particularly relates to a method, apparatus, device, and storage medium for determining an attack strategy. Background Art

[0002] With the rapid development of autonomous driving technology, the safety and reliability of vehicles have become the focus of research. Autonomous vehicles rely on a variety of sensors and algorithms to perceive the environment and make decisions. However, these systems also face potential cyber attack threats. To deal with cyber attacks, we can design attack experiments on autonomous vehicles to improve autonomous driving technology to detect cyber attacks. When designing attacks on general safety-critical cyber-physical systems, in related technologies, the focus is usually on several types of attacks, including Denial of Service (DoS) attacks, False Data Injection (FDI) attacks, and replay attacks, etc. Usually, these types of attacks are limited to linear systems, and the design of stealth attacks against non-linear, safety-critical autonomous vehicles is still limited. Summary of the Invention

[0003] In view of the above problems, the embodiments of this application provide a method, apparatus, device, and storage medium for determining an attack strategy, which can design an attack strategy against an autonomous vehicle through a target network model.

[0004] In a first aspect, the embodiments of this application provide a method for determining an attack strategy, including:

[0005] Obtain the current state of the driving trajectory of the target vehicle;

[0006] Based on the current state and a pre-established target network model, determine an execution action for attacking the target vehicle, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

[0007] In some embodiments, the method further includes:

[0008] Obtain the sample current state and sample next state of the driving trajectory of the vehicle, the sample execution action of the vehicle, and the sample residual;

[0009] Based on the sample current state and the sample next state, the sample execution action and the sample residual, determine training data;

[0010] Train the target network model based on the training data.

[0011] In some embodiments, obtaining the sample current state and sample next state of the driving trajectory, the sample execution action of the vehicle, and the sample residual includes:

[0012] Inputting the sample current state and the sample execution action into a pre-established first mapping model to obtain the sample next state.

[0013] In some embodiments, obtaining the sample current state and sample next state of the driving trajectory, the sample execution action of the vehicle, and the sample residual includes:

[0014] Inputting the sample current state and the sample execution action into a pre-established second mapping model to obtain the sample residual.

[0015] In some embodiments, the network model of the proximal policy optimization algorithm includes: a policy network, a value network, and an advantage network. The reward function of the policy network is: The value function of the value network is: The advantage function of the advantage network is A π (s,a) = Q π (s,a) - V π (s), where The loss function of the network model of the proximal policy optimization algorithm includes: a surrogate objective loss, a value function loss, and an entropy reward loss, or, the loss function of the neural network model includes: a surrogate objective loss, a value function loss, an entropy reward loss, and an auxiliary task loss. The auxiliary task loss includes: the prediction loss of the first mapping model and / or the prediction loss of the second mapping model, where is a three-dimensional vector, x k represents the sample current state at time step k, is a two-dimensional vector, d k represents the sample execution action, x k represents the x coordinate at time step k, y k represents the y coordinate at time step k, θ k represents the direction angle, v d represents the speed of the attack input, φ d represents the angle of the attack input, r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, x r,k is the reference state, Q π (s,a) represents the expected cumulative reward obtained by taking action a = d k in state s = x k and following policy π, γ is the discount factor, R t is the reward at time step t, E is the expected value, Vπ (s) represents the expected value of future cumulative rewards after state s starts following policy π.

[0016] In some embodiments, the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: an actor network and a critic network, and the reward function of the actor network is: The value function of the critic network is: The loss function of the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: the loss for optimizing the actor network, the loss for optimizing the critic network, and the temperature loss for adjusting the entropy coefficient, or, the loss for the critic network, the loss for the actor network, the temperature loss for adjusting the entropy coefficient, and the auxiliary task loss, where the auxiliary task loss includes: the prediction loss of the first mapping model and / or the prediction loss of the second mapping model, where is a three-dimensional vector, x k represents the sample current state at time step k, is a two-dimensional vector, d k represents the sample execution action, x k represents the x coordinate at time step k, y k represents the y coordinate at time step k, θ k represents the direction angle, v d represents the speed of the attack input, φ d represents the angle of the attack input, r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, x r,k is the reference state, Q π (s,a) represents the expected cumulative reward for taking action a = d k in state s = x k , and following policy π, γ is the discount factor, R t is the reward at time step t, is the expected value, V π (s) represents the expected value of future cumulative rewards after state s starts following policy π.

[0017] In some embodiments, determining the execution action for attacking the target vehicle based on the current state and a pre-established target network model includes:

[0018] Determining the amount of data of the acquired current state;

[0019] Based on the amount of data, determining a target network model from the network model of the proximal policy optimization algorithm and the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework;

[0020] Input the current state into the target network model to output an execution action for attacking the target vehicle.

[0021] In a second aspect, an apparatus for determining an attack strategy provided by an embodiment of the present application includes:

[0022] An acquisition module, configured to acquire the current state of the driving trajectory of the target vehicle;

[0023] A determination module, configured to determine an execution action for attacking the target vehicle based on the current state and a pre-established target network model, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

[0024] In a third aspect, an electronic device provided by an embodiment of the present application includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method provided in the first aspect is implemented.

[0025] In a fourth aspect, a computer-readable storage medium provided by an embodiment of the present application stores a computer program. When the computer program is executed by a processor, the method provided in the first aspect is implemented.

[0026] In a fifth aspect, a computer program product provided by an embodiment of the present application includes a computer program. When the computer program is executed by a processor, it is at least used to implement the method in any one of the first aspect.

[0027] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0028] The method for determining an attack strategy provided by the embodiments of the present application acquires the current state of the driving trajectory of the target vehicle; based on the current state and a pre-established target network model, determines an execution action for attacking the target vehicle, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework, and can implement the attack design on an autonomous driving vehicle through the network model based on the proximal policy optimization algorithm or the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

[0029] It can be understood that the beneficial effects of the second to fifth aspects above can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of the Drawings

[0030] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0031] Figure 1 Schematic flowchart of a method for determining an attack strategy provided by an embodiment of the present application;

[0032] Figure 2 Schematic diagram of a network model of a proximal policy optimization algorithm provided by an embodiment of the present application;

[0033] Figure 3 Schematic structural diagram of a network model of a proximal policy optimization algorithm with an auxiliary task introduced provided by an embodiment of the present application;

[0034] Figure 4 Schematic structural diagram of a network model of a deep reinforcement learning algorithm based on a maximum entropy reinforcement learning framework provided by an embodiment of the present application;

[0035] Figure 5 Schematic structural diagram of a network model of a deep reinforcement learning algorithm based on a maximum entropy reinforcement learning framework with an auxiliary task introduced provided by an embodiment of the present application;

[0036] Figure 6 Schematic structural diagram of a device for determining an attack strategy provided by an embodiment of the present application;

[0037] Figure 7 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0038] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0039] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0040] It should also be understood that the term "and / or" as used in the specification and claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0041] As used in the specification and claims of this application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if detected" can be interpreted as meaning "once determined" or "in response to determining" or "once detected" or "in response to detecting" depending on the context.

[0042] In addition, in the description of the specification and claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0043] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0044] Before introducing the embodiments of this application, a brief introduction to the related art is given:

[0045] With the emergence of autonomous vehicles equipped with advanced sensors and precise actuators, their application scenarios have expanded to a wide range of fields. These vehicles are gradually transitioning from restricted areas to urban open environments and are being increasingly applied in daily life. In recent years, people have paid increasing attention to the safety issues of these vehicles. Since autonomous vehicles highly rely on software and communication systems, they are vulnerable to various types of attacks, which may lead to serious accidents.

[0046] Previous studies usually focused on several types of attacks, including denial-of-service attacks, false data injection attacks, and replay attacks. A denial-of-service attack aims to exhaust system resources so that it cannot respond to legitimate requests. False data injection attacks involve maliciously tampering with data inputs to mislead the system's decision-making. Replay attacks record and then replay legitimate signals to induce the system to accept outdated or incorrect information. False data injection attacks are more common and easier to implement than replay attacks, and are more deceptive than denial-of-service attacks. In addition, attacks on autonomous vehicles are mainly divided into attacks on actuators and attacks on sensors. Compared with sensor attacks, actuator attacks are more difficult to detect because they do not directly affect the observations. This makes actuator attacks particularly dangerous because they can manipulate vehicle behavior without triggering typical detection mechanisms, posing a significant threat to the safety and reliability of autonomous vehicles. Therefore, there is no relevant technology publicly available on how to design false data injection attacks against actuators.

[0047] In addition, when designing attacks against general safety-critical information physical systems, researchers first proposed a closed-form linear attack strategy for discrete linear time-invariant systems and proved the optimality of this attack method. Subsequently, some scholars proposed an optimal injection strategy in which all actuators were compromised by the attacker and analyzed the dynamic response of the system under the optimal switched data injection attack. Then, the researchers designed an optimal hybrid data injection attack strategy that combines false data and its derivatives to degrade system performance. Although designing optimal FDI attacks is important, classical and advanced attack detectors can easily detect them and provide countermeasures. Therefore, how to design stealthy worst-case attacks has attracted the attention of researchers. Some scholars have studied an innovation-based stealthy attack strategy that uses the Kullback-Leibler divergence to measure the stealthiness of the attack and derived the worst-case attack strategy. A recent work considered strict stealth and ∈-stealth attack models, derived two types of conditions, proposed a method for strict stealth attacks, and showed that ∈-stealth attacks do not exist under certain conditions. However, although some studies have considered high-order and large-scale systems, most works, including the above studies, are limited to linear systems, and the research on the design of stealthy attacks against non-linear, safety-critical autonomous vehicles is still limited.

[0048] Based on the technical problems of related technologies, the embodiments of the present application provide a method for determining an attack strategy that can be applied to an electronic device. The electronic device may include: mobile phones, tablet computers, computer devices, etc., and the embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0049] The embodiments of the present application provide a method for determining an attack strategy, Figure 1 The schematic flow chart of a method for determining an attack strategy provided by the embodiments of the present application is asFigure 1 As shown, the method includes:

[0050] Step S101, obtaining the current state of the driving trajectory of the target vehicle.

[0051] In the embodiments of the present application, the target vehicle refers to an autonomous vehicle or an intelligent vehicle, and its driving trajectory and state can be monitored and analyzed in real time through sensors and algorithms. The target vehicle can be a real vehicle, a simulation vehicle, etc.

[0052] In the embodiments of the present application, the current state of the driving trajectory refers to state information such as the position, speed, and direction angle of the target vehicle at a certain moment, which is usually obtained through sensors (such as cameras, radars, etc.).

[0053] In the embodiments of the present application, the electronic device can obtain the current state of the driving trajectory of the target vehicle through sensors. In some embodiments, the user can also input the current state of the driving trajectory of the target vehicle.

[0054] Step S102, based on the current state and a pre-established target network model, determining an execution action for attacking the target vehicle, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

[0055] In the embodiments of the present application, the target network model refers to a neural network model used to determine the attack execution action, which can be a deep reinforcement learning algorithm based on the proximal policy optimization algorithm (PPO, Proximal Policy Optimization) or the maximum entropy reinforcement learning framework soft actor-critic (SAC, Soft Actor-Critic).

[0056] In the embodiments of the present application, the proximal policy optimization algorithm (PPO) is a reinforcement learning algorithm that ensures that the policy does not change drastically during the optimization process by restricting the amplitude of the policy update, thereby improving the stability and efficiency of training. The maximum entropy reinforcement learning framework (SAC) is a reinforcement learning algorithm that encourages exploration of different actions by maximizing the entropy of the policy, thereby improving the robustness and adaptability of the policy.

[0057] In the embodiments of the present application, during the training process, a large amount of vehicle driving trajectory data can be collected, including sample current states, sample next states, sample execution actions, and sample residuals. The target network model can be trained through these data.

[0058] In the embodiments of the present application, the execution action for the attack can be to obtain the false data to be injected, and the execution action can include: the speed of the attack input, the angle of the attack input, etc. The execution action for the attack can be referred to as an attack strategy.

[0059] In the embodiments of the present application, a pre-trained target network model (based on the PPO or SAC algorithm) can be used to calculate the optimal execution action of the attack according to the current state.

[0060] The method for determining the attack strategy provided by the embodiments of the present application includes obtaining the current state of the driving trajectory of the target vehicle; determining the execution action for attacking the target vehicle based on the current state and a pre-established target network model, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework, and can realize the attack design of the autonomous vehicle through the network model based on the proximal policy optimization algorithm or the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

[0061] In some embodiments, before step S102, the method further includes:

[0062] Step S1, obtaining the sample current state and sample next state of the driving trajectory of the vehicle, the sample execution action of the vehicle, and the sample residual.

[0063] In the embodiments of the present application, the sample current state of the driving trajectory of the vehicle refers to the state information of the vehicle at a certain moment, usually including the position of the vehicle (such as longitude and latitude coordinates), speed, direction (such as heading angle), etc. These data are collected in real time by sensors. The sample next state of the driving trajectory of the vehicle refers to the state information of the vehicle at the next moment, usually calculated or predicted through the current state and the execution action. For example, according to the current speed and direction, the position of the vehicle at the next moment is predicted. The sample execution action of the vehicle refers to the specific operation executed by the vehicle at a certain moment, usually including speed, angle, etc. The sample residual refers to the difference between the sensor observation value and the model prediction value. For example, the Extended Kalman Filter (EKF) generates residuals for detecting whether there are abnormalities or attacks in the system.

[0064] In the embodiments of the present application, the sample current state and sample next state of the driving trajectory of the vehicle, the sample execution action of the vehicle, and the sample residual can be generated through the actual driving data or simulation data of the vehicle. For example, different driving scenarios of the vehicle can be simulated in the simulation environment through simulation to generate a large amount of sample data.

[0065] In the embodiments of the present application, in the context of autonomous vehicles, collecting a large amount of attack data is both challenging and costly. Multiple auxiliary tasks are designed to improve the training speed, provide important guidance during the training process, and enhance the effectiveness of the attack strategy. The sample next state and / or the sample residual can be predicted through the service task.

[0066] In the embodiments of the present application, the current state of the sample and the action of the sample are input into a pre-established first mapping model to obtain the next state of the sample.

[0067] In the embodiments of the present application, the first mapping model can be expressed as [x k , d k → x' k+1 , where d k represents the action of the sample, x k is the current state of the sample, and x' k+1 is the next state of the sample.

[0068] In the embodiments of the present application, the first mapping model can simulate the behavior when under attack, so that when training the target network model, without the need for a large number of interactions with the environment, by predicting the next sample state, the attacker can better understand the dynamic behavior of the vehicle, thereby designing a more effective attack strategy.

[0069] In the embodiments of the present application, the mean squared error (MSE) loss function is used to compare the loss between the predicted next state x' k+1 and the true next state x k+1 . The loss function of the first mapping model can be defined as:

[0070]

[0071] where M is the number of samples in the batch.

[0072] In the embodiments of the present application, the first mapping network is a neural network for predicting the next state from the current state and attack data. This network approximately simulates the behavior of the vehicle under attack by learning the state update law of the vehicle.

[0073] The method provided in the embodiments of the present application can predict the next sample state through the first mapping model, and can learn state transition without the need for a large number of interactions with the environment, improving the training efficiency.

[0074] In some embodiments, obtaining the current state of the sample of the driving trajectory, the next state of the sample, the action of the sample of the vehicle, and the sample residual includes:

[0075] Inputting the current state of the sample and the action of the sample into a pre-established second mapping model to obtain the sample residual.

[0076] In the embodiments of the present application, the second mapping model can be expressed as: where is the residual.

[0077] In the embodiments of the present application, by minimizing the change in the residual to avoid detection, and by predicting the EKF residual as accurately as possible, the attacker can better understand the system's response to the attack, thereby designing a more covert attack strategy.

[0078] In the embodiments of the present application, the mean squared error (MSE) loss function can be used to compare the predicted EKF residual and the actual EKF residual r k for the loss therebetween.

[0079] In the embodiments of the present application, the loss function of the second mapping model is:

[0080] M is the number of samples in the batch. The second mapping network Nnaux is a neural network for predicting the EKF residual from the current state and the attack data (executed actions). This network approximates the system's response to the attack by learning the residual generation process of the system.

[0081] In the embodiments of the present application, by accurately predicting the EKF residual, the attacker can generate an attack signal that minimizes the change in the residual, thereby avoiding being detected. By learning the residual generation process of the system, the attacker can better understand the system's response to the attack, thereby designing a more effective attack strategy. By minimizing the change in the residual, the attacker can reduce the probability of the system detecting the attack, thereby increasing the success rate of the attack.

[0082] Step S2: Determine training data based on the current state of the sample, the next state of the sample, the executed action of the sample, and the residual of the sample.

[0083] In the embodiments of the present application, the training data includes the current state of the sample, the next state of the sample, the executed action of the sample, and the residual of the sample. The training data can be expressed as: (x k , x k+1 , d k , r k ), where: x k is the current state of the sample, x k+1 is the next state of the sample, d k is the executed action of the sample, and r k is the residual of the sample.

[0084] Step S3: Train the target network model based on the training data.

[0085] In the embodiments of the present application, the training data can be used to train a neural network model to optimize the parameters of the model, thereby obtaining the target network model.

[0086] In the embodiments of the present application, the target network model can be a model based on a reinforcement learning algorithm (such as PPO or SAC). The PPO model ensures that the policy does not change drastically during the optimization process by restricting the amplitude of policy updates, thereby improving the stability and efficiency of training. The SAC model encourages exploration of different actions by maximizing the entropy of the policy, thereby improving the robustness and adaptability of the policy.

[0087] The method provided in the embodiments of the present application can improve the accuracy of the attack policy by training the target network model.

[0088] In some embodiments, the network model of the proximal policy optimization algorithm Figure 2 is a schematic diagram of a network model of a proximal policy optimization algorithm provided in the embodiments of the present application, as Figure 2 shown, including: a policy network, a value network, and an advantage network. The reward function of the policy network is:

[0089] where is a three-dimensional vector, x k represents the current state of the sample at time step k, is a two-dimensional vector, d k represents the action executed by the sample, x k represents the x coordinate at time step k, y k represents the y coordinate at time step k, θ k represents the direction angle, v d represents the speed of the attack input, φ d represents the angle of the attack input, r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, x r,k is the reference state.

[0090] The value function of the value network is:

[0091] In the embodiments of the present application, the action value function Q π (s,a) represents the preset cumulative reward starting from the state s = x k , taking the action a = d k , and following the policy π(d k ∣x k ) = P(d k ∣x k ), where γ ∈ [0,1).

[0092] In the embodiments of the present application, in the attack design task, the attacker's goal is to avoid being detected by the detector, so it relies more on current and recent actions rather than long-term actions. Setting γ < 1 can make the agent pay more attention to short-term rewards, thereby improving the practicality and numerical efficiency of the algorithm.

[0093] In the embodiments of the present application, the state function V π (s) evaluates the expected value of the future cumulative reward after starting to follow the policy π from the state s = x k , and the cumulative reward is expressed as:

[0094]

[0095] is the discount factor, R t is the reward at time step t, E is the expected value, and V π (s) represents the expected value of the future cumulative reward after starting to follow the policy π from the state s.

[0096] In the embodiments of the present application, the relationship between V π (s) and Q π (s,a) is: The advantage function is defined as: A π (s,a) = Q π (s,a) - V π (s)

[0097] In the embodiments of the present application, the data for finding the optimal hidden attack strategy is equivalent to solving:

[0098] In the embodiments of the present application, the loss function of the network model of the proximal policy optimization algorithm includes: surrogate objective loss, value function loss, and entropy reward loss, or, the loss function of the neural network model includes: surrogate objective loss, value function loss, entropy reward loss, and auxiliary task loss, and the auxiliary task loss includes: prediction loss of the first mapping model and / or prediction loss of the second mapping model.

[0099] In the embodiments of the present application, the goal of the network model of the proximal policy optimization algorithm is to maximize the cumulative return while ensuring that the policy update does not deviate too far from the previous policy. The loss function consists of three main parts: surrogate objective loss function, value function loss function, and entropy reward loss function. Among them, the surrogate objective function loss is:

[0100]

[0101] Among them, ensures that the policy update is not too large, and π old (d k ∣x k ) is the previous policy.

[0102] In the embodiments of the present application, the value function loss can be expressed as: The entropy reward loss is expressed as: where u k is the executed action.

[0103] In the embodiments of the present application, when the loss function of the network model of the proximal policy optimization algorithm includes: surrogate objective loss, value function loss, and entropy reward loss, it can be expressed as:

[0104]

[0105] where λ VF is the weight of the value function loss, is the weight of the entropy reward loss.

[0106] In some embodiments, Figure 3 is a schematic structural diagram of a network model of a proximal policy optimization algorithm with an auxiliary task introduced according to an embodiment of the present application. As Figure 3 shown, since the policy update of the network model of the proximal policy optimization algorithm follows the policy gradient method, the loss of the auxiliary task can be directly added to the total loss function. If the first mapping model and / or the second mapping model are used, their losses can be directly added to the total loss function. Therefore, the loss function of the neural network model includes: surrogate objective loss, value function loss, entropy reward loss, and auxiliary task loss. The loss function can be expressed as:

[0107]

[0108] where λ aux is the weight of the auxiliary task loss, can be In some embodiments, can also be: In some embodiments, can be and the sum of.

[0109] In some embodiments, Figure 4 is a schematic structural diagram of a network model of a deep reinforcement learning algorithm based on a maximum entropy reinforcement learning framework according to an embodiment of the present application. As Figure 4 shown, the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: an actor network and a critic network. The reward function of the actor network is:

[0110] The value function of the critic network is:

[0111] Among them, is a three-dimensional vector, and x k represents the current state of the sample at time step k. is a two-dimensional vector, and d k represents the action executed by the sample, and x k represents the x coordinate at time step k, and y k represents the y coordinate at time step k, and θ k represents the direction angle, and v d represents the speed of the attack input, and φ d represents the angle of the attack input, and r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, and x r,k is the reference state, and Q π (s,a) represents the expected cumulative reward for taking action a = d k under the state s = x k , and following the policy π, γ is the discount factor, and R t is the reward at time step t. is the expected value, and V π (s) represents the expected value of the future cumulative reward after starting to follow the policy π for the state s.

[0112] In the embodiments of the present application, the loss function of the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: the loss for optimizing the actor network, the loss for optimizing the critic network, and the temperature loss for adjusting the entropy coefficient, or, the loss for the critic network, the loss for the actor network, the temperature loss for adjusting the entropy coefficient, and the auxiliary task loss, and the auxiliary task loss includes: the prediction loss of the first mapping model and / or the prediction loss of the second mapping model.

[0113] In the embodiments of the present application, the actor network can also be called: Actor network, and the critic network can also be called: Critic network.

[0114] In the embodiments of the present application, the loss for optimizing the critic network can be expressed as:

[0115]

[0116] Among them, θ is the network parameter, D is the replay buffer, and the target Q value is:

[0117]

[0118] Among them, α is the temperature parameter. is the policy of the actor network. is the target Q-network with parameters γ is the discount factor.

[0119] The loss used to optimize the actor network can be expressed as:

[0120]

[0121] In the embodiments of the present application, the temperature loss used to adjust the entropy coefficient can be expressed as:

[0122]

[0123] In the embodiments of the present application, the loss function of the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework can be expressed as:

[0124] In some embodiments, Figure 5 is a schematic structural diagram of the network model of a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework with an auxiliary task provided by the embodiments of the present application, as Figure 5 shown. If the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework introduces an auxiliary task, then the loss function can include the loss of the auxiliary task. Therefore, the loss function of the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework can be expressed as:

[0125]

[0126] where λ aux is the weight of the auxiliary task loss, can be In some embodiments, or can also be: In some embodiments, can be and the sum of.

[0127] In some embodiments, determining the execution action for attacking the target vehicle based on the current state and the pre-established target network model includes:

[0128] Step S1021, determining the amount of data of the acquired current state.

[0129] In the embodiments of the present application, the amount of data of the current state can be calculated, for example, the number of data points or the data dimension.

[0130] Step S1022: Determine the target network model from the network model of the proximal policy optimization algorithm and the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework according to the data volume.

[0131] In the embodiments of the present application, according to the size of the data volume, it is determined whether a more complex model (such as SAC) needs to be used to process high-dimensional data, or a simpler model (such as PPO) needs to be used to process low-dimensional data.

[0132] In the embodiments of the present application, if the data volume is large (such as a high-dimensional state space or complex time series data), a network model based on the maximum entropy reinforcement learning framework (SAC) is selected because SAC performs better in processing high-dimensional data and complex dynamics.

[0133] If the data volume is small (such as a low-dimensional state space or a simple dynamic system), a network model based on the proximal policy optimization algorithm (PPO) is selected because PPO has higher computational efficiency for low-dimensional data. The selection of the target network model is based on the trade-off between data volume and computational resources.

[0134] Step S1023: Input the current state into the target network model to output an execution action for attacking the target vehicle.

[0135] In the embodiments of the present application, the selected target network model is used to generate an attack execution action for the target vehicle based on the current state. The execution action can also be referred to as an attack strategy.

[0136] In the embodiments of the present application, the target network model can be selected according to different data situations. As a on-policy method based on gradient descent, PPO is easy to train, but has low data efficiency and weak exploration ability, sometimes resulting in conservative attack strategies. In contrast, as an off-policy method, SAC has high data efficiency and is very suitable for processing high-dimensional continuous spaces.

[0137] In some embodiments, the user can configure the selection strategy. For example, the acquisition speed of the attack strategy can be set to select the target network model based on the speed.

[0138] To train the proposed reinforcement learning strategy, an environment was constructed based on OpenAI Gym. This environment includes various path generators that can generate circular, U-shaped, and figure-eight trajectories for the tracking task of autonomous vehicles. During the trajectory tracking of autonomous vehicles, false data injection attacks with different periods and amplitudes were launched. Subsequently, trajectories were continuously collected for reinforcement learning training. For example, taking a circular trajectory example collected, in the absence of an attack, the controller can accurately track the circular trajectory. In the presence of an attack, the Extended Kalman Filter (EKF) shows significant changes at the attack point but does not diverge and can still track the trajectory well.

[0139] The model can be trained on an NVIDIA RTX 3070Ti GPU using PyTorch version 1.8.1 with a batch size of 64 for 30 epochs. Different types of auxiliary tasks are used to obtain data for training PPO and SAC until convergence. All reinforcement learning strategies can converge well after training. Algorithms with auxiliary tasks tend to achieve higher average rewards, thus indirectly improving the overall performance.

[0140] A simulation experiment platform can be constructed based on the widely used autonomous driving physical engine CARLA. In the simulation environment, attack tests are conducted.

[0141] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0142] According to the foregoing embodiments, an apparatus for determining an attack strategy is provided in an embodiment of the present application. Each module included in the apparatus, as well as each unit included in each module, can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. During the implementation process, the processor can be a central processing unit (CPU, Central Processing Unit), a microprocessor (MPU, Microprocessor Unit), a digital signal processor (DSP, Digital Signal Processing), or a field programmable gate array (FPGA, Field Programmable Gate Array), etc.

[0143] An embodiment of the present application provides an apparatus for determining an attack strategy. Figure 6 FIG. is a schematic structural diagram of an apparatus for determining an attack strategy provided in an embodiment of the present application. As Figure 6 shown, the apparatus 200 for determining an attack strategy includes:

[0144] An acquisition module 201, configured to acquire the current state of the driving trajectory of a target vehicle;

[0145] A determination module 202, configured to determine an execution action for attacking the target vehicle based on the current state and a pre-established target network model, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

[0146] In some embodiments, the attack policy determination device 200 further includes:

[0147] A sample acquisition module, configured to acquire the sample current state and the sample next state of the driving trajectory of the vehicle, the sample execution action of the vehicle, and the sample residual;

[0148] A training data determination module, configured to determine training data based on the sample current state, the sample next state, the sample execution action, and the sample residual;

[0149] A training module, configured to train the target network model based on the training data.

[0150] In some embodiments, the sample acquisition module includes:

[0151] A first prediction unit, configured to input the sample current state and the sample execution action into a pre-established first mapping model to obtain the sample next state.

[0152] In some embodiments, the sample acquisition module includes:

[0153] A second prediction unit, configured to input the sample current state and the sample execution action into a pre-established second mapping model to obtain the sample residual.

[0154] In some embodiments, the network model of the proximal policy optimization algorithm includes a policy network, a value network, and an advantage network, and the reward function of the policy network is: R k =[x k -x r,k T Q[x k -xr,k-rkTRrk;

[0155] The value function of the value network is: The advantage function of the advantage network is A π (s,a)=Q π (s,a)-V π (s), where ​The loss function of the network model of the proximal policy optimization algorithm includes: surrogate objective loss, value function loss, and entropy bonus loss; or, the loss function of the neural network model includes: surrogate objective loss, value function loss, entropy bonus loss, and auxiliary task loss, where the auxiliary task loss includes: prediction loss of the first mapping model and / or prediction loss of the second mapping model. Among them, is a three-dimensional vector, x k represents the current state of the sample at time step k, is a two-dimensional vector, d k represents the action executed by the sample, x k represents the x coordinate at time step k, y k represents the y coordinate at time step k, θ k represents the direction angle, v d represents the speed of the attack input, φ d represents the angle of the attack input, r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, x r,k is the reference state, Q π (s,a) represents the expected cumulative reward obtained by taking action a = d k under the state s = x k and following the policy π, γ is the discount factor, R t is the reward at time step t, is the expected value, V π V(s) represents the expected value of the future cumulative reward after starting to follow the policy π from state s.

[0156] In some embodiments, the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: an actor network and a critic network, and the reward function of the actor network is: The value function of the critic network is: The loss function of the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: the loss for optimizing the actor network, the loss for optimizing the critic network, and the temperature loss for adjusting the entropy coefficient; or, the loss for the critic network, the loss for the actor network, the temperature loss for adjusting the entropy coefficient, and the auxiliary task loss, where the auxiliary task loss includes: prediction loss of the first mapping model and / or prediction loss of the second mapping model. Among them, is a three-dimensional vector, x k represents the current state of the sample at time step k, is a two-dimensional vector, d k represents the action executed by the sample, x k represents the x coordinate at time step k, y k represents the y coordinate at time step k, θ kDenote the direction angle, v d Denote the speed of the attack input, φ d Denote the angle of the attack input, r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, x r,k is the reference state, Q π (s,a) represents the expected cumulative reward obtained by taking the action a = d k under the state s = x k and following the policy π, γ is the discount factor, R t is the reward at time step t, is the expected value, V π V(s) represents the expected value of the future cumulative reward after starting to follow the policy π from the state s.

[0157] In some embodiments, determining the execution action for attacking the target vehicle based on the current state and a pre-established target network model includes:

[0158] Determine the amount of data of the acquired current state;

[0159] Based on the amount of data, determine the target network model from the network model of the proximal policy optimization algorithm and the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework;

[0160] Input the current state into the target network model to output the execution action for attacking the target vehicle.

[0161] In addition, Figure 6 The determining device of the attack strategy shown can be a software unit, a hardware unit, or a unit combining software and hardware built into an existing electronic device, can also be integrated into the electronic device as an independent pendant, or can exist as an independent terminal device.

[0162] It should be noted that for the information interaction, execution process, etc. between the above-mentioned device / units, since they are based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and details are not described here again.

[0163] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0164] Figure 7 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of this application. As Figure 7 shown, the electronic device 3 in this embodiment may include: at least one processor 30 ( Figure 7 only one processor 30 is shown in the figure), a memory 31, and a computer program 32 stored in the memory 31 and executable on at least one processor 30. When the processor 30 executes the computer program 32, the steps in any of the foregoing method embodiments are implemented, or when the processor 30 executes the computer program 32, the functions of each module / unit in the foregoing device embodiments are implemented.

[0165] Exemplarily, the computer program 32 can be divided into one or more modules / units. One or more modules / units are stored in the memory 31 and executed by the processor 30 to complete this application. One or more modules / units can be a series of computer program 32 instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3.

[0166] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program 32, and when the computer program 32 is executed by the processor 30, the steps in any of the foregoing method embodiments can be implemented.

[0167] An embodiment of this application provides a computer program product. When the computer program product runs on an electronic device, it enables the electronic device to implement the steps in any of the foregoing method embodiments when executed.

[0168] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program 32 can be used to instruct relevant hardware to complete. The computer program 32 can be stored in a computer-readable storage medium. When the computer program 32 is executed by a processor 30, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program 32 includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the terminal, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0169] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0170] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0171] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0172] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0173] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

[0174] In each embodiment of the present application, the relevant user personal information that may be involved is all processed in strict accordance with the requirements of laws and regulations, following the principles of legality, legitimacy, and necessity, for reasonable purposes based on business scenarios, and is the personal information actively provided by the user during the use of the product / service or generated due to the use of the product / service, as well as the personal information obtained with the user's authorization.

[0175] The user personal information processed by the applicant will vary according to the specific product / service scenario, and it is necessary to be subject to the specific scenario of the user's use of the product / service. It may involve the user's account information, device information, driving information, vehicle information, or other relevant information. The applicant will treat the user's personal information and its processing with a high degree of diligence.

[0176] The applicant attaches great importance to the security of user personal information and has taken security protection measures that meet industry standards and are reasonable and feasible to protect the user's information and prevent personal information from being accessed, publicly disclosed, used, modified, damaged, or lost without authorization.

Claims

1. A method for determining an attack strategy, characterized in that, Including: Obtain the current state of the driving trajectory of the target vehicle; Based on the current state and a pre-established target network model, determine an execution action for attacking the target vehicle, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

2. The method according to claim 1, wherein The method further includes: Obtain the sample current state and sample next state of the driving trajectory of the vehicle, the sample execution action of the vehicle, and the sample residual; Determine training data based on the sample current state, the sample next state, the sample execution action, and the sample residual; Train the target network model based on the training data.

3. The method according to claim 2, wherein Obtaining the sample current state and sample next state of the driving trajectory, the sample execution action of the vehicle, and the sample residual includes: Input the sample current state and the sample execution action into a pre-established first mapping model to obtain the sample next state.

4. The method according to claim 2, wherein Obtaining the sample current state and sample next state of the driving trajectory, the sample execution action of the vehicle, and the sample residual includes: Input the sample current state and the sample execution action into a pre-established second mapping model to obtain the sample residual.

5. The method according to claim 4, characterized in that The network model of the proximal policy optimization algorithm includes: a policy network, a value network, and an advantage network. The reward function of the policy network is: The value function of the value network is: The advantage function of the advantage network is A π (s,a) = Q π (s,a) - V π (s), where The loss function of the network model of the proximal policy optimization algorithm includes: a surrogate objective loss, a value function loss, and an entropy bonus loss. Or, the loss function of the neural network model includes: a surrogate objective loss, a value function loss, an entropy bonus loss, and an auxiliary task loss. The auxiliary task loss includes: a prediction loss of the first mapping model and / or a prediction loss of the second mapping model, where is a three-dimensional vector, x k represents the current state of the sample at time step k, is a two-dimensional vector, d k represents the action executed by the sample, x k represents the x coordinate at time step k, y k represents the y coordinate at time step k, θ k represents the direction angle, v d represents the speed of the attack input, φ d represents the angle of the attack input, r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, x r,k is the reference state, Q π (s,a) represents the expected cumulative reward obtained by taking the action a = d k under the state s = x k , and following the policy π. γ is the discount factor, R t is the reward at time step t, is the expected value, V π (s) represents the expected value of the future cumulative reward after starting to follow the policy π from the state s.

6. The method according to claim 4, characterized in that The network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: an actor network and a critic network. The reward function of the actor network is: The value function of the critic network is: The loss function of the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework includes: the loss for optimizing the actor network, the loss for optimizing the critic network, and the temperature loss for adjusting the entropy coefficient, or, the loss for the critic network, the loss for the actor network, the temperature loss for adjusting the entropy coefficient, and the auxiliary task loss. The auxiliary task loss includes: the prediction loss of the first mapping model and / or the prediction loss of the second mapping model, where, is a three-dimensional vector, x k represents the current state of the sample at time step k, is a two-dimensional vector, d k represents the sample's executed action, x k represents the x coordinate at time step k, y k represents the y coordinate at time step k, θ k represents the direction angle, v d represents the speed of the attack input, φ d represents the angle of the attack input, r k is the residual, Q is the weight matrix of the state error, R is the weight matrix of the residual, x r,k is the reference state, Q π (s,a) represents the expected cumulative reward for taking action a = d k under the state s = x k , and following the policy π. γ is the discount factor, R t is the reward at time step t, is the expected value, V π (s) represents the expected value of the future cumulative reward after starting to follow the policy π from state s.

7. The method according to claim 1, wherein The determining, based on the current state and a pre-established target network model, an execution action for attacking the target vehicle includes: Determine the amount of data of the obtained current state; Based on the amount of data, determine a target network model from the network model based on the proximal policy optimization algorithm and the network model of the deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework; Input the current state into the target network model to output an execution action for attacking the target vehicle.

8. An apparatus for determining an attack strategy, characterized in that Including: An obtaining module, configured to obtain the current state of the driving trajectory of the target vehicle; A determining module, configured to determine an execution action for attacking the target vehicle based on the current state and a pre-established target network model, where the target network model is a network model based on the proximal policy optimization algorithm or a deep reinforcement learning algorithm based on the maximum entropy reinforcement learning framework.

9. An electronic device, characterized in that, Including: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that when the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.