Optimal data transmission method for unmanned ground vehicle
By constructing a Stackelberg dynamic game model and iteratively optimizing the defense and attack strategies of unmanned ground vehicles, the flexibility and security issues of the unmanned ground vehicle data transmission strategy are solved, and optimal data transmission and system security are achieved.
Patent Information
- Application Number
- CN202510963721.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-14
AI Technical Summary
The data transmission strategies of existing unmanned ground vehicles lack flexibility and dynamic adaptability, and cannot effectively respond to the attacker's changing attack strategies, resulting in increased system security risks and insufficient resource optimization.
A Stackelberg dynamic game model is constructed, and the defense objective function of the unmanned ground vehicle as the leader and the attack objective function of the attacker are established. The optimal attack and defense strategy structure is obtained through iterative optimization, and the defense strategy is dynamically adjusted to deal with the attacker. Sensors, state estimators and controllers are combined for data transmission.
It achieves optimal data transmission for unmanned ground vehicles when facing attacks, prevents data theft and tampering, ensures vehicle safety, reduces remote state errors, controls defense costs, and improves system security.
Smart Images

Figure CN120785602A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of unmanned vehicle control, and particularly relates to an optimal data transmission method for an unmanned ground vehicle. BACKGROUND
[0002] An unmanned ground vehicle (UGV) is a vehicle that can autonomously travel on the ground without human driving. As an important application field of cyber-physical systems (CPS), the UGV has developed rapidly in recent years and has wide application prospects in many fields such as military, agriculture and logistics. However, with the wide application of the UGV, the stability and safety of the UGV in complex environments have become increasingly prominent. In particular, the communication and control system of the UGV is vulnerable to network attacks, which may lead to tampering of vehicle control instructions or theft of sensor data, thereby affecting the safe operation of the vehicle.
[0003] The existing data transmission strategy of the unmanned ground vehicle is usually static, that is, the unmanned ground vehicle formulates fixed data transmission measures according to known threats. This strategy lacks flexibility and dynamic adaptability and cannot effectively respond to the changing attack strategies of attackers. At the same time, the existing data transmission strategy does not model the attackers. Attackers will constantly adjust their attack methods, and the unmanned ground vehicle often has difficulty in making optimal responses in real time, resulting in increased system security risks. In addition, the data transmission strategy is difficult to optimize for the actual attack behavior of the attacker, and cannot achieve the best data transmission effect. SUMMARY
[0004] In view of the above analysis, the embodiments of the present application aim to provide an optimal data transmission method for an unmanned ground vehicle, to solve the problems of the existing static defense strategy and insufficient resource optimization in the safety protection of the unmanned ground vehicle.
[0005] The main purpose of the present application is achieved by the following technical solutions:
[0006] The present application provides an optimal data transmission method for an unmanned ground vehicle, comprising the following steps:
[0007] A Stackelberg dynamic game model is constructed with the unmanned ground vehicle as the leader and the attacker as the follower. The defense target function of the leader and the attack target function of the attacker are constructed and iteratively optimized. When the Stackelberg equilibrium is reached, the optimal attack and defense strategy structure form of the unmanned ground vehicle is obtained. The defense target function is to minimize the remote state error covariance of the unmanned ground vehicle, the defense cost and maximize the attack state error covariance. The attack target function is to maximize the attack state error covariance and minimize the attack cost.
[0008] Based on the optimal attack and defense strategy structure, after the unmanned ground vehicle executes the optimal defense strategy at the current moment for data transmission, the attacker obtains the optimal attack strategy at the current moment based on the optimal defense strategy; the unmanned ground vehicle adjusts its own defense strategy according to the optimal attack strategy at the current moment, and obtains the optimal defense strategy at the next moment for data transmission.
[0009] Furthermore, the unmanned ground vehicle includes a sensor, a local state estimator, a remote state estimator and a controller; wherein,
[0010] The sensor is used to obtain the status of the unmanned ground vehicle;
[0011] The local state estimator is configured to use a Kalman filter to obtain a transmission data packet including a local state estimate of the unmanned ground vehicle and a corresponding local state error covariance based on the state of the unmanned ground vehicle, and transmit the transmission data packet via a wireless network;
[0012] The remote state estimator is configured to receive and parse the transmission data packet to obtain remote state estimation information including a remote state estimation and a corresponding remote state error covariance;
[0013] The controller is configured to generate control instructions based on the remote state estimation information to control the unmanned ground vehicle.
[0014] Furthermore, the defense strategy of the unmanned ground vehicle includes a transmission strategy and an encryption strategy;
[0015] Based on the optimal attack and defense strategy structure, the attacker obtains the optimal attack strategy at the current moment based on the optimal defense strategy, including:
[0016] When the transmission strategy in the optimal defense strategy at time k is not to transmit data, the optimal attack strategy at time k is not to attack;
[0017] Otherwise, use the following formula to get the optimal attack strategy at time k:
[0018]
[0019] in, represents the optimal attack strategy at time k; τ e,k represents the holding time from the moment the attacker successfully eavesdropped on the transmitted data to the moment k; Tr[] represents the trace of the matrix; represents the benchmark error covariance; δ a Indicates the attack energy weight; E a Indicates the energy consumed by the attacker in launching an attack; gk represents the encryption strategy; ε represents the encryption impact factor on the transmission success rate; λ represents the preset probability of the remote state estimator successfully receiving the transmission data packet in the attack-free state; λ a represents the preset probability of the remote state estimator successfully receiving the transmission data packet in the attack state; represents the attack state error covariance expectation weight at the i e,k-1 th attack failure moment at k-1 moment; h() represents the recursive update function of the error covariance; represents; m k-1,i represents the attack state error covariance expectation weight at the i th attack failure moment at k-1 moment; represents the value after i+1 times of recursion on the reference error covariance
[0020] Further, based on the optimal attack-defense strategy structure form, the unmanned ground vehicle adjusts the defense strategy of itself according to the optimal attack strategy at the current moment to obtain the optimal defense strategy at the next moment for data transmission, comprising: calculating the defender immediate reward values of three kinds of data transmission data, i.e. no transmission data, transmission data without encryption and transmission data with encryption, respectively, and based on each reward value, the following judgment is made to obtain the optimal data transmission strategy at the next moment:
[0021] When r d1,k -r d2,k ≤0 and r d1,k -r d3,k ≤0, then
[0022] When r d1,k -r d2,k >0 and r d2,k -r d3,k ≤0, then
[0023] When r d1,k -r d3,k >0 and r d2,k -r d3,k >0, then
[0024] Wherein, r d1,k represents the reward value of no transmission data at k moment; r d2,k represents the reward value of transmission data without encryption at k moment; r d3,k represents the reward value of transmission data with encryption at k moment; represents the optimal defense strategy at k moment; represents the optimal transmission strategy at k moment; represents the optimal encryption strategy at k moment.
[0025] Further, the defender immediate reward function at time k is:
[0026]
[0027] where β represents the system expected value weight; α represents the weight of the remote state error covariance; P k represents the remote state error covariance at time k; δ v represents the transmission energy weight; v k represents the transmission strategy at time k; E v represents the energy consumed for one transmission of data; P e,k represents the attack state error covariance at time k; δ e represents the encryption energy weight; g k represents the encryption strategy at time k; E e represents the energy consumed for one encryption.
[0028] Further, the defense objective function is:
[0029]
[0030] where β k represents the system expected value weight at time k; α represents the weight of the remote state error covariance; P k represents the remote state error covariance at time k; δ v represents the transmission energy weight; v k represents the transmission strategy at time k; E v represents the energy consumed for one transmission of data; P e,k represents the attack state error covariance at time k; δ e represents the encryption energy weight; g k represents the encryption strategy at time k; E e represents the energy consumed for one encryption.
[0031] Further, the remote state error covariance at time k is obtained using:
[0032]
[0033] where P k represents the remote state error covariance at time k; represents the steady state value of the error covariance; P k-1 represents the remote error state covariance at time k-1; h() represents the recursive update function of the error covariance; v k represents the transmission strategy at time k; γ k represents the received state of the remote state estimator at time k.
[0034] Further, the attack target function is:
[0035]
[0036] Wherein, N represents system running time; τ e,k represents the holding time from the last successful eavesdropping time to the k time; Tr[] represents the trace of the matrix; m k,i represents the attack state error covariance expectation weight of the i-th attack failure time; P e,i represents the attack state error covariance of the i-th attack failure time; δ a represents the attack energy weight; a k represents the attack strategy at the k time; E a represents the energy consumed by the attacker to launch an attack once.
[0037] Further, the attack state error covariance is obtained using the following formula:
[0038]
[0039] Wherein, P e,k represents the attack state error covariance at the k time; represents the steady-state value of the error covariance; P e,k-1 represents the attack state error covariance at the k-1 time; h() represents the recursive update function of the error covariance; γ e,k represents the eavesdropping state of the attacker at the k time.
[0040] Further, the optimal attack and defense strategy structure form comprises:
[0041] When the Stackelberg equilibrium is reached, the optimal defense strategy of the defender is:
[0042]
[0043] Wherein, a k represents the attack strategy of the attacker at the k time; η k represents the defense strategy of the defender at the k time; A1 represents the value range of the transmission strategy; A2 represents the value range of the encryption strategy; r d,k () represents the defense reward function;
[0044] The optimal attack strategy of the attacker is:
[0045]
[0046] Wherein, A3 represents the value range of the attack strategy; r a,k () represents the attack reward function.
[0047] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0048] 1. By constructing a Stackelberg game model and solving the Stackelberg equilibrium solution, the solution of the present invention can obtain the optimal data transmission strategy during the operation of the unmanned ground vehicle to cope with communication interference. When facing a DoS attack, it can accurately obtain a data transmission plan, effectively preventing data from being stolen, tampered with, or leaked during the transmission process, ensuring the safety of the vehicle and avoiding vehicle driving risks.
[0049] 2. The present invention models the adversarial relationship between the unmanned ground vehicle and the attacker as a Stackelberg dynamic game model, with the unmanned ground vehicle acting as both the defender and the leader, and the attacker as the follower, to reflect the strategic interaction and strategy sequence between the attacker and the defender. This allows the unmanned ground vehicle to dynamically adapt to the attacker's strategy changes, always be in an active position, formulate the optimal defense strategy in advance, effectively respond to the attacker's various attack methods, and improve the security of the system.
[0050] 3. While ensuring system security and reducing the remote state error of unmanned ground vehicles, the present invention assumes that the attacker obtains the optimal defense strategy under the optimal attack strategy, takes into account the control of defense costs, avoids the waste of resources caused by excessive defense, makes the strategies of both parties more reasonable, and ultimately ensures the information security of unmanned ground vehicles when facing attacks.
[0051] In the present invention, the above-mentioned technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of the present invention will be described in the following description, and some advantages will become apparent from the description or be learned through practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like parts throughout the drawings.
[0053] Figure 1 The figure is a flow chart of an optimal data transmission method for an unmanned ground vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.
[0055] One specific embodiment of the present application discloses an optimal data transmission method of unmanned ground vehicle, like Figure 1 As shown, comprising the following steps S1-S2:
[0056] Step S1, a Stackelberg dynamic game model is constructed with the unmanned ground vehicle as the leader and the attacker as the follower, the defense target function of the leader and the attack target function of the attacker are constructed respectively and iteratively optimized, and the optimal attack and defense strategy structure form of the unmanned ground vehicle is obtained when the Stackelberg equilibrium is reached; wherein the defense target function is to minimize the remote state error covariance of the unmanned ground vehicle, the defense cost and maximize the attack state error covariance; the attack target function is to maximize the attack state error covariance and minimize the attack cost.
[0057] The attacker is an entity that attempts to interfere, damage or obtain information of the unmanned ground vehicle system through illegal means without authorization, which can be a competitor for example.
[0058] Specifically, the unmanned ground vehicle is a linear time-invariant information physical system, and there is a linear relationship between the state change and the state variable and the control input, which makes the behavior of the unmanned ground vehicle easy to predict and control, and the parameters of the unmanned ground vehicle, including mass and translation friction coefficient, remain unchanged in time, so that the dynamic characteristics of the unmanned ground vehicle are stable, and it is convenient to design long-term control and transmission strategy, therefore, the system model of the unmanned ground vehicle is represented as follows:
[0059]
[0060] Wherein, p represents the position of the unmanned ground vehicle; v represents the speed of the unmanned ground vehicle; m represents the mass of the unmanned ground vehicle; μ represents the translation friction coefficient of the unmanned ground vehicle; u represents the control input acting on the unmanned ground vehicle; ω represents a random disturbance obeying a Gaussian distribution with mean zero and covariance represents the change rate of the position of the unmanned ground vehicle, i.e. the speed; represents the change rate of the speed of the unmanned ground vehicle, i.e. the acceleration.
[0061] More specifically, the system model of the unmanned ground vehicle is discretized with a sampling period of 0.1, and the discrete system representation of the unmanned ground vehicle is as follows:
[0062] x k+1 = Ax k + ωk
[0063] y k = Cx k + v k
[0064] wherein, xkdenotes the state vector of the unmanned ground vehicle system at time k, including n x state information of the unmanned ground vehicle system at time k; zkdenotes the measurement output of the unmanned ground vehicle system at time k, including n y system information observed by sensors; A denotes a system matrix, used to represent the dynamic change of system state; C denotes an output matrix, used to represent how the system state is mapped to the measurement output; ωkdenotes system noise, a random variable representing random disturbance affecting the system state at time k, whose value follows a zero-mean Gaussian distribution with covariance vkdennotes measurement noise, a random variable representing random error affecting the measurement output at time k, whose value follows a zero-mean Gaussian distribution with covariance∑ v .
[0065] It should be noted that the random variables ω k and v k are independent of each other, the system pair (A, ) is controllable, i.e., the system can be transferred from any initial state to any desired final state within a finite time through input such as control signal; at the same time, the system pair (A, C) is observable, i.e., the system state can be uniquely determined within a finite time through measurement output.
[0066] More specifically, since (A, ) is controllable and (A, C) is observable, for any initial state, the Kalman filter will enter a steady state after running for a certain time, denoted as xk= f°h(xk) , wherein, h(x) = AXA T +∑ ω is the Lyapunov operator, used to update the covariance matrix of the random process to reflect the influence of system dynamics and noise; f(x) = x-xC T (CXC T +R) -1 CZ, used to calculate the posterior covariance matrix of the state in the Kalman filter to reflect the uncertainty of the system state given the observation data. When the standard Kalman filter enters a steady state at the sensor end, the state error covariance
[0067] Furthermore, the unmanned ground vehicle includes a sensor, a local state estimator, a remote state estimator and a controller; wherein,
[0068] The sensor is used to obtain the status of the unmanned ground vehicle.
[0069] Specifically, the state of the unmanned ground vehicle includes real-time position and real-time speed.
[0070] The local state estimator is used to use Kalman filtering to obtain a transmission data packet including a local state estimation of the unmanned ground vehicle and a corresponding local state error covariance according to the state of the unmanned ground vehicle, and transmit the transmission data packet through a wireless network.
[0071] Specifically, since unmanned ground vehicles are subject to various interferences including system noise and measurement noise during operation, their true state is difficult to obtain directly. Therefore, it is necessary to use a standard Kalman filter and a standard Kalman filter equation to obtain local state estimation information. The Kalman filter is a recursive minimum variance estimation method that can estimate the vehicle state in real time based on the system's dynamic model and measurement model, combined with historical state estimation and current measurement data. The local state estimation information includes local minimum state estimation information and the corresponding local state error covariance
[0072] The remote state estimator is used to receive and parse the transmission data packet to obtain remote state estimation information including remote state estimation and corresponding remote state error covariance.
[0073] Specifically, the remote state estimator receives and parses the transmission data packet sent by the local state estimator through the wireless network, extracts the vehicle state information and error covariance therein, calculates the remote state estimation and the remote state error covariance to evaluate the uncertainty of the state estimation.
[0074] The controller is configured to generate control instructions based on the remote state estimation information to control the unmanned ground vehicle.
[0075] Specifically, based on the remote state estimation information, the controller uses an appropriate control algorithm to generate control instructions, for example, PID control or model predictive control, and sends the control instructions to the actuator to achieve control of the unmanned ground vehicle.
[0076] Further, in the iterative optimization of the Stackelberg dynamic game model with the unmanned ground vehicle as the leader and the attacker as the follower, the unmanned ground vehicle as the defender transmits data based on the optimal defense strategy at the current time; the defense strategy at time k includes the transmission strategy v k ={0,1} (0 means no data transmission, and 1 means data transmission) and the encryption strategy g k ={0,1} (0 means no data encryption, and 1 means data encryption).
[0077] wherein, at time k, when the transmission strategy v k =1 and the encryption strategy g k =0, the local state estimator directly transmits the local state estimation and the local state error covariance P k to the remote state estimator; when the transmission strategy v k =1 and the encryption strategy g k =1, the local state estimator transmits the state estimation and the local state error covariance P k to the remote state estimator after encryption; when the transmission strategy v k =0, the local state estimator does not transmit the data packet to the remote state estimator.
[0078] The holding time from the most recent successful transmission time to the current time k is defined as τ k , that is:
[0079] τ k =k-max 0≤l≤k {l:η l =1}
[0080] wherein η l is an indication variable, indicating whether the transmission data packet is successfully received at time l. If η l =1, it means that the transmission data packet is successfully received at time l; if η l =0, it means that the data packet is not successfully received at time l.
[0081] In the remote state estimator, γ k ={0,1} is used to describe the successful transmission of the local transmission data packet to the remote state estimator, wherein if the remote state estimator receives the transmission data packet, γ k =1; otherwise, γ k =0.
[0082] Therefore, based on the receiving state of the remote state estimator and the transmission strategy, the following formula is established:
[0083]
[0084] wherein, when v k γ k = 1, it means that the data packet is transmitted at k time and the remote state estimator successfully receives the transmitted data packet, τ k = 0, that is, the current time is the successful transmission time; when v k γ k = 0, it means that no data packet is transmitted at k time or the transmitted data packet is not successfully received by the remote state estimator, τ k = τ k-1 + 1, that is, the current time is the time of the previous time when the transmission is not successful + 1.
[0085] During the data transmission process of the unmanned ground vehicle, an attacker will launch a denial of service attack, that is, a DoS attack, on the unmanned ground vehicle, so that the transmitted data packet cannot be normally sent to the remote state estimator, resulting in that the remote state estimator cannot update the state estimation information in time, thereby affecting the strategy of the controller. Therefore, when the attacker launches a DoS attack on the unmanned ground vehicle, the probability of the remote state estimator receiving data will be greatly reduced. When the unmanned ground vehicle is attacked by the attacker, the following formula is used to represent the probability of the remote estimator receiving the transmitted data packet:
[0086]
[0087]
[0088] wherein, γ k represents the receiving state of the remote state estimator at k time; a k represents the attack strategy at k time; v k represents the transmission strategy at k time; g k represents the encryption strategy at k time; λ a represents a preset probability that the remote state estimator successfully receives the transmitted data packet in an attack state; and ε represents an encryption influence factor on the transmission success rate. represents the probability of obtaining γ k = 1 under given conditions. For example, the preset threshold can be set to 100 ms.
[0089] It should be noted that when the transmission strategy v k at k time is 0, that is, the local state estimator does not transmit data, the probability of the receiving state γ k = 1 of the remote state estimator at k time is 0, that is, the receiving state γ k = 0 of the remote state estimator at k time.
[0090] When the transmission strategy v at time k k is 1, that is, the local state estimator performs data transmission; encryption strategy g k is 0, that is, no data encryption is performed; and attack strategy a k is 1, that is, the attacker attacks, then the probability that the remote state estimator receives the data packet is λ a , that is, the receiving state γ of the remote state estimator at time k k =1 is λ a For example, a Can be set to 0.4.
[0091] When the transmission strategy v at time k k is 1, that is, the local state estimator performs data transmission; encryption strategy g k is 1, that is, data encryption is performed during transmission; and attack strategy a k is 1, that is, the attacker attacks, then the probability that the remote state estimator receives the data packet is ελ a ,, that is, the receiving state γ of the remote state estimator at time k k =1 is ελ a Among them, since the encryption process may introduce some additional complexity and overhead, affecting the transmission and reception of data packets, at this time 0≤ε≤1.
[0092] Otherwise, even if the UGV is not attacked, it may still fail to transmit data due to network congestion, but the probability is very small. Therefore, when the UGV is not attacked by an attacker, the probability of the remote estimator receiving the transmitted data packet is expressed as follows:
[0093]
[0094] Where λ represents the probability that the remote state estimator successfully receives the transmitted data packet in the preset no-attack state.
[0095] Specifically, when the transmission strategy v at time k k is 0, that is, when the local state estimator does not transmit data, the receiving state γ of the remote state estimator at time k k =1 has a probability of 0, that is, the receiving state γ of the remote state estimator at time k k =0.
[0096] When the transmission strategy v at time k k is 1, that is, the local state estimator performs data transmission; encryption strategy g kis 0, i.e. no data encryption during transmission, the probability that the remote state estimator receives the data packet is λ, i.e. the receiving state of the remote state estimator at time k is k 1 with probability λ. Exemplarily, λ can be set to 0.98.
[0097] When the transmission strategy v k at time k is 1, i.e. the local state estimator transmits data; the encryption strategy g k at time k is 1, i.e. data encryption during transmission, the probability that the remote state estimator receives the data packet is ελ, i.e. the receiving state of the remote state estimator at time k is k 1 with probability ελ. Here, 0≤ε≤1, since the encryption process can introduce some additional complexity and overhead, affecting the transmission and reception of data packets.
[0098] Further, define the data set collected by the remote state estimator from the initial time to time k as wherein, denotes the sequence of local state estimations from the initial time to time k; γ 1:k denotes the sequence of receiving states of the remote state estimator from the initial time to time k; v 1:k denotes the sequence of transmission strategies from the initial time to time k; based on the data set I k , the remote state estimator can obtain the remote state estimation and the corresponding remote state error covariance P k , respectively, as follows:
[0099]
[0100]
[0101] wherein, when v k γ k =1, i.e. data transmission at time k and successful reception by the remote state estimator, the remote state estimation obtained by the remote state estimator is the local state estimation at time k the remote state error covariance P k is the steady-state value of the error covariance when v k γ k =0, i.e. no data transmission at time k or no successful reception by the remote state estimator, the remote state estimation obtained by the remote state estimator is the predicted value of the prediction of the remote state estimation at time k-1 the remote state error covariance P kis the recursive value h(P k-1 ).
[0102] Furthermore, while the attacker is attacking the UGV, the attacker can obtain the system status information by stealing and decrypting data to more effectively attack the UGV. Therefore, the probability of the attacker successfully eavesdropping on the transmitted data packet and successfully decrypting it is obtained using the following formula:
[0103]
[0104] Among them, γ e,k Represents the eavesdropping state of the attacker at time k, when γ e,k =1, it means the attacker successfully eavesdrops on the data packet and successfully decrypts it. e,k =0, it means that the attacker has not successfully eavesdropped on the data packet or has not successfully decrypted it; e represents the probability that the preset attacker successfully eavesdrops on the transmitted data packet and successfully decrypts it; ε e represents the influence factor of encryption on the success rate of eavesdropping; v k represents the transmission strategy at time k; g k represents the encryption strategy at time k.
[0105] Specifically, when the transmission strategy v at time k k is 0, that is, when the local state estimator does not transmit data, the eavesdropping state γ of the attacker at time k e,k =1 is 0, which means that the attacker cannot eavesdrop on any data packet at time k, that is, γ e,k =0.
[0106] When the transmission strategy v at time k k is 1, that is, the local state estimator performs data transmission; encryption strategy g k is 0, that is, data encryption is not performed during transmission, then the probability that the attacker eavesdrops on the data packet is λ e , that is, the probability γ that the attacker eavesdrops on the data packet at time k e,k =1 is λ e , exemplary, λ e Can be set to 0.8.
[0107] When the transmission strategy v at time k k is 1, that is, the local state estimator performs data transmission; encryption strategy g k If it is 1, that is, data encryption is performed during transmission, then the probability that the attacker eavesdrops on the data packet and successfully decrypts it is ε e λ e , that is, the probability γ that the attacker eavesdrops on the data packet and successfully decrypts it at time ke,k =1 is ε e λ e , where 0≤ε e ≤1.
[0108] Furthermore, the dataset eavesdropped by the attacker from the initial moment to the kth moment is defined as in, represents the local state estimation sequence eavesdropped by the attacker from the initial moment to moment k; γ e,1:e,k Represents the attacker's eavesdropping state sequence from the initial moment to moment k; based on the dataset I e,k , use the following formula to get the attack state estimate and the corresponding attack state error covariance:
[0109]
[0110] Among them, when v k γ e,k = 1, it means that data is transmitted at time k and the attacker successfully attacks. The attacker obtains the attack state estimate is the local state estimate at time k Attack state error covariance P e,k is the steady-state value of the error covariance When v k γ e,k = 0, it means that there is no data transmission at time k or the attacker has not attacked successfully, and the attacker obtains the attack state estimate is the predicted value for the attack state estimation at time k-1 Attack state error covariance P e,k is the recursive value h(P e,k-1 ).
[0111] Furthermore, the Stackelberg dynamic game model is a non-cooperative game model in which one party (the leader) acts first, and the other party (the follower) responds optimally based on the leader's action. In this model, the leader has a first-mover advantage and can predict the follower's response and optimize its own strategy accordingly, choosing the one that is most beneficial to it. The follower, on the other hand, acts later and responds optimally based on the leader's chosen strategy.
[0112] Furthermore, the leader's defense objective function is:
[0113]
[0114] Among them, β kdenotes the system expected value weight at time k, used to balance the importance of different time steps; a denotes the weight of the remote state error covariance, used to balance the importance of the remote state error covariance and the attack state error covariance; P k denotes the remote state error covariance at time k; d v denotes the transmission energy weight, used to balance the cost of transmission decision; v k denotes the transmission strategy at time k; E v denotes the energy consumed by one transmission of data; P e,k denotes the attack state error covariance at time k; d e denotes the encryption energy weight, used to balance the cost of encryption decision; g k denotes the encryption strategy at time k; E e denotes the energy consumed by one encryption.
[0115] Specifically, under the optimal attack strategy of the attacker, the unmanned ground vehicle as the defender expects the estimation of the system state in the remote state estimator to be as accurate as possible, and at the same time, while achieving the defense goal, the cost of the defense measure is reduced as much as possible, and at the same time, the attacker expects the estimation of the system state to be as inaccurate as possible.
[0116] More specifically, by minimizing the remote state error covariance aTr[P k ], the defender expects the estimation of the system state in the remote state estimator to be as accurate as possible; by minimizing the transmission cost d v v k E v and the encryption cost d e g k E e , the defender expects to reduce the defense cost as much as possible while achieving the defense goal; by maximizing the attack state error covariance (1-a)Tr[P e,k ], the defender expects the estimation of the system state by the attacker to be as inaccurate as possible.
[0117] The expected values of the remote state error covariance and the attack state error covariance are calculated using the following formula:
[0118]
[0119]
[0120] Further, the attack target function of the attacker is:
[0121]
[0122] where N denotes the system running time; t e,kdenotes the holding time from the moment of the last successful eavesdropping to the moment k; Tr[] denotes the trace of a matrix; m k,i denotes the attack state error covariance expectation weight at the i th moment of attack failure; P e,i denotes the attack state error covariance at the i th moment of attack failure; δ a denotes the attack energy weight; a k denotes the attack strategy at the moment k; E a denotes the energy consumed by the attacker to launch an attack.
[0123] Specifically, denotes the sum of the attack effects of the attacker at different time points, and the attacker maximizes this value to represent that the greater the interference of the attack on the system state estimation; meanwhile, the attack cost δ is minimized during the attack process a a k E a .
[0124] In the policy iteration process, the unmanned ground vehicle randomly selects an initial defense strategy, i.e., whether to perform data transmission and whether to perform data encryption when performing data transmission; the attacker formulates an initial attack strategy according to the initial defense strategy of the unmanned ground vehicle, i.e., whether to perform an attack; the unmanned ground vehicle updates the defense strategy at the next moment according to the initial attack strategy of the attacker; wherein at each iteration moment, the instantaneous reward of the unmanned ground vehicle and the attacker is calculated using the following formula respectively, and the strategy at the next moment is updated based on the instantaneous reward to optimize the respective objective functions:
[0125] The instantaneous reward function of the unmanned ground vehicle at the moment k is:
[0126]
[0127] The instantaneous reward function of the attacker at the moment k is:
[0128]
[0129] Specifically, the instantaneous reward functions of the unmanned ground vehicle and the attacker respectively reflect the measurement of the benefits and costs of different strategies at each moment, and reflect the direct results brought by the strategy selection of the unmanned ground vehicle and the attacker at each moment.
[0130] The above iteration process is repeated until a Stackelberg equilibrium state is reached, i.e., the unmanned ground vehicle and the attacker cannot obtain better results by changing their own strategies unilaterally.
[0131] Further, the optimal attack and defense strategy structure form includes:
[0132] When the Stackelberg equilibrium is reached, the optimal defense strategy of the unmanned ground vehicle is:
[0133]
[0134] Among them, a k represents the attacker's attack strategy at time k; η k represents the defense strategy of the defender at time k; A1 represents the value range of the transmission strategy; A2 represents the value range of the encryption strategy; r d,k () represents the defense reward function.
[0135] The attacker's optimal attack strategy is:
[0136]
[0137] Among them, A3 represents the value range of the attack strategy; r a,k () represents the attack reward function.
[0138] Specifically, the optimal attack and defense strategy structure is an expression of the optimal strategies of the leader (unmanned ground vehicle) and the follower (attacker) when the Stackelberg equilibrium is reached, that is, the optimal strategy of the leader and the optimal strategy of the follower form a mutually dependent relationship. The leader's strategy is selected based on the follower's possible strategy, and the follower's strategy is selected based on the leader's strategy. This mutually dependent relationship enables the strategies of both parties to form a structurally stable equilibrium state.
[0139] Step S2: Based on the optimal attack and defense strategy structure, the unmanned ground vehicle executes the optimal defense strategy at the current moment for data transmission; based on the optimal defense strategy, the attacker obtains the optimal attack strategy at the current moment; the unmanned ground vehicle adjusts its own defense strategy according to the optimal attack strategy at the current moment, obtains the optimal defense strategy at the next moment for data transmission.
[0140] Specifically, since the Stackelberg dynamic game model is a non-cooperative game model, in the decision-making process, the unmanned ground vehicle and the attacker will not jointly optimize the overall benefits through cooperation, but independently select their own strategies to maximize their own benefits, while the unmanned ground vehicle as the leader of the game, first selects its optimal encryption transmission strategy, usually has the advantage of first mover, can layout and optimize its own strategy in advance, so that the attacker needs to weigh more factors when selecting the attack strategy, thereby reducing the success rate and benefits of the attack; the attacker is a follower, selects its optimal attack strategy according to the strategy obtained by eavesdropping from the unmanned ground vehicle, and needs to consider the strategy of the defender when formulating the optimal attack strategy, which makes the attacker cannot attack at will, but must weigh between cost and benefit.
[0141] Further, based on the optimal attack and defense strategy structure form, the attacker obtains the optimal attack strategy at the current time based on the optimal defense strategy, comprising:
[0142] When the transmission strategy in the optimal defense strategy at the k time is not to transmit data, the optimal attack strategy at the k time is not to attack.
[0143] Specifically, since the unmanned ground vehicle selects not to transmit data, because there is no data transmission, the attack is meaningless, so the optimal attack strategy of the attacker at this time is not to attack, that is,
[0144] Otherwise, the optimal attack strategy at the k time is obtained using the following formula:
[0145]
[0146]
[0147] Wherein, represents the optimal attack strategy at the k time, represents not to attack, 1 represents to attack; τ e,k represents the holding time from the successful time of the attacker eavesdropping transmission data to the k time; Tr[] represents the trace of the matrix; represents the steady-state value of the error covariance; δ a represents the attack energy weight; E a represents the energy consumed by the attacker to launch an attack once; g k represents the encryption strategy; ε represents the influence factor of encryption on the transmission success rate; λ represents the preset probability that the remote state estimator successfully receives the transmission data packet in the no attack state; λ arepresents the probability of the remote state estimator successfully receiving the transmission data packet in the preset attack state; represents the attack state error covariance expectation weight at the i-th attack failure moment at time k-1; e,k-1 represents the recursive update function of the error covariance. represents the attack state error covariance expectation weight at the i-th attack failure moment at time k-1; k-1,i represents the attack state error covariance expectation weight at the i-th attack failure moment at time k-1; represents the steady-state value of the error covariance after i+1 recursions.
[0148] Specifically, the function f(τ e,k ) is a value related to the error covariance, which is used to evaluate the state of the unmanned ground vehicle at time k, and can more comprehensively reflect the dynamic changes and stability of the unmanned ground vehicle, which comprehensively considers the holding time τ e,k from the most recent successful eavesdropping transmission data moment of the attacker to time k, the weight of the historical attack failure moment, and the dynamic change of the reference error covariance , wherein the steady-state value of the error covariance describes the error level of the unmanned ground vehicle under the condition of no attack and no encrypted transmission, which is a scalarized measure for evaluating the overall error level of the unmanned ground vehicle, the higher the error level, the more unstable the state, and the greater the possibility of success of the attacker, thereby the function f(τ e,k ) can dynamically evaluate the state of the unmanned ground vehicle at the current moment through τ e,k and , so that the attacker can decide whether to attack according to the real-time state of the unmanned ground vehicle, instead of based on static or outdated information.
[0149] The threshold value thereafter includes the trace of the steady-state value of the error covariance The trade-off between the cost and benefit of the attack of the attacker and the cumulative effect M of the historical attack failure on the current state of the unmanned ground vehicle reflect the error level of the unmanned ground vehicle system, the trade-off between the cost and benefit of the attack of the attacker considering the encryption strategy and the transmission success rate, and the cumulative effect of the historical attack failure on the current state of the unmanned ground vehicle.
[0150] Based on the above factors, the attacker can obtain the optimal attack strategy under the optimal defense strategy at the current moment.
[0151] Further, based on the optimal attack-defense strategy structure form, the unmanned ground vehicle adjusts its defense strategy according to the optimal attack strategy at the current time to obtain an optimal defense strategy at the next time for data transmission, including: calculating defender immediate reward values of three kinds of data transmission data, i.e., no transmission of data, transmission of data without encryption, and transmission of data with encryption, respectively, and based on each reward value, making the following judgments to obtain the optimal data transmission strategy at the next time:
[0152] when r d1,k -r d2,k > 0 and r d1,k -r d3,k > 0, then
[0153] when r d1,k -r d2,k > 0 and r d2,k -r d3,k > 0, then
[0154] when r d1,k -r d3,k > 0 and r d2,k -r d3,k > 0, then
[0155] wherein r d1,k represents a reward value of no transmission of data at k time; r d2,k represents a reward value of transmission of data without encryption at k time; and r d3,k represents a reward value of transmission of data with encryption at k time. represents an optimal defense strategy at k time. represents an optimal transmission strategy at k time. represents an optimal encryption strategy at k time.
[0156] Specifically, the unmanned ground vehicle adjusts its defense strategy according to the optimal attack strategy of the attacker to obtain an optimal data transmission strategy at the next time, and a calculation formula of the immediate reward value of the unmanned ground vehicle is:
[0157]
[0158] wherein, represents a trace of an expected value of a system error covariance at k time, reflecting an error level of the system at k time, and the higher the error level, the worse the performance of the unmanned ground vehicle, i.e., the smaller the value, the better the performance of the unmanned ground vehicle, and the larger the value, the worse the performance of the unmanned ground vehicle; The trace of the expectation value of the attack state error covariance of the attacker at the k moment, reflects the error level of the attacker at the k moment, the higher the error level, the worse the attack effect, that is, the smaller the value, the better the attack performance, and the larger the value, the worse the attack performance.
[0159] Therefore, when the unmanned ground vehicle selects a transmission strategy, the smaller the immediate reward value represents the better defense performance.
[0160] Therefore, the unmanned ground vehicle can effectively cope with the optimal attack strategy of the attacker by using the optimal defense strategy at the next moment, and optimal data transmission is achieved.
[0161] In summary, the optimal data transmission method of the unmanned ground vehicle has the following beneficial effects:
[0162] 1. The scheme of the present application can obtain the optimal data transmission strategy by constructing a Stackelberg game model and solving the Stackelberg equilibrium solution, so as to cope with communication interference and obtain the data transmission scheme when facing DoS attack, effectively prevent data from being stolen, tampered or leaked during transmission, ensure the safety of the vehicle, and avoid the risk of vehicle driving.
[0163] 2. The present application models the antagonistic relationship between the unmanned ground vehicle and the attacker as a Stackelberg dynamic game model, the unmanned ground vehicle as the defender and the leader, and the attacker as the follower, to reflect the strategy interaction and sequence of the attack and defense sides, so that the unmanned ground vehicle can dynamically adapt to the strategy change of the attacker, always in the active position, and formulate the optimal defense strategy in advance, effectively cope with various attack means of the attacker, and improve the security of the system.
[0164] 3. The present application guarantees the safety of the system and reduces the remote state error of the unmanned ground vehicle, while assuming that the attacker obtains the optimal defense strategy under the optimal attack strategy, taking into account the control of the defense cost, avoiding the waste of resources caused by excessive defense, making the strategies of both sides more reasonable, and finally ensuring the information security of the unmanned ground vehicle when facing attacks.
[0165] The above is only the preferred embodiment of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. An optimal data transmission method for an unmanned ground vehicle, characterized in that: The steps include: A Stackelberg dynamic game model is constructed with an unmanned ground vehicle as a leader and an attacker as a follower. The leader's defense objective function and the attacker's attack objective function are constructed and iteratively optimized. When a Stackelberg equilibrium is reached, the optimal attack and defense strategy structure of the unmanned ground vehicle is obtained. The defense objective function is to minimize the long-range state error covariance and defense cost of the unmanned ground vehicle and maximize the attack state error covariance; the attack objective function is to maximize the attack state error covariance and minimize the attack cost. Based on the optimal attack and defense strategy structure, after the unmanned ground vehicle executes the optimal defense strategy at the current moment for data transmission, the attacker obtains the optimal attack strategy at the current moment based on the optimal defense strategy; the unmanned ground vehicle adjusts its own defense strategy according to the optimal attack strategy at the current moment, and obtains the optimal defense strategy at the next moment for data transmission.
2. The method according to claim 1, characterized in that The unmanned ground vehicle includes a sensor, a local state estimator, a remote state estimator and a controller; wherein, The sensor is used to obtain the status of the unmanned ground vehicle; The local state estimator is configured to use a Kalman filter to obtain a transmission data packet including a local state estimate of the unmanned ground vehicle and a corresponding local state error covariance based on the state of the unmanned ground vehicle, and transmit the transmission data packet via a wireless network; The remote state estimator is configured to receive and parse the transmission data packet to obtain remote state estimation information including a remote state estimation and a corresponding remote state error covariance; The controller is configured to generate control instructions based on the remote state estimation information to control the unmanned ground vehicle.
3. The method according to claim 2, characterized in that The defense strategy of the unmanned ground vehicle includes a transmission strategy and an encryption strategy; Based on the optimal attack and defense strategy structure, the attacker obtains the optimal attack strategy at the current moment based on the optimal defense strategy, including: When the transmission strategy in the optimal defense strategy at time k is not to transmit data, the optimal attack strategy at time k is not to attack; Otherwise, use the following formula to get the optimal attack strategy at time k: in, represents the optimal attack strategy at time k; τ e,k represents the holding time from the moment the attacker successfully eavesdropped on the transmitted data to the moment k; Tr[] represents the trace of the matrix; represents the benchmark error covariance; δ a Indicates the attack energy weight; E a Indicates the energy consumed by the attacker in launching an attack; g k represents the encryption strategy; ε represents the influence factor of encryption on the transmission success rate; λ represents the probability that the remote state estimator successfully receives the transmission data packet in the preset non-attack state; λ a It represents the probability that the remote state estimator successfully receives the transmitted data packet under the preset attack state; represents the τth time at time k-1 e,k-1 The expected weight of the attack state error covariance at the moment of attack failure; h() represents the recursive update function of the error covariance; Indicates; m k-1,i represents the expected weight of the covariance of the attack state error at the time of the i-th attack failure at time k-1; Denotes the covariance of the benchmark error The value after i+1 recursions.
4. The method according to claim 3, characterized in that Based on the optimal attack and defense strategy structure, the unmanned ground vehicle adjusts its own defense strategy according to the optimal attack strategy at the current moment to obtain the optimal defense strategy for data transmission at the next moment, including: calculating the defender's instant reward value for three types of data transmission: not transmitting data, transmitting data without encryption, and transmitting data with encryption, and performing the following judgment based on each reward value to obtain the optimal data transmission strategy at the next moment: When r d1,k -r d2,k ≤0 and r d1,k -r d3,k When ≤0, When r d1,k -r d2,k >0 and r d2,k -r d3,k When ≤0, When r d1,k -r d3,k >0 and r d2,k -r d3,k >0, then Among them, r d1,k represents the reward value for not transmitting data at time k; r d2,k represents the reward value for transmitting data without encryption at time k; r d3,k represents the reward value for transmitting and encrypting data at time k; represents the optimal defense strategy at time k; represents the optimal transmission strategy at time k; represents the optimal encryption strategy at time k.
5. The method according to claim 4, characterized in that: The defender's immediate reward function at time k is: Among them, β represents the system expected value weight; α represents the weight of the remote state error covariance; P k represents the remote state error covariance at time k; δ v represents the transmission energy weight; v k represents the transmission strategy at time k; E v Indicates the energy consumed in transmitting data once; P e,k represents the covariance of the attack state error at time k; δ e represents the encryption energy weight; g k represents the encryption strategy at time k; E e Indicates the energy consumed by one encryption.
6. The method according to claim 2, characterized in that: The defense objective function is: Among them, β k represents the weight of the system expected value at time k; α represents the weight of the remote state error covariance; P k represents the remote state error covariance at time k; δ v represents the transmission energy weight; v k represents the transmission strategy at time k; E v Indicates the energy consumed in transmitting data once; P e,k represents the covariance of the attack state error at time k; δ e represents the encryption energy weight; g k represents the encryption strategy at time k; E e Indicates the energy consumed by one encryption.
7. The method according to claim 6, characterized in that The remote state error covariance at time k is obtained using the following formula: Among them, P k represents the remote state error covariance at time k; represents the steady-state value of the error covariance; P k-1 represents the long-range error state covariance at time k-1; h() represents the recursive update function of the error covariance; v k represents the transmission strategy at time k; γ k represents the received state of the remote state estimator at time k.
8. The method according to claim 6, characterized in that The attack target function is: Where N represents the system operation time; τ e,k represents the holding time from the most recent successful eavesdropping moment to moment k; Tr[] represents the trace of the matrix; m k,i represents the expected weight of the attack state error covariance at the moment of the i-th attack failure; P e,i represents the covariance of the attack state error at the moment of the i-th attack failure; δ a Indicates the attack energy weight; a k represents the attack strategy at time k; E a Indicates the energy consumed by the attacker to launch an attack.
9. The method according to claim 8, characterized in that The attack state error covariance is obtained using the following formula: Among them, P e,k represents the covariance of the attack state error at time k; represents the steady-state value of the error covariance; P e,k-1 represents the attack state error covariance at time k-1; h() represents the recursive update function of the error covariance; γ e,k represents the eavesdropping state of the attacker at time k.
10. The method according to any one of claims 1 to 9, characterized in that: The optimal attack and defense strategy structure includes: When the Stackelberg equilibrium is reached, the optimal defense strategy of the defender is: Among them, a k represents the attacker's attack strategy at time k; η k represents the defense strategy of the defender at time k; A1 represents the value range of the transmission strategy; A2 represents the value range of the encryption strategy; r d,k () represents the defense reward function; The attacker's optimal attack strategy is: Among them, A3 represents the value range of the attack strategy; r a,k () represents the attack reward function.
Citation Information
Cited By
Two-vehicle target defense intelligent game playing method based on reachable set
CN121069999A