Hidden data poisoning attack method, system, program and equipment for offline reinforcement learning and storage medium
By identifying and perturbing key decision sequences in offline reinforcement learning, and using a dual-objective optimization algorithm to maximize TD errors, the problem of insufficient effectiveness and invisible attacks in the existing technology of offline reinforcement learning data poisoning attacks is solved, and efficient and concealed attack effects are achieved.
Patent Information
- Application Number
- CN202510064154.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to achieve hidden and efficient data poisoning attacks in offline reinforcement learning, especially in continuous tasks. The attack effectiveness is insufficient and not hidden enough, which limits the research of defense strategies.
By identifying and perturbing key decision sequences, reducing the diversity of decision sequences in the dataset, a two-objective optimization algorithm is used to maximize TD errors while meeting the concealment requirements, thereby improving attack efficiency and concealment.
In the case of low toxicity (such as 1% and 5%), the performance of agents obtained by toxic data training can be significantly reduced, with an average drop of 84% and 90%, while improving the concealment of the attack.
Smart Images

Figure CN120031097A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of offline reinforcement learning, and specifically relates to a covert data poisoning attack method, system, program, device and storage medium for offline reinforcement learning. Background Art
[0002] Offline reinforcement learning, also known as batch reinforcement learning, is a paradigm in which a learning agent learns from a previously collected set of experience data. Compared to online reinforcement learning, which requires obtaining feedback from the environment in real time to improve the strategy, offline RL can learn without online interaction with the environment. It is suitable for exploring expensive, time-consuming or risky scenarios, such as autonomous driving, healthcare decision-making, intelligent robot control, and game design.
[0003] Data poisoning attacks have been widely used in the field of machine learning. In the early days, they mainly targeted basic models such as logistic regression and support vector machines (Marco Barreno, Blaine Nelson, Anthony D. Joseph, and JD Tygar. The security of machine learning. Mach. Learn., 81(2): 121–148, November 2010. ISSN 0885-6125. doi: 10.1007 / s10994-010-5188-5.). In recent years, with the development of deep learning, attackers have gradually turned to attacking deep networks (Feng, J., Cai, Q.-Z., and Zhou, Z.-H. Learning to confuse: generating training time adversarial data with autoencoder. Advances in Neural Information Processing Systems, 32, 2019.). These works show that data security has become an issue that cannot be ignored. As an important branch of the current machine learning field, RL also has data security issues. At present, there are many works on data poisoning attacks on online RL, and they are quite in-depth (Rakhsha, A., Zhang, X., Zhu, X., and Singla, A. Reward poisoning in reinforcement learning: Attacks against unknown learners in unknown environments. arXiv preprint arXiv:2102.08492, 2021.)
[0004] Offline RL also faces the same threat of data poisoning as online RL. The above work on data poisoning for online RL provides many ideas for the research of offline RL. There are currently a few studies on data poisoning attacks for offline reinforcement learning. Ma (Ma, Y., Zhang, X., Sun, W., and Zhu, J. Policy poisoning in batch reinforcement learning and control. Advances in Neural Information Processing Systems, 2019.) et al. first conducted a poisoning attack on an offline reinforcement learning dataset. However, this method attacks each time step, and the poisoning ratio is too high to be detected. Rakhsha (Rakhsha, A., Radanovic, G., Devidze, R., Zhu, X., and Singla, A. Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning. In International Conference on Machine Learning, pp. 7974–7984. PMLR, 2020.) et al. proposed an attack method that balances effect and cost, but only focused on simple discrete tasks and did not target continuous tasks that are closer to real scenarios (such as robots, autonomous driving, etc.). Gong (GONG C, YANG Z, BAI Y, et al. BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets [C] / / 2024 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 2024: 218-218.) et al. proposed an attack method that inserts triggers in offline datasets, but it requires the attacker to be able to perform the attack in both the training and testing phases, which requires too high permissions for the attacker. The above methods have the problems of high poisoning ratio and large disturbance amplitude, which result in insufficient attack effectiveness and lack of concealment. It is difficult to simulate the scenarios of covert attacks in the real world, which limits the research on offline reinforcement learning defense strategies. Summary of the invention
[0005] The purpose of the present invention is to provide a data poisoning method for offline reinforcement learning, which can reduce the diversity of decision sequences in a data set by identifying and perturbing key decision sequences, thereby increasing the concealment of the attack while improving the attack efficiency.
[0006] The present invention provides a hidden data poisoning attack method for offline reinforcement learning, comprising the following steps:
[0007] Step 1: Get the state space S, action space A and reward space R in the clean offline dataset D;
[0008] Step 2: Use the clean dataset to train a clean agent model, using the action-value function to represent each state-action pair (s t ,a t ) performs TD error calculation to obtain the TD error value of each time step;
[0009] Step 3: Sort all time steps in the data set according to the TD error value, select the time steps with larger TD errors as key time steps, and form a key time step set C;
[0010] Step 4: For the state-action pairs (s in the critical time step set t ,a t ) Add disturbances, and the disturbances are optimized through a dual-objective optimization algorithm to maximize the TD error while meeting the requirements of concealment;
[0011] Step 5: The optimized perturbation η t * The state-action pair (s) of each data added to the key time step set t ,a t ) to obtain the poisoned data set D' and complete the entire poisoning process.
[0012] Furthermore, the step 2 specifically includes the following steps:
[0013] Step 2.1: Transform the trajectory τ at each moment t =(s t ,a t ,r r ) to process the corresponding state and input the state data s t Flatten into 1 dimension and convert into integer type, then input action data a t Flatten into 1 dimension and splice behind the state data along dimension 1 to form a new state-action pair (s t ,a t )enter;
[0014] Step 2.2: Use a clean dataset to pre-train an agent with normal performance. First, based on the offline reinforcement learning algorithm, use the clean dataset D to train the agent, and evaluate the performance of the training model in the test environment. According to the evaluation results, adjust the training parameters so that the average cumulative reward can reach the average level in existing research.
[0015] Step 2.3: Transform the new state-action pair (s t ,a t ) is input into the action value Q network of the clean agent to calculate the state-action pair Q at time t θ (s t, a t ), get the state-action pair (s) at time t+1 t+1, a t+1 ) and calculate Q θ (s t+1, a t+1 );
[0016] Step 2.4: Calculate the TD error δ for each time step in the dataset t :
[0017] Q θ (s t ,a t )←Q θ (s t ,a t )+α·δ t
[0018]
[0019] Among them, α is the learning rate and γ is the discount factor.
[0020] Furthermore, the key time step set C of step 3 is: the data set is divided according to the TD error δ t Sort from large to small and select error δ according to the poisoning ratio p t The larger time step is taken as the key time step, and the key time step set C is obtained as the target of subsequent data poisoning attacks.
[0021] Furthermore, the step 4 specifically includes the following steps:
[0022] Step 4.1: For each state-action pair (s corresponding to each time step t in the key time step set C, t ,a t ), using a constrained bi-objective programming problem to find the optimal perturbation
[0023]
[0024] in, represents the state-action pair (s t ,a t )Add disturbance The L2 norm of δ t ' is the TD error after adding disturbance, ∈ is the upper limit of the disturbance amplitude;
[0025] Step 4.2: Use the weighted sum method to transform the above dual-objective optimization problem into a single-objective optimization problem:
[0026]
[0027] Among them, β 1 and β 2 is the weight parameter, β 1 and β 2 Take 1 for both;
[0028] Step 4.3: Solve the bi-objective optimization problem and use the sequential least squares programming algorithm as the optimization solver to solve the objective function To solve.
[0029] Furthermore, the step 5 specifically includes the following steps:
[0030] Step 5.1: Obtain the state-action pairs (s) at each critical time step t t ,a t );
[0031] Step 5.2: Solve the optimal perturbation The state-action pairs (s) corresponding to the key time steps added to the original dataset D t ,a t ) to generate a poisoned data set D';
[0032] Step 5.3: Use D' to train the agent and obtain the poisoned agent A';
[0033] Step 5.4: Evaluate the degradation of the performance of the poisoned agent A' relative to the clean agent A in the test environment to quantify the negative impact of the attack on the offline reinforcement learning agent.
[0034] The present invention also provides a computer device / equipment / system, comprising a memory, a processor, and a computer program stored in the memory, wherein when the processor executes the computer program, the steps of any of the above-mentioned data poisoning attack methods for offline reinforcement learning are implemented.
[0035] The present invention also provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned data poisoning attack methods for offline reinforcement learning.
[0036] The present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of any of the above-mentioned data poisoning attack methods for offline reinforcement learning.
[0037] The beneficial effects of the present invention are:
[0038] 1. The present invention proposes a data poisoning method for offline reinforcement learning, which can be used to analyze the offline reinforcement learning process to find that the time step with a larger time difference error is more important, representing a weaker link in the learning process, and determine it as a critical time step that has a significant impact on the learning task. At the same time, dynamic disturbances are generated for the critical time step, and the dual-objective optimization idea is used to specifically design the maximization and minimization objectives, thereby improving the effectiveness of the attack while ensuring concealment; and the present invention is universal to different offline reinforcement learning algorithms and different reinforcement learning tasks. The attack method proposed in the present invention has been verified in four offline reinforcement learning algorithms: Batch-Constrained Q-learning (BCQ), Batch-Ensemble Actor-Critic with Retrace (BEAR), Conservative Q-Learning (CQL) and BC (Behavioural Cloning); Walker2D, Hopper and Half-Cheetah in the MuJoCo robot simulator, and Carla-Lane autonomous driving in the Carla simulator. When the poisoning ratio is only 1%, the performance of the intelligent agent trained with poisonous data is reduced by an average of 84% compared with the performance of the intelligent agent trained with clean data sets; when the poisoning ratio is 5%, the performance of the intelligent agent is reduced by 90%.
[0039] 2. The disturbance generation method based on dual-objective optimization proposed in the present invention can generate tiny disturbances that have a significant impact on the performance of the intelligent agent. Compared with the original data, the change is very small and imperceptible, with a magnitude of only 0.05, which improves the concealment of the attack. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is an overview diagram of the covert data poisoning method for offline reinforcement learning provided by the present invention. DETAILED DESCRIPTION
[0041] The present invention is further described below in conjunction with the accompanying drawings.
[0042] The hidden data poisoning attack system for offline reinforcement learning of the present invention includes a key time step positioning module and a poisoning attack module based on dual-objective optimization;
[0043] Key time step location module: This module is responsible for identifying key time steps with significant TD errors from offline reinforcement learning datasets. By accurately locating these key time steps, it can maximize the attack effect and reduce the overall decision diversity of the dataset. Its function is to find the key time steps in the reinforcement learning process through TD error analysis and perform subsequent poisoning attack operations on these time steps. The core parts of this module are the TD error calculation module and the key time step extraction module.
[0044] TD error calculation module: This module measures the impact of each time step on policy learning by calculating the TD error of each time step. Its function is to perform TD error analysis on each state-action pair in the data set and sort the time steps by the TD error size, thereby screening out the most critical time steps.
[0045] Critical time step extraction module: This module screens the time steps according to the TD error size, extracts the time steps that have the greatest impact on model training, and generates a set of critical time steps. Its function is to ensure that only the most critical time steps are poisoned and reduce interference with irrelevant time steps.
[0046] Poisoning attack module based on dual-objective optimization: This module is responsible for poisoning the identified key time steps. Through the dual-objective optimization algorithm, the TD error perturbation is maximized while maintaining the concealment of the attack. Its role is to apply perturbations to the key time steps in the offline data set to achieve the poisoning effect while maintaining the similarity between the poisoned data and the original data to avoid being detected. The core parts of this module are the dual-objective optimization module, the perturbation generation module, and the data replacement module.
[0047] Dual-objective optimization module: This module defines the optimization objectives of poisoning attacks, including minimizing the amount of disturbance to maintain data concealment and maximizing the TD error to enhance the attack effect. Its role is to determine the optimal disturbance amplitude through the optimization algorithm, making the attack both effective and concealed.
[0048] Perturbation generation module: This module is responsible for generating poisoning perturbations and adding them to the state-action pairs in the key time step to generate poisoned data. Its role is to generate the optimal perturbation based on the results of the dual-objective optimization and ensure that it has both concealment and attack effectiveness.
[0049] Data replacement module: This module replaces the corresponding time steps in the original data with the generated poisoned data and generates the final poisoned data set. Its function is to replace the key time steps after adding disturbances to the original data set to ensure that the poisoned data set can be used to train poisoned agents. The present invention provides a data poisoning method for offline reinforcement learning, which reduces the diversity of decision sequences in the data set by identifying and disturbing key decision sequences, and can increase the concealment of the attack while improving the efficiency of the attack.
[0050] like Figure 1 As shown, the present invention discloses a hidden data poisoning method for offline reinforcement learning, and the present invention comprises the following steps:
[0051] 1) Obtain the state space and action space in a clean offline dataset.
[0052] 2) Use the clean dataset to train the clean agent model, and use the action value function to calculate the state-action pair at each time step.<state,action> Perform TD error calculation to obtain the TD error value for each time step.
[0053] 3) Sort all time steps in the data set according to the TD error value, select the time steps with larger TD errors as key time steps, and form a set of key time steps.
[0054] 4) Add perturbations to the state-action pairs in the set of critical time steps, and maximize the TD error while meeting the hidden requirements through a bi-objective optimization algorithm.
[0055] 5) Add the optimized perturbation to the state-action pair of each data in the key time step set<state,action> On the , obtain the poisoned data set and complete the entire poisoning process.
[0056] In step 2), a clean agent model A is trained using a clean dataset, and the state-action pair at each time step is computed using the action-value function<state,action> Perform TD error calculation to obtain the TD error value for each time step, specifically:
[0057] 201) Process the state corresponding to the trajectory at each moment, flatten the input state into 1 dimension, and convert it into an integer type, then flatten the input action data into 1 dimension, and splice it behind the state data along the dimension 1 direction to form a new state-action pair input;
[0058] 202) Use a clean dataset to pre-train an agent with normal performance. The training process can use an algorithm different from the target task. First, based on the offline reinforcement learning algorithm, use the clean dataset to train the agent, and evaluate the performance of the training model in the test environment. According to the evaluation results, adjust the training parameters so that the average cumulative reward can reach the average level in existing research;
[0059] 203) inputting the new state-action pair value obtained in 201) into the Q network of the clean agent in 202), calculating the parameter value at time t, obtaining the state-action pair at time t+1 and calculating the parameter value;
[0060] 204) After that, the TD error is calculated for each time step in the dataset:
[0061] In step 3), all time steps in the data set are sorted according to the TD error value, and the time steps with larger TD errors are selected as key time steps, and a key time step set is formed, which is specifically:
[0062] 301) Sort the data set from large to small according to the TD error;
[0063] 302) Determine the number of key points according to the poisoning ratio p. For example, if the total number of data sets is N, the number of key points is Choose the one with larger error time steps as key time steps, and obtain the key time step set as the attack target for subsequent data poisoning.
[0064] In step 4), perturbations are added to the state-action pairs in the critical time step set. The perturbations are optimized by a dual-objective optimization algorithm to maximize the TD error while meeting the requirements of concealment. Specifically,
[0065] 401) Definition of dual-objective optimization problem: The optimization goal is to find the optimal perturbation for each state-action pair corresponding to each time step t in the critical time step set C, which satisfies two optimization goals at the same time: one is to minimize the perturbation amplitude to maintain the attack concealment; the other is to maximize the TD error to increase the impact of the attack. It can be formalized as a constrained dual-objective optimization problem
[0066] 402) Solution of dual-objective optimization problem: In the above optimization problem, limiting the disturbance within the range of ∈ is a boundary constraint problem. In the optimization solution process, it is necessary to ensure that the disturbance size does not exceed the constraint range. The value function prediction and discount factor calculation in the TD error calculation process involve nonlinear functions and belong to nonlinear constraints. In order to facilitate the solution, the weighted sum method is used to transform the above dual-objective optimization problem into a single-objective optimization problem:
[0067] 403) solves the dual-objective optimization problem in 402) and uses the Sequential Least Squares Programming (SLSQP) algorithm as the optimization solver to solve the objective function To solve.
[0068] In step 5), the optimized perturbation is added to the state-action pair of each data in the key time step set to obtain the poisoned data set and complete the entire poisoning process, specifically:
[0069] 501) Implementation of poisoning attack: Add the optimal perturbation obtained by solving to the state-action pair corresponding to the key time step in the original data set to generate a poisoned data set;
[0070] 502) training the agent to obtain a poisoned agent;
[0071] 503) Evaluate the degradation of the performance of poisoned agents relative to clean agents in a test environment, and quantify the negative impact of the attack on offline reinforcement learning agents.
[0072] Example 1
[0073] A covert data poisoning attack method for offline reinforcement learning includes the following steps:
[0074] Step 1: Get the state space S, action space A, and reward space R in the clean offline dataset D.
[0075] Step 2: Use the clean dataset to train a clean agent model, using the action-value function to represent each state-action pair (s t ,a t ) performs TD error calculation to obtain the TD error value of each time step;
[0076] Step 2.1: Transform the trajectory τ at each moment t =(s t ,a t ,r r ) to process the corresponding state and input the state data s t Flatten into 1 dimension and convert into integer type, then input action data a t Flatten into 1 dimension and splice behind the state data along dimension 1 to form a new state-action pair (s t ,a t )enter;
[0077] Step 2.2: Use a clean dataset to pre-train an agent with normal performance. The training process can use an algorithm different from the target task. First, based on the offline reinforcement learning algorithm, use the clean dataset D to train the agent, and evaluate the performance of the training model in the test environment. According to the evaluation results, adjust the training parameters so that the average cumulative reward can reach the average level in existing research;
[0078] Step 2.3: Input the new state-action pair value obtained in step 2.1 into the action value Q network parameter θ of the clean agent in step 2.2, and pass the state-action pair (s t ,a t ) Calculate Q θ (s t, a t ), get the state-action pair (s) at time t+1 t+1, a t+1 ) and calculate Q θ (s t+1, a t+1 );
[0079] Step 2.4: Calculate the TD error δ for each time step in the dataset t :
[0080] Q θ (s t ,a t )←Q θ (s t ,a t )+α·δ t
[0081]
[0082] Among them, α is the learning rate and γ is the discount factor.
[0083] Step 3: Sort all time steps in the dataset according to the TD error value, select the time step with the largest TD error as the key time step, and sort the dataset according to the TD error δ t Sort from large to small and select error δ according to the poisoning ratio p t The larger time step is taken as the key time step, and the key time step set C is obtained as the attack target of subsequent data poisoning and forms the key time step set C;
[0084] Step 3.1: Transform the dataset according to the TD error δ t Sort from largest to smallest;
[0085] Step 3.2: Determine the number of key points based on the poisoning ratio p. For example, if the total number of data sets is N, the number of key points is Select δ t Larger time steps as key time steps, and obtain the key time step set C as the attack target for subsequent data poisoning;
[0086] Step 4: For the state-action pairs (s in the critical time step set t ,a t ) Add disturbances, and the disturbances are optimized through a dual-objective optimization algorithm to maximize the TD error while meeting the requirements of concealment;
[0087] Step 4.1: For each state-action pair (s corresponding to each time step t in the key time step set C, t ,a t ), using a constrained bi-objective programming problem to find the optimal perturbation
[0088]
[0089] in, represents the state-action pair (s t ,a t )Add disturbance The L2 norm of is used to quantify the magnitude of the disturbance; δ t ' is the TD error after adding disturbance, ∈ is the upper limit of the disturbance amplitude, which is used to limit the magnitude of the disturbance amplitude;
[0090] Step 4.2: Solving the dual-objective optimization problem: In the above optimization problem, the perturbation Limiting to the range of ∈ is a boundary constraint problem. During the optimization process, it is necessary to ensure that the disturbance size does not exceed the constraint range. TD error δ t The prediction of the value function Q(s,a) and the calculation of the discount factor γ in the calculation process involve nonlinear functions and belong to nonlinear constraints. To facilitate the solution, the weighted sum method is used to transform the above dual-objective optimization problem into a single-objective optimization problem:
[0091]
[0092] Among them, β 1 and β 2 is a weight parameter used to balance the importance of the two objectives. In this method, the two objectives have the same importance. 1 and β 2 Take 1 for both;
[0093] Step 4.3: Solve the dual-objective optimization problem in step 4.2, using the Sequential Least Squares Programming (SLSQP) algorithm as the optimization solver for the objective function To solve.
[0094] Step 5: The optimized perturbation η t The state-action pair (s) of each data added to the key time step set t ,a t ) to obtain the poisoned data set D' and complete the entire poisoning process;
[0095] Step 5.1: Obtain the state-action pairs (s) at each critical time step t t ,a t );
[0096] Step 5.2: Solve the optimal perturbation The state-action pairs (s) corresponding to the key time steps added to the original dataset D t ,a t ) to generate a poisoned data set D';
[0097] Step 5.3: Use D' to train the agent and obtain the poisoned agent A';
[0098] Step 5.4: Evaluate the degradation of the performance of the poisoned agent A' relative to the clean agent A in the test environment to quantify the negative impact of the attack on the offline reinforcement learning agent.
[0099] In particular, in some preferred embodiments of the present invention, a computer device is also provided, including a memory and a processor and a computer program stored in the memory, and when the processor executes the computer program, the steps of the covert data poisoning attack method for offline reinforcement learning described in any of the above embodiments are implemented.
[0100] In some other preferred embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program / instructions are stored. When the computer program is executed by a processor, the steps of the covert data poisoning attack method for offline reinforcement learning described in any of the above embodiments are implemented.
[0101] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiment method can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the process of the embodiment of the covert data poisoning attack method for offline reinforcement learning, which will not be repeated here.
[0102] The hidden data poisoning method for offline reinforcement learning proposed in this patent has a good attack effect, has a great impact on different offline reinforcement learning algorithms and different reinforcement learning tasks, and has high attack efficiency and is relatively hidden. The attack method proposed in this invention can achieve an average decrease of 84% in the performance of the intelligent agent trained with poisoned data compared to the performance of the intelligent agent trained with clean data sets when the poisoning ratio is only 1%; when the poisoning ratio is 5%, the performance of the intelligent agent decreases by 90%.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A hidden data poisoning attack method for offline reinforcement learning, characterized by: The following steps are involved: Step 1: Get the state space S, action space A and reward space R in the clean offline dataset D; Step 2: Use the clean dataset to train a clean agent model, using the action-value function to represent each state-action pair (s t ,a t ) performs TD error calculation to obtain the TD error value of each time step; Step 3: Sort all time steps in the data set according to the TD error value, select the time steps with larger TD errors as key time steps, and form a key time step set C; Step 4: For the state-action pairs (s in the critical time step set t ,a t ) Add a disturbance η t , the perturbation is carried out by a dual-objective optimization algorithm to maximize the TD error while satisfying the concealment requirement; Step 5: Optimize the perturbation The state-action pair (s) of each data added to the key time step set t ,a t ) to obtain the poisoned data set D' and complete the entire poisoning process.
2. The hidden data poisoning attack method for offline reinforcement learning according to claim 1, characterized in that: The step 2 specifically includes the following steps: Step 2.1: Transform the trajectory τ at each moment t =(s t ,a t ,r r ) to process the corresponding state and input the state data s t Flatten into 1 dimension and convert into integer type, then input action data a t Flatten into 1 dimension and splice behind the state data along dimension 1 to form a new state-action pair (s t ,a t )enter; Step 2.2: Use a clean dataset to pre-train an agent with normal performance. First, based on the offline reinforcement learning algorithm, use the clean dataset D to train the agent, and evaluate the performance of the training model in the test environment. According to the evaluation results, adjust the training parameters so that the average cumulative reward can reach the average level in existing research. Step 2.3: Transform the new state-action pair (s t ,a t ) is input into the action value Q network of the clean agent to calculate the state-action pair Q at time t θ (s t, a t ), get the state-action pair (s) at time t+1 t+1, a t+1 ) and calculate Q θ (s t+1, a t+1 ); Step 2.4: Calculate the TD error δ for each time step in the dataset t : Q θ (s t ,a t )←Q θ (s t ,a t )+a·d t Among them, α is the learning rate and γ is the discount factor.
3. The hidden data poisoning attack method for offline reinforcement learning according to claim 1, characterized in that: The key time step set C of step 3 is: t Sort from large to small and select error δ according to the poisoning ratio p t The larger time step is taken as the key time step, and the key time step set C is obtained as the target of subsequent data poisoning attacks.
4. The hidden data poisoning attack method for offline reinforcement learning according to claim 1, characterized in that: The step 4 specifically comprises the following steps: Step 4.1: For each state-action pair (s corresponding to each time step t in the key time step set C, t ,a t ), using a constrained bi-objective programming problem to find the optimal perturbation in, represents the state-action pair (s t ,a t )Add disturbance The L2 norm of δ t ' is the TD error after adding disturbance, ∈ is the upper limit of the disturbance amplitude; Step 4.2: Use the weighted sum method to transform the above dual-objective optimization problem into a single-objective optimization problem: Among them, β1 and β2 are weight parameters; Step 4.3: Solve the bi-objective optimization problem and use the sequential least squares programming algorithm as the optimization solver to solve the objective function To solve.
5. The hidden data poisoning attack method for offline reinforcement learning according to claim 1, characterized in that: The step 5 specifically comprises the following steps: Step 5.1: Obtain the state-action pairs (s t ,a t ); Step 5.2: Solve the optimal perturbation The state-action pairs (s) corresponding to the key time steps added to the original dataset D t ,a t ) to generate a poisoned data set D'; Step 5.3: Use D' to train the agent and obtain the poisoned agent A'; Step 5.4: Evaluate the degradation of the performance of the poisoned agent A' relative to the clean agent A in the test environment to quantify the negative impact of the attack on the offline reinforcement learning agent.
6. A covert data poisoning attack system for offline reinforcement learning, characterized in that: It includes a module for locating key time steps and a poisoning attack module based on dual-objective optimization; The positioning key time step module includes a TD error calculation module and a key time step extraction module; the TD error calculation module is used to perform TD error analysis on each state-action pair in the data set and sort the time steps by the TD error size; the key time step extraction module is used to filter the time steps according to the TD error size, extract the time steps that have the greatest impact on model training, and generate a key time step set; The poisoning attack module based on dual-objective optimization includes a dual-objective optimization module, a disturbance generation module and a data replacement module; the dual-objective optimization module is used to determine the optimal disturbance amplitude through an optimization algorithm to make the attack both effective and concealed; the disturbance module is used to generate the optimal disturbance according to the result of the dual-objective optimization and add it to the state-action pair in the key time step to generate poisoned data; the data replacement module is used to replace the corresponding time step in the original data with the poisoned data and generate the final poisoned data set.
7. A computer device / equipment / system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.