Poison attack method, system, program and equipment for offline reinforcement learning decision mode diversity and storage medium
The key decision sequences in offline reinforcement learning are identified through feature encoder and clustering methods, and the perturbation of hidden constraints is added, which solves the problem of insufficient data poisoning attack efficiency and concealment in the prior art, and achieves an efficient and concealed attack effect.
Patent Information
- Application Number
- CN202510064156.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology is difficult to effectively target data poisoning attacks in offline reinforcement learning, especially in terms of reducing the cost of attacks and increasing the efficiency of attacks while increasing the concealment of attacks.
Advanced representations are extracted through feature encoder, decision-making patterns are identified using clustering methods, and key decision sequences are determined based on the poisoning ratio, and perturbations are added to the hidden constraints, so that the poisoned key decision sequences may appear or even disappear in the data set, thereby reducing the diversity of decision-making patterns.
With less attack cost, the attack efficiency is significantly improved, and the attack concealment is improved by dynamically adding slight perturbations, so that when the poisoning ratio is low, the performance of the agent will drop by more than 83%.
Smart Images

Figure CN120031098A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of offline reinforcement learning, and specifically relates to a poisoning attack method, system, program, device and storage medium for the diversity of offline reinforcement learning decision modes. Background Art
[0002] Offline reinforcement learning, also known as batch reinforcement learning, is a paradigm in which a learning agent learns from a previously collected set of experience data. Compared to online reinforcement learning, which requires obtaining feedback from the environment in real time to improve the strategy, offline RL can learn without online interaction with the environment. It is suitable for exploring expensive, time-consuming or risky scenarios, such as autonomous driving, healthcare decision-making, intelligent robot control, and game design.
[0003] Data poisoning attacks have been widely used in the field of machine learning. In the early days, they mainly targeted basic models such as logistic regression and support vector machines (Marco Barreno, Blaine Nelson, Anthony D. Joseph, and JD Tygar. The security of machine learning. Mach. Learn., 81 (2): 121–148, November 2010. ISSN 0885-6125. doi: 10.1007 / s10994-010-5188-5.). In recent years, with the development of deep learning, attackers have gradually turned to attacking deep networks (Feng, J., Cai, Q.-Z., and Zhou, Z.-H. Learning to confuse: generating training time adversarial data with autoencoder. Advances in Neural Information Processing Systems, 32, 2019.). These works show that data security has become an issue that cannot be ignored. As an important branch of the current machine learning field, RL also has data security issues. At present, there are many works on data poisoning attacks on online RL, and they are quite in-depth (Rakhsha, A., Zhang, X., Zhu, X., and Singla, A. Reward poisoning in reinforcement learning: Attacks against unknown learners in unknown environments. arXiv preprint arXiv:2102.08492, 2021.)
[0004] Online RL has the problem of data poisoning, and offline RL also faces such a threat. The above work on data poisoning for online RL provides many ideas for the research of offline RL. In 2019, Ma (Ma, Y., Zhang, X., Sun, W., and Zhu, J. Policy poisoning in batch reinforcement learning and control. Advances in Neural Information Processing Systems, 2019.) attacked the reinforcement learning training dataset for the first time, taking batch reinforcement learning and controller as victims, and attacking at each time step to affect the performance of the agent, but the cost of attacking each time step is high; in 2020, Rakhsha (Rakhsha, A., Radanovic, G., Devidze, R., Zhu, X., and Singla, A. Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning. In International Conference on Machine Learning, pp.7974–7984.PMLR, 2020.) converts the attack into an optimization problem and finds the optimal attack for different attack costs, but the proposed attack does not target more realistic continuous tasks, and the attack modifies the reward and state transfer functions at the same time, increasing the attack cost; in 2022, Gong (Gong, C., Yang, Z., Bai, Y., He, J., Shi, J., Sinha, A., Xu, B., Hou, X., Fan, G., and Lo, D. Mind your data! Hiding backdoors in offline reinforcement learning datasets. arXiv preprint arXiv:2210.04688,2022.) proposed a backdoor attack against offline reinforcement learning, which controls the attack by adding triggers, but the poisoning method is random poisoning, and the poisoning ratio needs to reach 10% to have a good attack effect. The cost is high, and the backdoor attack needs to be continuously operated during the model training and testing phases to be effective, which has high requirements for the attacker; In 2022, Wu (Wu, F., Li, L., Xu, C., Zhang, H., Kailkhura, B., Kenthapadi, K., Zhao, D., and Li, B.Copa: Certifying robust policies for offline reinforcement learning against poisoning attacks.arXiv preprint arXiv:2203.08398,2022.) proposed a certification framework for the robustness of offline reinforcement learning strategies, which can judge the robustness of the algorithm; in 2022, Rang (Rangi, A., Xu, H., Tran-Thanh, L., and Franceschetti, M. Understanding the limits of poisoning attacks in epiisodic reinforcement learning.arXiv preprint arXiv:2208.13663,2022.) proved the limits of poisoning attacks in reinforcement learning (both offline and online) through pure theory, but the existing work lacks analysis of offline data sets and analysis of the impact of offline algorithms and models. It is necessary to study low-cost and efficient poisoning attacks and deeply analyze the impact of attacks. . Summary of the invention
[0005] The purpose of the present invention is to provide a data poisoning method for offline reinforcement learning, which can increase the concealment of the attack while improving the attack efficiency by adding small disturbances to the points with greater influence in the learning process.
[0006] The present invention provides a poisoning attack method for the diversity of offline reinforcement learning decision modes, comprising the following steps:
[0007] Step 1: Get the state space S, action space A and reward space R in the clean offline dataset D;
[0008] Step 2: Use the feature encoder to encode each state-action pair (s t ,a t ) extract high-level representations and use clustering methods to spatially partition these representations to identify different decision-making patterns;
[0009] Step 3: From multiple consecutive time steps of state-action pairs (s t ,a t ) to extract the core decision sequence Get the core decision sequence set According to the poisoning ratio p, the core decision sequence with the least number of occurrences in the data set is determined as the key decision sequence, and a key decision sequence set M is formed;
[0010] Step 4: Determine the perturbation ξ subject to the hiddenness constraint t , so that the poisoned key decision sequence disappears in the key decision sequence set M;
[0011] Step 5: Set the constrained perturbation ξ t The state-action pair (s) of each data added to the key decision sequence set M t ,a t ) to obtain the poisoned data set D' and complete the entire poisoning process.
[0012] Furthermore, the step 2 specifically includes the following steps:
[0013] Step 2.1: Transform the trajectory τ at each moment t =(s t ,a t ,r t ) corresponding to the state, flatten the input state into 1 dimension and convert it into an integer type, then flatten the input action data into 1 dimension and splice it behind the state data along dimension 1 to form a new state-action pair (s t ,a t )enter;
[0014] Step 2.2: Transform the new state-action pair (s t ,a t ) is input into the feature encoder to obtain the output feature array, which is the encoding result of the state-action pair in the data set. The encoding result is expressed as F = f 1 ,f 2 ,...,f T ;
[0015] Step 2.3: Use the elbow rule to calculate the optimal k value for all original state-action pairs in the data set, which will be used as the number of clusters in the subsequent k-means clustering step;
[0016] Step 2.4: According to the k value obtained by the solution, randomly select k cluster centers from the encoding result F, and then use the k-means method to convert F = f 1 ,f 2 ,...,f T All encoding results in are divided into k clusters, and a cluster label c is assigned to each state-action pair at each moment t , where t∈{1,2,…,K}.
[0017] Furthermore, in the step 2.2, the feature encoder acquisition method is to train an intelligent agent model with normal behavior using a clean offline reinforcement learning data set to obtain model parameters of the model.
[0018] Furthermore, the step 3 specifically includes the following steps:
[0019] Step 3.1: Multiple state-action pairs (s t ,a t ) is a sequence composed of i =(s i ,a i ),(s i+1 ,a i+1 ),...,(s i+l ,a i+l ); After the space is divided by clustering method, the state-action pair (s t ,a t ) are assigned a cluster label c t , then the sequence is T i =c i ,c i+1 ,...,c i+l ;
[0020] Step 3.2: After clustering, i Remove consecutive repeated labels to obtain the core decision sequence
[0021] Step 3.3: Core Decision Sequence The number of times it occurs in the entire dataset All core decision sequences Decision sequence set
[0022] Step 3.4: Determine the number of key points Z based on the poisoning ratio p and the total number of data sets N.
[0023] Step 3.5: Sequence the new data set into the core decision sequence Number of occurrences The values of are arranged from small to large, and the indexes of the key sequences that meet the number are obtained from small to large according to the number of key points, and the data corresponding to all the selected indexes constitute the key sequence set M.
[0024] Furthermore, the step 4 specifically includes the following steps:
[0025] Step 4.1: Set the disturbance ratio η and get the upper limit of disturbance ε:
[0026] ε=(s t ,a t )×η
[0027] Step 4.2: For each critical time step, randomly generate multiple sets of state-action pairs (st ,a t ) The perturbation factor with the same shape The infinite norm n of the perturbation factor is less than the upper perturbation limit;
[0028] Step 4.3: Add multiple sets of perturbation factors to the original sequence T before deduplication i , and perform deduplication to generate multiple sets of potential poisoning sequences
[0029] Step 4.4: Calculate the counts of multiple sets of potential poisoning sequences in the data set, and select the ones that can make the original decision sequence count The maximum increase in poisoning time replaces the original time T i , determine the corresponding optimal perturbation factor.
[0030] The present invention also provides a hidden data poisoning attack system for offline reinforcement learning, a module for locating key decision sequences and a module for reducing decision mode diversity;
[0031] The key decision sequence positioning module includes a state-action space division module, a decision sequence extraction module and a decision sequence counting module; the state-action space division module is used to divide the state-action space in the data set into different clusters; the decision sequence extraction module is used to extract the core decision sequence from the sequence composed of continuous state-action pairs; the decision sequence counting module calculates the number of times all decision sequences appear in the data set, and obtains a rare decision sequence set according to the poisoning ratio; the decision pattern diversity reduction module is used to add the disturbance under the constraint to the state-action pair of each data in the key decision sequence set, thereby reducing the diversity of decision patterns in the data set, obtaining a poisoned data set, and completing poisoning.
[0032] The present invention also provides a computer device / equipment / system, comprising a memory, a processor, and a computer program stored in the memory, wherein when the processor executes the computer program, the steps of any of the above-mentioned poisoning attack methods for the diversity of offline reinforcement learning decision models are implemented.
[0033] The present invention also provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned poisoning attack methods for the diversity of offline reinforcement learning decision modes.
[0034] The present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of any of the above-mentioned poisoning attack methods for the diversity of offline reinforcement learning decision models.
[0035] The beneficial effects of the present invention are:
[0036] 1. The present invention proposes a data poisoning method for offline reinforcement learning. For the first time, starting from the sequence level, it analyzes the impact of decision pattern diversity on offline reinforcement learning. First, in the offline reinforcement learning process, the decision pattern sequence that has a significant impact on the learning task is determined, and the key decision sequence is poisoned, which greatly reduces the diversity of decision patterns in the data set and improves the attack efficiency while using less attack cost. The present invention is also applicable to different offline reinforcement learning algorithms and different reinforcement learning tasks. The attack method proposed in the present invention has been verified in four offline reinforcement learning algorithms: Batch-Constrained Q-learning (BCQ), Batch-Ensemble Actor-Critic with Retrace (BEAR), Conservative Q-Learning (CQL) and BC (Behavioural Cloning); Walker2D, Hopper and Half-Cheetah in the MuJoCo robot simulator, and Carla-Lane autonomous driving in the Carla simulator. When the poisoning ratio is only 1%, the performance of the intelligent agent trained with poisoned data is reduced by more than 83% on average compared with the performance of the intelligent agent trained with clean data sets; when the poisoning ratio is 5%, the performance of the intelligent agent is reduced by more than 90%.
[0037] 2. The perturbation method proposed in the present invention can dynamically add tiny perturbations that do not exceed a certain proportion of the original data in the data set according to the size of the original data. Compared with the original data, the change is very small and difficult to detect, which improves the concealment of the attack. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 An overview of the decision-making model diversity poisoning method for offline reinforcement learning provided by the present invention;
[0039] Figure 2 Diagram of the feature encoder model. DETAILED DESCRIPTION
[0040] The present invention is further described below in conjunction with the accompanying drawings.
[0041] The present invention discloses a poisoning attack system for offline reinforcement learning decision mode diversity, including a key decision sequence positioning module and a decision mode diversity reduction module;
[0042] Positioning key decision sequence module: This module is responsible for identifying different decision patterns from continuous high-dimensional offline RL datasets and determining which decision patterns have a greater impact on the decision diversity in the entire dataset. Its function is to find the key decision sequences in the offline reinforcement learning dataset and reduce the attack cost and improve the attack efficiency by attacking the key sequences. The core parts of this module are feature encoder, state-action space partition module, decision sequence extraction module and decision sequence counting module.
[0043] Feature Encoder: This module uses the trained clean agent as an encoder to extract features from the input data. Its function is to convert the data in the offline dataset into model data that the agent can understand.
[0044] State-action space partitioning module: This module is responsible for partitioning all state-action spaces in the data set into different clusters. Its function is to provide a basis for extracting decision patterns.
[0045] Decision sequence extraction module: This module is responsible for extracting the core decision sequence from the sequence of continuous state-action pairs. Its function is to extract the core decision change pattern of these behavior sequences and remove redundant information.
[0046] Decision sequence counting module: Counts the number of times all decision sequences appear in the data set, and obtains a set of rare decision sequences based on the poisoning ratio. Its function is to measure the prevalence or rarity of decision sequences in the data set.
[0047] Decision mode diversity reduction module: This module is responsible for adding disturbances to the state-action pairs in the key time step. Its function is to make small modifications to the data in the offline dataset to reduce the diversity of the decision mode in the dataset, thereby poisoning the data. The core part of this module is the disturbance generator.
[0048] Perturbation Generator: This module obtains the state-action value in the key time step and adds a perturbation that is smaller than a certain multiple of the value to the data, generating multiple sets of potentially toxic data. It then selects the data that can maximize the increase in the decision sequence count as toxic data, thereby modifying the data set.
[0049] like Figure 1 As shown, the present invention further discloses a data poisoning method for offline reinforcement learning, comprising the following steps:
[0050] 1) Obtain the state space, action space, and reward space in a clean offline dataset;
[0051] 2) Use the feature encoder to decode the state-action pair at each time step<state,action> Extract high-level representations from the dataset and use clustering methods to spatially partition these representations to identify different decision-making patterns;
[0052] 3) Extract decision sequences from state-action pairs of multiple consecutive time steps to obtain a set of decision sequences. According to the poisoning ratio, the decision sequence with the least number of occurrences in the data set is determined as the key decision sequence, and the key decision sequence set is formed;
[0053] 4) By adding constrained perturbations to the key decision sequences, the goal of reducing the number of occurrences of poisoned key decision sequences in the data set or even eliminating them is achieved, thereby reducing the diversity of decision patterns in the offline data set;
[0054] 5) Add the perturbation under the constraint to each state-action pair of data in the key decision sequence set<state,action> On the , obtain the poisoned data set and complete the entire poisoning process.
[0055] In step 2), the state-action pair corresponding to the trajectory at each moment is encoded into the feature encoder, and the features extracted after encoding are spatially divided, specifically:
[0056] 201) Process the state corresponding to the trajectory at each moment, flatten the input state into 1 dimension, and convert it into an integer type, then flatten the input action data into 1 dimension, and splice it behind the state data along the dimension 1 direction to form a new state-action pair input;
[0057] 202) Use a clean offline reinforcement learning dataset to train an intelligent agent model with normal behavior, obtain the model parameters of the model, and use them as feature encoders, such as Figure 2 As shown;
[0058] 203) inputting the new state-action pair value obtained in 201) into the encoder trained in 202) to obtain an output feature array, which is the encoding result of the state-action pair in the data set;
[0059] 204) Using the elbow rule, the optimal k value is calculated for all original state-action pairs in the data set, which is used as the number of clusters in the subsequent k-means clustering step;
[0060] 205) According to the k value obtained by the solution, k cluster centers are randomly selected from the encoding results, and then all the encoding results are divided into k clusters according to the k-means method, which means that the features of all state-action pairs in the data set are divided into k different spaces, and cluster labels are assigned to the state-action pairs at each moment;
[0061] In step 3), decision sequences are extracted from the state-action pairs of multiple consecutive time steps to obtain a decision sequence set. The decision sequence with the least number of occurrences in the data set is determined as the key decision sequence according to the poisoning ratio, and a key decision sequence set is formed, which is specifically:
[0062] 301) After the sequence of state-action pairs of multiple consecutive time steps has been divided into two parts in 205), each state-action pair corresponding to each time step is assigned a cluster label;
[0063] 302) For T at this time i Remove the consecutive repeated labels and get the core decision mode of the sequence, which is expressed as a decision sequence
[0064] 303) Decision sequence The number of times it appears in the entire data set is counted, and the set of all decision sequences in the data set is called the decision sequence set;
[0065] 304) Determine the number of key points according to the poisoning ratio. For example, if the total number of data sets is 1 million, when the poisoning ratio is 1%, the number of key points is 10,000;
[0066] 305) Arrange the new data set from small to large according to the value of , and obtain the index of the key sequence that meets the number from small to large according to the number of key points determined in 304), and form a key sequence set with the data corresponding to all the selected indexes.
[0067] In step 4), a perturbation constrained by concealment is added to the key decision sequence, so that the poisoned key decision sequence appears less frequently or even disappears in the data set, thereby reducing the diversity of decision patterns in the offline data set. Specifically,
[0068] 401) Set the disturbance ratio to η, and calculate the disturbance upper limit ε for each key time step. The disturbance upper limit is the state-action pair multiplied by the disturbance ratio
[0069] 402)402)For each critical time step, randomly generate n groups of perturbation factors that are consistent with the shape of the state-action pair, and the infinite norm of the perturbation factor is less than the perturbation upper limit;
[0070] 403) adding the multiple groups of perturbation factors generated in 402) to the original sequence before deduplication, and performing the deduplication process in 302) to generate multiple groups of potential poisoning sequences;
[0071] 404) Calculate the counts of the above multiple sets of potential poisoning sequences in the data set, and select the ones that can make the original decision sequence count C TiThe poisoned time series with the largest increase replaces the original time series, and the corresponding optimal perturbation factor is determined.
[0072] In step 5), the perturbation under the constraint is added to the state-action pair of each data in the key decision sequence set to obtain the poisoned data set and complete the entire poisoning process, specifically:
[0073] 501) first obtain the key decision sequence index corresponding to step 305), and obtain the value of the state-action pair in the corresponding key decision sequence;
[0074] 502) For each key decision sequence, generate a disturbance factor according to the method in step 4;
[0075] 503) The disturbance factor generated in step 502) is added to the state-action pair in each key decision in the key decision sequence set to obtain a poisoned data set and implement a data poisoning attack.
[0076] Example 1
[0077] This embodiment discloses a poisoning attack method for the diversity of offline reinforcement learning decision modes, including the following steps:
[0078] Step 1: Get the state space S, action space A and reward space R in the clean offline dataset D;
[0079] Step 2: Use the feature encoder to encode each state-action pair (s t ,a t ) extracts high-level representations and uses clustering methods to spatially partition these representations to identify different decision-making modes; specifically:
[0080] Step 2.1: Transform the trajectory τ at each moment t =(s t ,a t ,r t ) to process the corresponding state, flatten the input state into 1 dimension and convert it into an integer type, then flatten the input action data into 1 dimension and splice it behind the state data along the dimension 1 direction;
[0081] Step 2.2: Use a clean offline reinforcement learning dataset to train a normal-behaving agent model, obtain the model parameters of the model, and use them as the feature encoder;
[0082] Step 2.3: Input the new state-action pair value in step 2.1 into the encoder obtained in step 2.2 to obtain the output feature array, which is the encoding result of the state-action pair in the data set. The encoding result is expressed as F = f 1 ,f 2,...,f T ;
[0083] Step 2.4: Use the elbow rule to calculate the optimal k value for all original state-action pairs in the data set, which will be used as the number of clusters in the subsequent k-means clustering step;
[0084] Step 2.5: According to the k value obtained by the solution, randomly select k cluster centers from the encoding result F, and then use the k-means method to convert F = f 1 ,f 2 ,...,f T All encoding results in are divided into k clusters, and a cluster label c is assigned to each state-action pair at each moment t , where t∈{1,2,…,K}.
[0085] Step 3: From multiple consecutive time steps of state-action pairs (s t ,a t ) to extract the decision sequence Get a set of decision sequences According to the poisoning ratio p, the decision sequence with the least number of occurrences in the data set is determined as the key decision sequence, and a key decision sequence set M is formed; specifically:
[0086] Step 3.1: Multiple state-action pairs (s t ,a t ) is a sequence composed of i =(s i ,a i ),(s i+1 ,a i+1 ),...,(s i+l ,a i+l ); After the space is divided by clustering method, the state-action pair (s t ,a t ) are assigned a cluster label c t , then the sequence is T i =c i ,c i+1 ,...,c i+l ;
[0087] Step 3.2: T i Remove the consecutive repeated labels and get the core decision mode of the sequence, which is expressed as a decision sequence
[0088] Step 3.3: Decision sequence The number of occurrences in the entire data set is counted and expressed as And all decision sequences in the data set The set composed of is called a decision sequence set, expressed as
[0089] Step 3.4: Determine the number of key points Z based on the poisoning ratio p and the total number of data sets N. For example, if the total number of data sets is 1 million, when the poisoning ratio is 1%, the number of key points is 10,000;
[0090] Step 3.5: Transform the new dataset into The values of are arranged from small to large. According to the number of key points determined in step 3.4, the indexes of the key sequences that meet the number are obtained from small to large, and the data corresponding to all the selected indexes constitute the key sequence set M.
[0091] Step 4: Add a perturbation ζ to the key decision sequence subject to the hiddenness constraint t , generating poisoned state-action pairs (s t ',a t '), specifically:
[0092] Step 4.1: Set the perturbation ratio to 0.05 and calculate the perturbation upper limit ε for each key time step. The perturbation upper limit is the state-action pair multiplied by the perturbation ratio.
[0093] ε=(s t ,a t )×η
[0094] For example, η = 5% means that the magnitude of the realized disturbance value does not exceed 5% of the magnitude of the original state-action pair value itself;
[0095] Step 4.2: For each critical time step, randomly generate multiple sets of state-action pairs (s t ,a t ) The perturbation factor with the same shape Where n = 50, and the infinite norm of the perturbation factor is less than the upper perturbation limit;
[0096] Step 4.3: Add the multiple sets of perturbation factors generated in step 4.2 to the original sequence T before deduplication i and perform the deduplication step in step 3.2 to generate multiple sets of potential poisoning sequences
[0097] Step 4.4: Calculate the counts of the above multiple sets of potential poisoning sequences in the data set, and select the ones that can make the original decision sequence count The maximum increase in poisoning time replaces the original time T i , determine the corresponding optimal perturbation factor.
[0098] Step 5: Add the perturbation under the constraint to each state-action pair (s t ,a t ) to obtain the poisoned data set and complete the entire poisoning process; specifically:
[0099] Step 5.1: First, obtain the key decision sequence index corresponding to step 3.5, and obtain the value of the state-action pair in the corresponding key decision sequence;
[0100] Step 5.2: For each critical decision sequence, generate a perturbation factor according to the method in step 4;
[0101] Step 5.3: Add the perturbation factor generated in step 5.2 to the state-action pair in each key decision in the key decision sequence set M to obtain a poisoned data set and implement a data poisoning attack.
[0102] In particular, in some preferred embodiments of the present invention, a computer device is also provided, including a memory and a processor and a computer program stored in the memory, and when the processor executes the computer program, the steps of the poisoning attack method for the diversity of offline reinforcement learning decision patterns described in any of the above embodiments are implemented.
[0103] In other preferred embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program / instructions are stored. When the computer program is executed by a processor, the steps of the poisoning attack method for the diversity of offline reinforcement learning decision modes described in any of the above embodiments are implemented.
[0104] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiment method can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the process of the above-mentioned poisoning attack method embodiment for the diversity of offline reinforcement learning decision models, which will not be repeated here.
[0105] In general, the data poisoning method for offline reinforcement learning proposed in this application has a good attack effect. The method fully reflects the negative impact of insufficient diversity of decision patterns in sequence-level data sets on offline reinforcement learning models. The method has a great impact on different offline reinforcement learning algorithms and different reinforcement learning tasks. The method is highly efficient and relatively hidden.
[0106] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A poisoning attack method targeting the diversity of offline reinforcement learning decision-making models, characterized by: The following steps are involved: Step 1: Get the state space S, action space A and reward space R in the clean offline dataset D; Step 2: Use the feature encoder to encode each state-action pair (s t ,a t ) extract high-level representations and use clustering methods to spatially partition these representations to identify different decision-making patterns; Step 3: From multiple consecutive time steps of state-action pairs (s t ,a t ) to extract the core decision sequence Get the core decision sequence set According to the poisoning ratio p, the core decision sequence with the least number of occurrences in the data set is determined as the key decision sequence, and a key decision sequence set M is formed; Step 4: Determine the perturbation ξ subject to the hiddenness constraint t , so that the poisoned key decision sequence disappears in the key decision sequence set M; Step 5: Set the constrained perturbation ξ t The state-action pair (s) of each data added to the key decision sequence set M t ,a t ) to obtain the poisoned data set D' and complete the entire poisoning process.
2. The poisoning attack method for diversity of offline reinforcement learning decision-making modes according to claim 1, characterized in that: The step 2 specifically includes the following steps: Step 2.1: Transform the trajectory τ at each moment t =(s t ,a t ,r t ) corresponding to the state, flatten the input state into 1 dimension and convert it into an integer type, then flatten the input action data into 1 dimension and splice it behind the state data along dimension 1 to form a new state-action pair (s t ,a t )enter; Step 2.2: Transform the new state-action pair (s t ,a t ) is input into the feature encoder to obtain the output feature array, which is the encoding result of the state-action pair in the data set. The encoding result is expressed as F = f1, f2, ..., f T ; Step 2.3: Use the elbow rule to calculate the optimal k value for all original state-action pairs in the data set, which will be used as the number of clusters in the subsequent k-means clustering step; Step 2.4: According to the k value obtained by the solution, randomly select k cluster centers from the encoding result F, and then use the k-means method to transform F = f1, f2, ..., f T All encoding results in are divided into k clusters, and a cluster label c is assigned to each state-action pair at each moment t , where t∈{1,2,…,K}.
3. The hidden data poisoning attack method for offline reinforcement learning according to claim 2 is characterized in that: In the step 2.2, the feature encoder acquisition method is to train an intelligent agent model with normal behavior using a clean offline reinforcement learning data set to obtain model parameters of the model.
4. The poisoning attack method for diversity of offline reinforcement learning decision-making modes according to claim 2, characterized in that: The step 3 specifically includes the following steps: Step 3.1: Multiple state-action pairs (s t ,a t ) is a sequence composed of i =(s i ,a i ),(s i+1 ,a i+1 ),...,(s i+l ,a i+l ); After the space is divided by clustering method, the state-action pair (s t ,a t ) are assigned a cluster label c t , then the sequence is T i =c i ,c i+1 ,...,c i+l ; Step 3.2: After clustering, i Remove consecutive repeated labels to obtain the core decision sequence Step 3.3: Core Decision Sequence The number of times it appears in the entire dataset All core decision sequences Core decision sequence set Step 3.4: Determine the number of key points Z based on the poisoning ratio p and the total number of data sets N. Step 3.5: Sequence the new data set into the core decision sequence Number of occurrences The values of are arranged from small to large, and the indexes of the key sequences that meet the number are obtained from small to large according to the number of key points, and the data corresponding to all the selected indexes constitute the key sequence set M.
5. The poisoning attack method for the diversity of offline reinforcement learning decision-making modes according to claim 4 is characterized in that: The step 4 specifically comprises the following steps: Step 4.1: Set the disturbance ratio η and get the upper limit of disturbance ε: ε=(s t ,a t )×h Step 4.2: For each critical time step, randomly generate multiple sets of state-action pairs (s t ,a t ) The perturbation factor with the same shape The infinite norm n of the perturbation factor is less than the upper perturbation limit; Step 4.3: Add multiple sets of perturbation factors to the original sequence T before deduplication i , and perform deduplication to generate multiple sets of potential poisoning sequences Step 4.4: Calculate the counts of multiple sets of potential poisoning sequences in the data set, and select the ones that can make the original decision sequence count The maximum increase in poisoning time replaces the original time T i , determine the corresponding optimal perturbation factor.
6. A covert data poisoning attack system for offline reinforcement learning, characterized by: Locate key decision sequence modules and reduce decision mode diversity modules; The key decision sequence positioning module includes a state-action space division module, a decision sequence extraction module and a decision sequence counting module; the state-action space division module is used to divide the state-action space in the data set into different clusters; the decision sequence extraction module is used to extract the core decision sequence from the sequence composed of continuous state-action pairs; the decision sequence counting module calculates the number of times all decision sequences appear in the data set, and obtains a rare decision sequence set according to the poisoning ratio; the decision pattern diversity reduction module is used to add the disturbance under the constraint to the state-action pair of each data in the key decision sequence set, thereby reducing the diversity of decision patterns in the data set, obtaining a poisoned data set, and completing poisoning.
7. A computer device / equipment / system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.