Anti-robustness improvement method for agent target-oriented reinforcement learning

By introducing semi-contrasting characterization attacks and adversarial representation strategies in agent goal-oriented reinforcement learning, the shortcomings of improving the robustness of adversarial attacks in the existing technology are solved, and the agent's stronger robustness and adaptability in complex environments are achieved.

CN120197648APending Publication Date: 2025-06-24INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510155048.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has shortcomings in improving the robustness of adversarial attacks in reinforcement learning algorithms, especially in goal-oriented reinforcement learning. Traditional methods rely on pseudo-labels and critic networks, lack flexibility and face multiple limitations in practical applications.

Method used

A method of adversarial robustness improvement for target-oriented reinforcement learning of agents is proposed. Through the semi-contraspectral characterization attack mechanism and a general adversarial characterization strategy defense framework, the performance and security of agents in complex environments are enhanced. The semi-contraspectral characterization attack maximizes the representation distance between the original input tuple and the negative sample, and enhances robustness independently of the evaluation function; the adversarial characterization strategy improves the agent's ability to resist adversarial attacks by optimizing the weights of the encoder, actor network and evaluater network.

Benefits of technology

It effectively improves the adversarial robustness of the agent, allowing it to have stronger adaptability when facing potential extreme situations, and significantly improves its application performance in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197648A_ABST
    Figure CN120197648A_ABST
Patent Text Reader

Abstract

The invention discloses an adversarial robustness improvement method for agent target-oriented reinforcement learning. The adversarial robustness improvement method comprises the following steps: 1) collecting a group of training data from interaction of a target condition reinforcement learning agent and an environment; wherein each training data in the group is represented as lt; s, g, r, a, s'gt; s represents the state, g represents the target, r represents the reward, a represents the adopted action, and s'represents the next state; constructing a plurality of negative samples for increasing the diversity of characterization disturbance; 2) maximizing the original input tuple lt in the collected training data; s, gt; obtaining a disturbed confrontation sample according to the characterization distance between the negative sample and the corresponding negative sample; 3) using the disturbed adversarial sample to enhance a value function and a strategy function of a target condition reinforcement learning agent, and optimizing an encoder network, a behavior network and an evaluator network; and 4) based on the optimized encoder network, the behavioral network and the evaluator network, constructing a target-oriented reinforcement learning agent with improved robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and relates to a method for improving the adversarial robustness of agent goal-oriented reinforcement learning. Background Art

[0002] In recent years, adversarial attacks have become an important research direction in the field of deep learning, especially in reinforcement learning algorithms. An attacker can interfere with the decision-making process of an agent by imposing small perturbations on the environmental state or the target. Such attacks not only affect the performance of the model but also may pose serious security risks.

[0003] Traditional methods mainly focus on the vulnerability of reinforcement learning, with relatively few studies on goal-oriented reinforcement learning, and many existing methods rely on pseudo-labels and critic networks. Although these methods have achieved certain results to some extent, they lack flexibility and face various limitations in practical applications. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a method for improving the adversarial robustness of agent goal-oriented reinforcement learning.

[0005] The present invention relates to a new semi-contrastive representation attack mechanism and a general adversarial representation strategy defense framework; the semi-contrastive representation attack effectively avoids the dependence on approximate labels and can operate independently of the evaluation function; the adversarial representation strategy aims to enhance the performance and security of the agent in complex environments and improve its robustness by more effectively dispersing the agent's attention.

[0006] The technical solution of the present invention is as follows:

[0007] A method for improving the adversarial robustness of agent goal-oriented reinforcement learning, the steps of which include:

[0008] 1) Collect a set of training data from the interaction between the goal-conditioned reinforcement learning agent and the environment; wherein, each piece of training data in the set is represented as <s, g, r, a, s ′ >, s represents the state, g represents the goal, r represents the reward, a represents the action taken, and s ′ represents the next state; construct multiple negative samples to increase the diversity of representation perturbations, and the negative samples corresponding to each set of training data are {<-s, g>, <s, -g>, <-s, -g>};

[0009] 2) Maximize the original input tuple <s, g> in the collected training data <s, g, r, a, s ′ > and the corresponding negative samples

[0010] The representation distance between {<-s, g>, <s, -g>, <-s, -g>} is obtained to get the perturbed adversarial sample;

[0011] 3) Use the perturbed adversarial sample to enhance the value function and policy function of the target conditional reinforcement learning agent, and optimize the encoder network ψ(·), the actor network and the critic network

[0012] 4) Based on the optimized encoder network ψ(·), the actor network and the critic network Construct a goal-oriented reinforcement learning agent with improved robustness.

[0013] Furthermore, maximize the representation distance between the original input tuple <s, g> in the collected training data <s, g, r, a, s ′ > and the corresponding negative samples {<-s, g>, <s, -g>, <-s, -g>} through the semi-contrastive representation attack method; the semi-contrastive representation attack method is: for the given feature extraction function f(·) and the input tuple <s, g>, by calculating Maximize the representation distance between the original tuple <s, g> and the corresponding negative samples {<-s, g>, <s, -g>, <-s, -g>}; where, <s, g> - represents the negative tuple, is a concave function, respectively represent the encodings of the state s and the goal g, is the set of goals within the norm ball centered at the goal g with radius ∈, is the set of states within the norm ball centered at the state s with radius ∈, is the expected value of the tuple <s, g>.

[0014] Furthermore, by calculating to minimize the similarity-based loss to achieve maximizing the representation distance between the original tuple <s, g> and the corresponding negative samples {<-s, g>, <s, -g>, <-s, -g>}.

[0015] Furthermore, use the projected gradient descent method to iteratively adjust the encoded states V(s) and the goal V(g) of the perturbed state s and the goal g to minimize the similarity-based loss function Finally, obtain the perturbed adversarial sample The method is:

[0016] 21) At the (i + 1)-th iteration, calculate the semi-contrastive representation attack of state s

[0017] The semi-contrastive representation attack of target g

[0018] where

[0019] α is the learning rate, is the gradient calculated by backpropagation;

[0020] 22) When the maximum number of iterations I is reached, obtain the perturbed adversarial example where proj(·) represents the projection operation.

[0021] Furthermore, for the input state or target the projection operation proj(·) confines the input information within a norm ball centered at the origin with a radius of ∈; where The norm is defined as the maximum value of the absolute values of the elements in a vector, and the absolute value of each element in the vector does not exceed ∈.

[0022] Furthermore, adopt the adversarial representation strategy and use the adversarial example obtained after perturbation to enhance the value function and policy function of the target-conditioned reinforcement learning agent. The method is as follows:

[0023] 31) At any time step t, retrieve a small batch of tuples from the replay buffer and generate semi-contrastive enhanced samples according to the negative tuples corresponding to the tuples and V and V scr (g m ), is the state corresponding to the m-th tuple in time step t, is the action taken in state g m is the target corresponding to the m-th tuple;

[0024] 32) Use the enhanced value function and the enhanced policy function to optimize the weights of the encoder network ψ(·), the actor network and the critic network :

[0025] where, update the enhanced policy function The method for the weights of the encoder network ψ(·) in is as follows: For the states and target samples <s m , g1> and <s j , g2> retrieved from the replay buffer and the corresponding state perturbations target perturbations and the trade-off factor β, define a sensitivity-aware regularizer for the encoder network ψ(·)

[0026]

[0027] to optimize the weights of the encoder network ψ(·).

[0028] Furthermore, by performing a negation operation on the original s-tuple <s, g>, that is, taking the negation of each element, a negative tuple (<s, g> - ) is obtained.

[0029] Furthermore, the target-conditioned reinforcement learning agent is an intelligent robot, an autonomous vehicle, or an intelligent system.

[0030] A server, comprising a memory and a processor, wherein the memory stores a computer program configured to be executed by the processor, and the computer program includes instructions for performing the above method.

[0031] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above method.

[0032] The present invention focuses on adversarial representation attacks and their defense mechanisms for agent goal-oriented reinforcement learning algorithms. With the rapid progress of artificial intelligence technology, especially its wide application in fields such as autonomous decision-making, robot control, and autonomous driving, the importance of reinforcement learning algorithms has become even more prominent. In diverse and complex application scenarios, improving the robustness of the model is crucial for maintaining the reliability and security of the system. The innovative method proposed by the present invention can effectively identify and cope with adversarial perturbations, helping to improve the adaptability of the goal-oriented reinforcement learning agent model in the face of potential extreme situations, thereby promoting the application performance of intelligent systems in unknown and dynamic environments. The application fields of this method cover, but are not limited to, goal-oriented tasks such as intelligent robots, autonomous vehicles, and intelligent systems (i.e., the target-conditioned reinforcement learning agent is an intelligent robot, an autonomous vehicle, or an intelligent system), and explore various edge cases through semi-contrastive representation attacks.

[0033] 1: Semi-contrastive Representation Attack (SCR)

[0034] The present invention proposes a new type of attack method, which is specifically customized for the agent target conditional reinforcement learning algorithm. Different from traditional methods, this attack method does not rely on specific pseudo-labels and does not require access to the critic network. This attack and defense mechanism not only has the unique characteristics of target conditional reinforcement learning, but also can be widely applied to a variety of reinforcement learning algorithms, showing good generality. This representation-based adversarial attack can be seamlessly integrated in the deployment stage. The present invention adopts the representation-based adversarial attack, which can be seamlessly integrated in the deployment stage and effectively overcomes the problems faced by the state-based adversarial attack when lacking gradient information of labeled examples. The specific definition and implementation method are as follows:

[0035] 1.1: Definition of semi-contrastive representation attack

[0036] The semi-contrastive representation attack aims to maximize the representation distance between the original state and its corresponding perturbed version. Given the feature extraction function f(·) and the input tuple <s, g>, the semi-contrastive representation attack can be defined as follows:

[0037]

[0038] where, Logical function applicable to any (v). <s, g> - Denotes a negative tuple. In the present invention, the negative sample tuple is generated by a negation operation, specifically including <-s, g>, <s, -g> and <-s, -g>. Here, s represents a state vector in the state set S, and g represents a target vector in the target set G. The negation operation is achieved by taking the inverse of each element of the state s or the target g. Since is a concave function, we can utilize subadditivity to optimize the calculation, which is expressed as follows:

[0039]

[0040] Specifically, the challenge of calculating the expected supremum can be reformulated as determining the expectation over the supremum. This means that our goal is to optimize the given negative tuple <s, g> - of i.e., the above convex function. In addition, by utilizing the non-increasing property of , the present invention can approximate the attack target by directly minimizing the similarity-based loss . In the present invention, the function f(·) is equivalent to the subsequent encoder ψ(·). Based on the above analysis, the present invention proposes an approximation method based on projected gradient descent, which is applicable to the state and the target In the training data, s represents the state and g represents the target. Here, the state and the target represent the encodings of the state s and the target g respectively.

[0041] 1.2: Approximate method based on projected gradient

[0042] The adversarial attack method for the original input tuple <s, g> uses the encoder ψ(·) and the predefined negative tuple <s, g> - and the step size α. The iterative update rule is expressed as follows:

[0043] At each iteration step i+1, first calculate the loss function Then calculate the gradient through backpropagation Finally, scale the gradient using the learning rate α and update the perturbation state V scr (s) and the target V scr (g). The semi-contrastive representation of the state attacks V scr (s) and the semi-half representation of the target attacks V scr (g). The initial perturbations are set to and The iteration is defined as follows:

[0044]

[0045] More specifically, it is defined as:

[0046]

[0047] In the case of the iteration number being (I, (≥1)), the finally generated adversarial tuple can be defined by the following projection function:

[0048]

[0049] where, (proj(·)) represents the projection operation. According to the definition of the present invention, (proj(·)) is expressed as a (∈-) bounded norm ball. Specifically, this means that for any input state or target the projection operation (proj(·)) will limit them within a norm ball centered at the origin with a radius of ∈ so as to ensure that the generated adversarial samples are within a certain range. Among them the norm (also known as the maximum norm) is defined as the maximum value of the absolute values of the elements in a vector, which means that the absolute value of each element of the vector does not exceed ∈.

[0050] In constructing the negative tuple (<s, g> -) When we do this, we adopt a simple and efficient method to generate the following combinations by negating the original tuple <s, g>: <-s, -g>, <-s, g>, and <s, -g>. This construction method can effectively introduce semi-contrastive representation attacks at each time step in T steps of target-conditioned reinforcement learning, aiming to distract the agent from the established goal. The advantage of this method is that through this adversarial attack, the reward sequence ((r0,…, r T-1})) can be maintained as sparse as possible, thereby achieving the efficiency and adversarial robustness of agent training.

[0051] 2. Adversarial Representation Strategy

[0052] The present invention also proposes an innovative hybrid defense strategy called the adversarial representation strategy, aiming to dynamically enhance the adversarial attack robustness of the reinforcement learning algorithm under specific target conditions of the agent. This strategy combines semi-contrastive adversarial enhancement with a sensitivity-aware regularizer to enhance the resistance of the underlying reinforcement learning agent to various perturbations. Through this combination, the adversarial representation strategy effectively improves the robustness of the target-conditioned reinforcement learning algorithm against adversarial attacks, enabling it to better adapt to and handle potential threats in complex environments. The algorithm is divided into two parts, as follows:

[0053] 2.1: Semi-contrastive Adversarial Enhancement

[0054] The semi-contrastive adversarial enhancement method proposed in the present invention generates adversarial samples through SCR attacks on input tuples during the training process, and uses these samples to enhance the value function and policy function, thereby optimizing the weights of the encoder, actor network, and critic network to study the impact of data augmentation.

[0055] Specifically, in each training cycle of the basic target-conditioned reinforcement learning agent, we retrieve a small batch of tuples from the replay buffer and use projected gradient descent to construct negative tuples to generate semi-contrastive enhanced samples and V scr (g n ). Subsequently, we use the enhanced value function represented as and the enhanced policy function represented as to optimize the weights of the encoder network θ ψ , actor network and critic network .

[0056]

[0057] g nThey are a set of states and target samples in a small batch of retrieved tuples, where the subscript t represents the time step and n represents the sample index. is an action sample in the mini-batch of tuples, which is the action taken under the state and V scr (g n ) is to generate semi-contrastive augmented samples by using the constructed negative tuples.

[0058] 2.2: Sensitivity-Aware Regularizer

[0059] To make up for the deficiencies of simple state representation in goal-conditioned reinforcement learning, the present invention proposes a new regularization method to measure its sensitivity by considering the Lipschitz constants of states and goals. By given two pairs of observed data and Due to the sparsity of the reward sequence, usually In this context, we can derive from the method of updating the state-goal tuple similarity in the policy function by using the above pair of transitions, and iteratively update the augmented policy function using the mean squared loss for the encoder ψ(·) in

[0060]

[0061] Here, largely depends on the next state s m+1 and s j+1 , and it is usually difficult to extract effective information from the absolute reward difference . This situation limits the evaluation of the difference between state-goal tuples in the current step and hinders the policy and critic functions from obtaining valuable insights in the interaction between the agent and the environment. To address this limitation of SimSR used in GCRL, we propose a sensitivity-aware regularizer to replace so as to enhance the effect of policy learning.

[0062] Given the states and target samples <s m , g1> and <s j , g2> in two sets of tuples retrieved from the replay buffer, and the corresponding state perturbations target perturbations and the trade-off factor β, the sensitivity-aware regularizer of the encoder ψ(·) can be defined as:

[0063]

[0064] In the present invention, to optimize the encoder parameter θ ψ。We constructed an experience-based loss function. This loss function introduces a robustness term M(·) / ||δ||2, which is used to characterize the implicit differences between tuples. By definition, physically close tuples should have similar Lipschitz constants in terms of state and target, while physically distant tuples should exhibit larger differences. This means that tuples that are physically far apart tend to have significantly different Lipschitz constants. We can further strengthen this property by introducing a simple state representation operator, where the absolute reward difference in the simple state representation operator can effectively quantify the differences between tuples. Therefore, by designing this loss function and the robustness term, we aim to improve the robustness and accuracy of the model, thereby achieving more robust performance in practical applications. The simple state representation operator is expressed as follows:

[0065]

[0066] where π represents the policy function, T π represents the simple state representation operator.

[0067] The advantages of the present invention are as follows:

[0068] The semi-contrastive representation attack proposed by the present invention can enhance the robustness of the target-conditioned reinforcement learning algorithm when experiencing adversarial attacks. This method abandons the dependence on pseudo-labels and the critic network, providing a more flexible and effective strategy for adversarial attacks. This method introduces adversarial perturbations through a feature extraction function and input tuples (such as states and targets), and its core goal is to maximize the representation distance between the original state and its perturbed version at each time step. This mechanism can effectively guide the attention of the agent for more precise goal guidance, thereby improving the adaptability and robustness of the algorithm and ensuring its excellent performance in various complex environments.

[0069] To enhance the performance of the target-conditioned reinforcement learning algorithm, the present invention further proposes an adversarial representation strategy. This strategy uses the adversarial samples generated by the semi-contrastive representation attack to optimize the weights of the agent's encoder, actor network, and critic network, thereby improving the learning ability and performance of the model in adversarial environments. In this way, the agent can not only effectively cope with adversarial interference but also improve the training efficiency.

[0070] The present invention also introduces a sensitivity-aware regularizer, which takes into account the Lipschitz constant between the state and the target to improve the differential representation between state-target tuples. This regularization method enables the agent to gain more valuable insights when interacting with the environment, thereby significantly enhancing the effect of policy learning.

[0071] In summary, the attack and defense mechanism of the present invention not only demonstrates good generality and applicability, but also helps to improve the ability of reinforcement learning algorithms to cope with various disturbances in complex dynamic environments. This method performs excellently in a variety of advanced goal-conditioned reinforcement learning algorithms, significantly enhancing the robustness against adversarial attacks compared to the current state-of-the-art algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 is a flowchart of a method for enhancing adversarial robustness in goal-directed reinforcement learning designed according to the present invention.

[0073] Figure 2 is a structural diagram of an adversarial representation strategy framework for robustness enhancement according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] The present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0075] The present invention introduces adversarial perturbations through a feature extraction function and a state-goal tuple, maximizing the representation distance between the original state and the perturbed version to enhance the attack effect. It avoids relying on pseudo-labels and critic networks, providing a more flexible and effective attack method. Subsequently, the proposed adversarial representation strategy uses the generated adversarial samples to optimize the weights of the agent's encoder, actor network, and critic network, enhancing the model's learning ability and performance in adversarial environments. The sensitivity-aware regularizer included therein considers the Lipschitz constant between the state and the goal, improving the differential representation between the state-goal tuples and significantly enhancing the policy learning effect. The attack and defense mechanism of the present invention not only has good generality and applicability, but also can improve the anti-perturbation ability of reinforcement learning algorithms in complex dynamic environments. The following is the detailed process of the method:

[0076] (I) Data collection and sample construction:

[0077] (1) Collect a set of training data from the interaction between a goal-conditioned reinforcement learning agent and the environment. The data tuple form is:

[0078] <s, g, r, a, s ′ >]

[0079] where s represents the state, g represents the goal, r represents the reward, a represents the action taken, and s ′ represents the next state.

[0080] (2) Use the above data tuples to construct negative sample tuples:

[0081] {<-s, g>, <s, -g>, <-s, -g>}

[0082] Increase the diversity of representation perturbations through these negative samples.

[0083] (2) Sample generation based on semi-contrastive representation attack:

[0084] (1) Through the semi-contrastive representation attack method, optimize the input state and the perturbed version of the target at each time step to maximize the representation distance between the original tuple <s, g> and the perturbed tuple, i.e., the negative sample tuple.

[0085] (2) Use the projected gradient descent method to iteratively adjust the perturbed state V(s) and the target V(g) to minimize the loss function of the representation distance. In this way, we gradually approach the optimal perturbed version and finally obtain the perturbed adversarial samples. These perturbations can maximize the representation distance between the original tuple and the negative sample tuple, thereby improving the robustness of the agent.

[0086] (3) Optimization of adversarial representation strategy:

[0087] (1) Semi-contrastive adversarial enhancement, using the adversarial samples generated in step (2). Enhance the value function and policy function of the target-conditioned reinforcement learning agent, and optimize the parameters of the following networks: the encoder network ψ(·), by minimizing the difference between the adversarial samples and the original samples in the encoding space, enhance the robustness of the encoder, loss function; the actor network Enhance the robustness of the actor network by minimizing the policy loss on the adversarial samples; the critic network Enhance the robustness of the critic network by minimizing the value function loss on the adversarial samples.

[0088] (2) Sensitivity-aware regularizer optimization. To consider the sensitivity relationship between the state and the target, introduce a sensitivity-aware regularizer. Specifically, by minimizing the rate of change of the state and the target on the adversarial samples, enhance the robustness of the network.

[0089] (4) Construction and verification of a robust reinforcement agent:

[0090] (1) Agent construction, based on the optimized network parameters ψ(·), and Construct a goal-oriented reinforcement learning agent with improved robustness.

[0091] (2) Performance verification, test the performance of the agent in a complex dynamic environment, and verify its adaptability and performance improvement in an adversarial environment.

[0092] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.

Claims

1. A method for improving adversarial robustness of intelligent agent goal-oriented reinforcement learning, the steps comprising: 1) Collect a set of training data from the interaction between the target-conditional reinforcement learning agent and the environment; each training data in the set is represented as <s,g,r,a,s ′ >, s represents the state, g represents the goal, r represents the reward, a represents the action taken, s ′ Indicates the next state; constructs multiple negative samples to increase the diversity of representation disturbances. The negative samples corresponding to each set of training data are {<-s,g>,<s,-g> ,<-s,-g>}; 2) Maximize the amount of training data collected <s,g,r,a,s ′ > The original input tuple<s,g> and the corresponding negative samples {<-s,g>,<s,-g> ,<-s,-g>}, and obtain the perturbed adversarial sample; 3) Use the perturbed adversarial samples to enhance the value function and strategy function of the target conditional reinforcement learning agent, optimize the encoder network ψ(·), the actor network and reviewer network 4) Based on the optimized encoder network ψ(·), actor network and reviewer network Building goal-oriented reinforcement learning agents with improved robustness.

2. The method according to claim 1, characterized in that Maximizing the collected training data through semi-contrastive representation attack method <s,g,r,a,s ′ > The original input tuple<s,g> and the corresponding negative samples {<-s,g>,<s,-g> ,<-s,-g>}; the semi-contrastive characterization attack method is: for a given feature extraction function f(·) and an input tuple<s,g> , by calculating Maximize the original tuple<s,g> and the corresponding negative samples {<-s,g>,<s,-g> ,<-s,-g>}; where, <s,g> - represents a negative tuple, is a concave function, Represent the encoding of state s and target g respectively, represents the target g as the center and the radius ∈ -The target set in the norm ball, represents the state s as the center and the radius ∈ -The set of states within the norm ball, To represent a tuple<s,g> expected value.

3. The method according to claim 2, characterized in that By calculation To minimize the loss based on similarity Maximizing the original tuple<s,g> and the corresponding negative samples {<-s,g>,<s,-g> ,<-s,-g>}.

4. The method according to claim 3, characterized in that The projected gradient descent method is used to iteratively adjust the encoding state V(s) and target V(g) of the perturbation state s and the target g to minimize the similarity-based loss function Finally, we get the perturbed adversarial sample The method is: 21) At the i+1th iteration, calculate the semi-contrast representation attack of state s The half-and-half of target g indicates attack in, α is the learning rate, Calculate gradients for backpropagation; 22) When the maximum number of iterations I is reached, the perturbed adversarial sample is obtained in, proj(·) represents a projection operation.

5. The method according to claim 4, characterized in that For the input status or target The projection operation proj(·) restricts the input information to a region with the origin as the center and a radius of ∈ -norm sphere; The -norm is defined as the maximum absolute value of each element in a vector, where the absolute value of each element in the vector does not exceed ∈.

6. The method according to claim 4, characterized in that Adopting the adversarial representation strategy and using the adversarial samples obtained after perturbation to enhance the value function and policy function of the target conditional reinforcement learning agent, the method is as follows: 31) At any time step t, retrieve a small batch of tuples from the replay buffer And according to the tuple The corresponding negative tuple generates a semi-contrast enhanced sample and V scr (g m ), is the corresponding state in the mth tuple in time step t, is in state The action taken next, g m is the corresponding target in the mth tuple; 32) Using the Enhanced Value Function And the enhanced strategy function Optimize the encoder network ψ(·) and the actor network and reviewer network Weight: Among them, the update enhancement strategy function The weights of the encoder network ψ(·) in ψ(·) are calculated as follows: given two tuples of state and target samples retrieved from the replay buffer, m ,g1> and j ,g2> and the corresponding state disturbance Target perturbation and a trade-off factor β, defining a sensitivity-aware regularizer for the encoder network ψ(·)​​ The weights used to optimize the encoder network ψ(·).

7. The method according to claim 2 or 3, characterized in that: By the original s tuple<s,g> Performing the negation operation means inverting each element to obtain a negative tuple (<s,g> - ).

8. The method according to claim 1, characterized in that The target condition reinforcement learning agent is an intelligent robot, an autonomous driving car, or an intelligent system.

9. A server, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.