Safety reinforcement learning system and method based on empirical modification

By introducing the security experience priority playback method SPER, the state is reshaped and the priority is adjusted, which solves the problem of sparse loss in the existing security reinforcement learning algorithm, and improves the performance and system stability of the agent in security tasks.

CN120470593APending Publication Date: 2025-08-12重庆中科汽车软件创新中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510613406.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing security reinforcement learning algorithms do not pay enough attention to security-related experience in security-related fields, which leads to sparse loss problems, leading to large approximation errors in state action loss networks, unstable policy network updates, and even generate unsafe policies.

Method used

The security experience priority playback method SPER is introduced, which reshapes the state, increases safety-related signals, adjusts priority, combines PER and DRB, updates network parameters, and optimizes the learning process.

Benefits of technology

Improve the performance of the agent in safety-critical tasks, reduce safety accidents, enhance the safety and stability of the system, optimize decision-making strategies, and reduce collision rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470593A_ABST
    Figure CN120470593A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent system safety, in particular to a safety reinforcement learning system and method based on empirical modification. The method comprises the steps that an intelligent agent interacts with the environment to generate experience, the state is reshaped, and the experience is enhanced to be (phi (s), a, phi (s'), r, c and d); a safety related signal f is increased through a safety experience priority playback method SPER, and experience (phi (s), a, phi (s'), r, c, d and f) is obtained; judging and processing the safety related signal f according to the loss signal; modifying the priority according to the safety related signal f; updating the priority of the batch samples; and updating the network according to a secure deep reinforcement learning method. According to the technical scheme, the approximation error of the state action loss network can be smaller, and the sparse loss problem can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent system security technology, and in particular to a security reinforcement learning system and method based on experience modification. Background Art

[0002] Currently, the state-action loss network is updated in the same way as the state-action reward network. Therefore, the method for handling sparse loss can refer to the method for handling sparse rewards. Methods for addressing sparse rewards include reward shaping, double repaly buffer (DRB), and prioritized experience replay (PER). The core ideas are as follows: Reward Shaping The core idea of profit reshaping is to redesign the profit function based on the potential energy function, so as to ensure that the optimal strategy of the new profit function is the same as the optimal strategy of the original profit function.

[0003] Double Replay Buffer (DRB) Double Replay Buffer (DRB) is a technique used in improved reinforcement learning algorithms, particularly in experience replay-based algorithms such as the Deep Q-Network (DQN). DRB uses two buffer pools, one storing important experiences and the other storing normal experiences. DRB updates the network by sampling from both buffers with specific probabilities, sampling important experiences with a greater probability and using them to update the state-action-reward network.

[0004] Prioritized Experience Replay (PER) Prioritized Experience Replay (PER) is an improved experience replay technique used to improve the performance and efficiency of reinforcement learning algorithms. Its core concept is to use TD-error (Temporal Difference Error) to define the priority of experiences in the experience pool. A higher priority means a higher probability of being sampled. TD-error represents the difference between the current estimated reward and the actual observed reward and is often used to measure the prediction error of the current policy.

[0005] However, in the complete collision avoidance task, the loss obtained from the environment must be zero. However, loss reshaping methods introduce non-zero loss, causing the predictions of the state-action loss network to be greater than 0. This affects the update of the policy network. PER prioritizes experience by calculating the TD-error of the state-action loss network. In other words, experience with a higher TD-error relative to the loss is more important and more likely to be sampled. However, in safety-critical fields (such as the Safety Gym simulation platform, a simulation platform for reinforcement learning research focused on safety-related tasks), safety is the primary concern, and safety-related experience should be sampled more frequently than experience with a higher TD-error.

[0006] Existing safety reinforcement learning algorithms lack sufficient attention to safety-critical experience, resulting in a low probability of collisions between agents and obstacles in certain tasks. Furthermore, the better the collision avoidance policy learned by the agent, the lower the probability of a non-zero loss, a phenomenon known as the sparsity loss problem. This results in large errors in the state-action loss network's approximation of the state-action loss function, resulting in poor approximation. This can lead to "incorrect" policy network updates, causing policy divergence and even unsafe policy networks. Summary of the Invention

[0007] The purpose of the present invention is to propose a secure reinforcement learning system and method based on experience modification, which can reduce the approximation error of the state-action loss network and handle the sparse loss problem.

[0008] To achieve the above-mentioned objectives, in a first aspect, the present invention discloses a safety reinforcement learning method based on experience modification, including: an intelligent agent generates experience through interaction with the environment, reshapes the state, and enhances the experience to (Φ(s), a, Φ(s'), r, c, d); through the safety experience priority replay method SPER, a safety-related signal f is added to obtain experience (Φ(s), a, Φ(s'), r, c, d, f); the safety-related signal f is judged and processed according to the loss signal; the priority is modified according to the safety-related signal f; the priority of the batch samples is updated; and the network is updated according to the safety deep reinforcement learning method.

[0009] Beneficial Effects of the Basic Solution: This technical solution provides a more comprehensive description of the interaction between the agent and the environment. By reshaping the state, more valuable features can be extracted, helping the agent better understand the environmental state and make more accurate decisions. Furthermore, the increased information dimension provides more basis for the learning process, improving both accuracy and efficiency.

[0010] By introducing the safety-focused experience replay method (SPER) and adding a safety-related signal f to obtain (Φ(s), a, Φ(s'), r, c, d, f), the agent can focus more on safety-related experience. Prioritizing the replay of safety-related experience allows the agent to more quickly learn how to make decisions while ensuring safety, improving its performance in safety-critical tasks and reducing the occurrence of safety incidents.

[0011] By evaluating and processing the safety-related signal f based on the loss signal, the agent can dynamically adjust its focus on safety based on the actual situation. When the loss signal indicates a safety issue, the safety-related signal can be promptly processed, guiding the agent to take measures to avoid or mitigate the risk, thereby enhancing the security and stability of the system.

[0012] By modifying priorities based on the security-related signal f and updating the priority of batches of samples, the agent can more flexibly adjust its learning focus. Highly security-relevant experiences are given higher priority, allowing them to be learned and utilized more frequently. This accelerates the agent's learning of security policies and helps it better adapt to complex and changing security environments.

[0013] Updating the network using secure deep reinforcement learning methods combines the previous steps to effectively incorporate safety-related information into the network's learning process. By continuously updating network parameters, the agent can gradually optimize its decision-making strategy, not only improving task completion efficiency but also achieving better performance while ensuring safety, enabling the agent to achieve optimal or near-optimal behavior within safety constraints.

[0014] This application improves the problem of sparse loss by modifying experience in spatial states and probabilistically. Specifically, through a state reshaping method, the state is reshaped so that the state-action loss network approximates the state-action loss function. Then, PER and DRB are combined to unify them and a secure experience priority replay method SPER is proposed to achieve the purpose of addressing the sparsity loss problem.

[0015] As an implementable preferred solution, the agent interacts with the environment to generate experience and reshape the state. The experience enhancement is (Φ(s), a, Φ(s'), r, c, d), which includes the following: The agent interacts with the environment. The agent reaches state s' by taking action a in state s, obtains reward signal r, loss c and termination signal d, and generates an experience (s, a, s', r, c, d); The state is reshaped by the method of income reshaping and state aggregation to obtain the reshaped state Φ(s); Select actions based on the reshaped state Φ(s) Reach the next state , and obtain experience (Φ(s),a, Φ(s'),r,c,d).

[0016] As an implementable preferred solution, the state is reshaped. The state reshaping formula is as follows:

[0017] Among them, s represents the state before reshaping, Indicates status The state of remodeling, is the enhancement function of state s, is a mapping of the state s; the 1-norm is used to measure whether the agent is near the obstacle / goal.

[0018] As an implementable preferred solution, the agent Safety-related signals for step interaction:

[0019] in, Indicates that the agent is Safety-related signals for step-by-step interaction, Indicates that the agent is The loss of step interaction.

[0020] As an implementable preferred solution, the safety experience priority replay method SPER is used to add a safety-related signal f, and the experience (Φ(s), a, Φ(s'), r, c, d, f) is obtained, which also includes the following: The current safety-related signal Set to 0, get the experience (Φ(s), a, Φ(s'), r, c, d, f=0), and use the replay buffer of length k as the secondary buffer to temporarily store the current k-step interaction experience .

[0021] As an implementable preferred solution, judging and processing the safety-related signal f according to the loss signal includes the following: If the current loss signal , Set the experienced safety signals in all auxiliary buffers Secondary Buffer to 1, otherwise do not make any changes.

[0022] As an implementable preferred solution, the priority is modified according to the safety-related signal f, including the following: Calculate its priority based on the state-action loss network: if Step interaction , all experienced safety-related signals in the Secondary Buffer are set to Otherwise, no change is made; or the priority is modified according to the safety-related signal f by setting the experience priority to a constant; and placing the experience into the primary buffer.

[0023] As a feasible and preferred solution, the priority of updating batch samples includes the following: According to the state-action loss network, the TD-error and priority calculated according to the PER experience are calculated. If yes, the priority is increased; otherwise, the priority remains unchanged. The priority of the batch samples is updated and stored in the primary buffer.

[0024] As an implementable preferred solution, the network is updated according to the secure deep reinforcement learning method, including the following: 8. According to the secure deep reinforcement learning method, the method of updating the network is to batch sample from the primary buffer according to priority, and update the state-action value network, state-action loss network and policy network according to the secure deep reinforcement learning method.

[0025] In the second aspect, the present invention discloses a security reinforcement learning system based on experience modification, which is installed on an intelligent body and uses the above-mentioned security reinforcement learning method based on experience modification, including an interactive experience generation unit, a state reshaping unit, an experience enhancement unit, a SPER increase in security-related signals to obtain experience unit, a security-related signal judgment processing unit, a priority modification unit, a batch sample priority update unit, and an update network unit. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A logical diagram of a secure reinforcement learning method based on experience modification.

[0027] Figure 2 Reshape the state for loss-related states.

[0028] Figure 3 Schematic diagram of the state-action loss function under sparse loss.

[0029] Figure 4 Schematic diagram of the state-action loss function after state reshaping.

[0030] Figure 5 Schematic diagram of the combination of state reshaping and SPER with algorithms such as DDPG-ALM.

[0031] Figure 6 Schematic diagram comparing the convergence of the three algorithms under state reshaping.

[0032] Figure 7 Schematic diagram of DDPG-ALM's safety-first experience replay convergence comparison.

[0033] Figure 8 Schematic diagram for DDPG-ALM state reshaping comparison.

[0034] Figure 9 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to make the technical solution and advantages of the present application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It will be understood that the specific embodiments described herein are only partial embodiments of the present invention, which are only used to explain the present application, rather than to limit the present application. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered to be isolated, and they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the drawings of the following embodiments represent the same features or components, which can be applied to different embodiments.

[0036] In addition, unless otherwise defined, technical or scientific terms used in the description of the present invention should have the common meanings understood by those skilled in the art in the art to which the present invention belongs.

[0037] The present invention will be further described in detail below with reference to the accompanying drawings: Reference numerals: electronic device 500 , processor 501 , communication interface 502 , memory 503 , bus 504 .

[0038] Reference Figure 1 and Figure 5 The present disclosure provides a secure reinforcement learning method based on experience modification, including: In step S1, the agent is placed in a simulated indoor navigation environment consisting of multiple rooms, corridors, and obstacles (such as furniture). The environment can be represented by a three-dimensional grid, with each grid cell containing attributes such as location and the presence of obstacles. The agent has state information such as position, speed, and direction, and can perform actions such as forward, backward, left, and right turns.

[0039] The agent interacts with the environment. From the initial state s, the agent interacts with the environment through action a and reaches the state s', obtains the reward signal r, loss c and termination signal d, and generates an experience (s, a, s', r, c, d).

[0040] Reward signal r, if the agent successfully reaches the target position, the reward r is set to a large positive value; if the agent approaches the target position during the movement, a certain positive reward is given according to the degree of proximity; if the agent is far away from the target position, the reward r is set to a small negative value.

[0041] Loss c: When the agent collides with an obstacle, loss c is set to a large positive value. During normal movement, loss c is set to a smaller positive value for each grid unit moved. In the complete collision avoidance task, the loss obtained from the environment is required to be zero. However, this invention takes into account practical circumstances and incurs losses when the agent interacts with obstacles.

[0042] Termination signal d: When the agent reaches the target position or collides seriously with an obstacle (for example, the number of collisions exceeds a certain threshold), the termination signal d is set to 1, indicating that the current interaction process ends; otherwise, d is set to 0.

[0043] Step S2, reshape the state through the method of income reshaping and state aggregation to obtain the reshaped state Φ(s).

[0044] In order to make the approximation error of the state-action loss network smaller, the idea of benefit reshaping and state aggregation is used to reshape the state. The state reshaping formula is as follows:

[0045] Among them, s represents the state before reshaping, Indicates status The state of remodeling, is the enhancement function of state s, is a mapping of state s.

[0046] If the agent is near an obstacle, it should focus on loss-related features, that is, loss-related features should be enhanced. is the enhancement function, which can be a linear or nonlinear function, such as ,in .

[0047] In addition, the target-related features are reshaped. The 1-norm can be used to measure whether the agent is near an obstacle. Figure 2 , where the blue dotted circle indicates whether the agent is near the obstacle; the green rectangle indicates the obstacle area; the red circle indicates the agent near the obstacle, and its detection of the loss-related state of the obstacle Enhanced to ; The blue cross represents the loss-related state of the agent outside the obstacle range Mapped to .

[0048] according to Figure 2 It can be seen that if there are no obstacles near the agent, it only needs to focus on the benefits; if the agent is near an obstacle, it needs to be more careful. State reshaping is to map the state corresponding to being away from the obstacle to the state After state reshaping, the state-action loss function The state space becomes smaller. Indicates that the status Take the following State-action loss function for an action.

[0049] Reference Figure 3 and Figure 4 , becomes , the state when away from obstacles also changes to ,Right now The range is the blue area in the figure, and the yellow area in the figure is mapped to .so, The solution space is smaller, the proportion of non-zero loss is higher, and it is smoother rather than a "spiky" function. Therefore, the approximation error of the state-action loss network is smaller.

[0050] Step S3, select action according to the reshaped state Φ(s) Reach the next state , and obtain experience (Φ(s),a, Φ(s'),r,c,d).

[0051] Step S4, add the safety-related signal f to obtain experience (Φ(s), a, Φ(s'), r, c, d, f) through the safety experience priority playback method SPER. Set to 0, get the experience (Φ(s), a, Φ(s'), r, c, d, f=0), and use the replay buffer of length k as the secondary buffer to temporarily store the current k-step interaction experience .

[0052] By combining PER and DRB into a unified framework, we propose the Safe Prioritized Experience Replay (SPER) method, which increases the probability of obtaining loss experience.

[0053] SPER inherits PER's experience prioritization mechanism, using the TD-error of an experience to indicate its priority. However, in the CMDP model (Continuous Markov Decision Process, a variant of the Markov decision process), safety-related experience should be given higher priority. Therefore, if the experience is safety-related, SPER will increase its priority, increasing the probability of its sampling, thereby making the state-action loss network corresponding to this experience more closely approximate the state-action loss function. Experience stored in the Replay Buffer in SPER is defined as: ; This formula indicates that Step interaction, the agent is in state Take action , reach the next state , and gain benefits ,loss , termination signal and safety-related signals ,in, Indicates that the agent is step-by-step interactive experience, Indicates that the agent is Step interaction status, Indicates that the agent is Take action when the step interaction state is reached, Indicates that the agent is The benefits gained from step interaction, Indicates that the agent is The loss of step interaction, Indicates that the agent is The termination signal of the step interaction, Indicates that the agent is Safety-related signals for step-by-step interaction.

[0054] Replay Buffer is a technique used in reinforcement learning to store data observed by the agent in the environment for later learning.

[0055] The agent in the Safety-related signals for step interaction:

[0056] in, Indicates that the agent is Safety-related signals for step-by-step interaction, Indicates that the agent is The loss of step interaction.

[0057] In addition, based on the DRB method, the safety-related signals are set retroactively. , backtracking along the trajectory Step forward The safety-related signals obtained from the experience of the first step interaction are all set to 1. That is: ,in, . when ,experience It is called safety-related experience.

[0058] SPER is a static security assessment technique that cannot update the security signal of the experience pool as training progresses. Specifically, based on state reshaping, SPER is combined with a secure deep reinforcement learning algorithm similar to DDPG (Deep Deterministic Policy Gradient, a deep reinforcement learning algorithm that combines the policy gradient method with the concept of deep Q-network). Two replay buffers are used to store experience e: the primary buffer and the secondary buffer. The basic process is as follows: Step S41: The agent interacts with the environment and gains experience through state reshaping. .

[0059] Step S42: Use the Replay Buffer with a length of k as the Secondary Buffer to temporarily store the current k-step interaction experience. .

[0060] Step S43, if Step interaction , all experienced safety-related signals in the Secondary Buffer are set to ; otherwise, no changes are made.

[0061] In step S44, according to the FIFO principle (First In First Out (FIFO) principle), the experience stored in the secondary buffer earliest is popped into the primary buffer, and the currently explored experience is placed into the secondary buffer.

[0062] Step S45: Calculate experience based on PER TD-error and priority, if , increase the priority; otherwise, the priority remains unchanged.

[0063] Step S46: Store the experience into the Replay Buffer (Primary Buffer) with the same structure as PER. In step S47, samples are taken from the Replay Buffer, and the state-action value network, state-action loss network, and policy network are updated. Then, the process goes to step S45.

[0064] Step S5: judging and processing the safety-related signal f according to the loss signal.

[0065] In step S5, the method for judging and processing the safety-related signal f according to the loss signal is as follows: if the current loss signal , Set the experienced safety signals in all auxiliary buffers Secondary Buffer to 1, otherwise do not make any changes.

[0066] In step S6, the priority is modified based on the safety-related signal f. The experience stored earliest in the secondary buffer is retrieved and its priority is set. Based on the first-in-first-out (FIFO) principle, the experience stored earliest in the secondary buffer is popped into the primary buffer, and the currently explored experience is placed in the secondary buffer.

[0067] In step S6, the priority is modified according to the safety-related signal f, and the specific method is as follows: Calculate its priority based on the state-action loss network: if Step interaction , all experienced safety-related signals in the Secondary Buffer are set to Otherwise, no changes are made.

[0068] Alternatively, set the experience priority to a constant; place the experience in the primary buffer.

[0069] Step S7: Update the priority of the batch samples. The method for updating the priority of the batch samples is as follows: According to the state action loss network, the PER calculation experience TD-error and priority, if , increase the priority. Otherwise, the priority remains unchanged, update the priority of the batch samples, and store them in the primary buffer.

[0070] S8: Update the network according to the secure deep reinforcement learning method.

[0071] According to the secure deep reinforcement learning method, the network is updated by sampling in batches according to priority from the primary buffer, and the state-action value network, state-action loss network, and policy network are updated according to the secure deep reinforcement learning method.

[0072] A safety reinforcement learning system based on experience modification is installed on an intelligent agent and includes an interactive experience generation unit, a state reshaping unit, an experience enhancement unit, a SPER unit for gaining experience by adding safety-related signals, a safety-related signal judgment and processing unit, a priority modification unit, a batch sample priority update unit, and an update network unit, wherein: The interaction experience generation unit is used for the agent to interact with the environment. The agent reaches state s' by interacting with the environment through action a in state s, obtains a reward signal r, loss c and termination signal d, and generates an experience (s, a, s', r, c, d).

[0073] The state reshaping unit is used to reshape the state through the method of income reshaping and state aggregation to obtain the reshaped state Φ(s).

[0074] The experience augmentation unit is used to select actions based on the reshaped state Φ(s) Reach the next state , and obtain experience (Φ(s),a, Φ(s'),r,c,d).

[0075] The SPER adds safety-related signals to obtain experience. The unit is used to add safety-related signals f to obtain experience (Φ(s), a, Φ(s'), r, c, d, f) through the safety experience priority playback method SPER. Set to 0, get the experience (Φ(s), a, Φ(s'), r, c, d, f=0), and use the replay buffer of length k as the secondary buffer to temporarily store the current k-step interaction experience .

[0076] The safety-related signal judgment and processing unit is used to judge and process the safety-related signal f according to the loss signal.

[0077] The priority modification unit is used to modify the priority based on the safety-related signal f. It retrieves the most recent experience stored in the secondary buffer and sets its priority. Based on the first-in-first-out (FIFO) principle, the oldest experience stored in the secondary buffer is popped into the primary buffer, and the currently explored experience is placed in the secondary buffer.

[0078] The batch sample priority updating unit is used to update the priority of the batch samples.

[0079] The update network unit updates the network according to the secure deep reinforcement learning method.

[0080] Finally, this example uses SPER and DDPG-like secure deep reinforcement learning algorithms based on state reshaping to better handle the problem of sparse loss. The demonstration process is as follows Figure 5 .

[0081] The agent will reach state s' through action a and interact with the environment, obtain reward signal r, loss c and termination signal d, and generate an experience of (s, a, s', r, c, d). The policy network is based on the state of the environment feedback. , reshape the state to , select actions based on the reshaped state Reach the next state , and gain experience (Φ(s), a, Φ(s'), r, c, d); the current safety-related signal Set to 0, get experience (Φ(s), a, Φ(s'), r, c, d, f=0), and store the experience in the Secondary Buffer; if the current loss signal , set the safety signals of all experiences in the secondary buffer to 1, otherwise make no changes; remove the experience first stored in the secondary buffer and set its priority. There are two methods: first, calculate its priority according to step 3 of SPER based on the state-action loss network, and second, set its priority to a large constant; put this experience into the primary buffer; batch sample from the primary buffer according to priority, and update the state-action value network, state-action loss network, and policy network according to a safe deep reinforcement learning algorithm such as DDPG-ALM (DDPG-ALM (Augmented Lagrangian Method) is an improved DDPG algorithm based on the augmented Lagrangian method), DDPG-Lagrange (DDPG-Lagrange is another improved DDPG algorithm based on Lagrangian multipliers), and DDPG-Lambda; update the priority of the batch samples according to step 5 of SPER based on the state-action loss network and store them in the primary buffer; The agent re-interacts with the environment and explores it, and so on.

[0082] Reference Figures 6 to 8 Compared to the original DDPG algorithm, the DDPG-ALM algorithm reduces the average loss rate from 11.280 to 2.513. With the support of state reshaping and safety-first experience replay technology, DDPG-ALM further generates a safer strategy, reducing the average loss rate from 2.513 to 0.178. While ensuring a safer strategy, DDPG-ALM better balances losses and benefits, demonstrating the superior performance of this invention.

[0083] This method can be combined with existing safety reinforcement learning algorithms, achieving good integration. It also improves the security of existing algorithms by, for example, reducing the collision rate through extensive training experience with collision losses. Furthermore, the proposed method is easy to implement and understand.

[0084] Those skilled in the art will understand that all or part of the processes in a secure reinforcement learning method based on experience modification can be implemented by instructing related hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of various embodiments of a secure reinforcement learning method based on experience modification. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0085] The present application also provides an electronic device 500 that utilizes the aforementioned experience-based security reinforcement learning system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the aforementioned experience-based security reinforcement learning method are implemented. In the present application, the processor serves as the control center of the computer method and can be a processor of a physical machine or a processor of a virtual machine.

[0086] Reference Figure 9 The electronic device 500 includes: at least one processor 501, at least one communication interface 502, at least one memory 503, and at least one bus 504. Bus 504 is used to implement communication between these components, communication interface 502 is used to communicate signaling or data with other node devices, and memory 503 stores machine-readable instructions executable by processor 501. When the electronic device 500 is running, processor 501 communicates with memory 503 via bus 504. When the machine-readable instructions are called by processor 501, the steps of the above-mentioned experience-based secure reinforcement learning method are executed.

[0087] The above contents are merely embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. A person of ordinary skill in the art is aware of all common technical knowledge in the technical field to which the invention belongs before the filing date or priority date, is able to obtain all existing technologies in the field, and has the ability to apply conventional experimental means before that date. A person of ordinary skill in the art can, under the guidance of this application, improve and implement this scheme in combination with his or her own abilities. Some typical known structures or known methods should not become an obstacle for a person of ordinary skill in the art to implement this application. It should be pointed out that for a person of ordinary skill in the art, several variations and improvements can be made without departing from the structure of the present invention, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection claimed in this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A secure reinforcement learning method based on experience modification, characterized in that: include: The intelligent agent generates experience through interaction with the environment, reshapes its state, and enhances the experience to (Φ(s), a, Φ(s'), r, c, d); through the safe experience priority replay method SPER, the safety-related signal f is added to obtain experience (Φ(s), a, Φ(s'), r, c, d, f); the safety-related signal f is judged and processed based on the loss signal; the priority is modified based on the safety-related signal f; the priority of the batch samples is updated; and the network is updated according to the safe deep reinforcement learning method.

2. A secure reinforcement learning method based on experience modification according to claim 1, characterized in that: The agent interacts with the environment to generate experience, reshapes the state, and the experience enhancement is (Φ(s), a, Φ(s'), r, c, d), which includes the following: The agent interacts with the environment. The agent reaches state s' by interacting with the environment through action a in state s, obtains reward signal r, loss c and termination signal d, and generates an experience (s, a, s', r, c, d); The state is reshaped by the method of income reshaping and state aggregation to obtain the reshaped state Φ(s); Select actions based on the reshaped state Φ(s) Reach the next state , and obtain experience (Φ(s),a, Φ(s'),r,c,d).

3. The method for secure reinforcement learning based on experience modification according to claim 2, characterized in that: Reshape the state. The state reshaping formula is as follows: Among them, s represents the state before reshaping, Indicates status The state of remodeling, is the enhancement function of state s, is a mapping of the state s; the 1-norm is used to measure whether the agent is near the obstacle / goal.

4. The method for secure reinforcement learning based on experience modification according to claim 1, characterized in that: The agent in the Safety-related signals for step interaction: in, Indicates that the agent is Safety-related signals for step interaction, Indicates that the agent is The loss of step interaction.

5. The secure reinforcement learning method based on experience modification according to claim 1, characterized in that: Through the safety experience priority replay method SPER, the safety-related signal f is added, and the experience (Φ(s), a, Φ(s'), r, c, d, f) is obtained, which also includes the following: The current safety-related signal Set to 0, get the experience (Φ(s), a, Φ(s'), r, c, d, f=0), and use the replay buffer of length k as the secondary buffer to temporarily store the current k-step interaction experience .

6. The method of secure reinforcement learning based on experience modification according to claim 1, characterized in that: The safety-related signal f is judged and processed according to the loss signal, including the following: If the current loss signal , Set the experienced safety signals in all auxiliary buffers to 1, otherwise do not make any changes.

7. The method of secure reinforcement learning based on experience modification according to claim 1, characterized in that: The priority level is modified according to the safety-related signal f, including the following: Calculate its priority based on the state-action loss network: if Step interaction , all experienced safety-related signals in the Secondary Buffer are set to Otherwise, no change is made; or the priority is modified according to the safety-related signal f by setting the experience priority to a constant; and placing the experience into the primary buffer.

8. The secure reinforcement learning method based on experience modification according to claim 1, characterized in that: Update the priority of batch samples, including the following: According to the state-action loss network, the empirical TD-error and priority are calculated according to PER. If yes, the priority is increased; otherwise, the priority remains unchanged. The priority of the batch samples is updated and stored in the primary buffer.

9. The secure reinforcement learning method based on experience modification according to claim 1, characterized in that: Update the network according to the secure deep reinforcement learning method, including the following:

8. Update the network by sampling priority batches from the primary buffer using secure deep reinforcement learning methods. Update the state-action value network, state-action loss network, and policy network using secure deep reinforcement learning methods.

10. A secure reinforcement learning system based on experience modification, installed on an intelligent agent, characterized in that A safety reinforcement learning method based on experience modification as described in claims 1-9 is used, including an interactive experience generation unit, a state reshaping unit, an experience enhancement unit, an SPER safety-related signal increase experience acquisition unit, a safety-related signal judgment processing unit, a priority modification unit, a batch sample priority update unit, and an update network unit.