Robot behavior decision-making method and device for simulating memory priority playback function of hippocampus

By simulating the robotic human brain behavior decision model that prioritizes playback mechanisms in the hippocampus, the existing robotic decision-making efficiency and accuracy are solved, and a faster and more efficient learning and decision-making process is achieved.

CN119940397APending Publication Date: 2025-05-06ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112617.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing robot behavior decision-making model has not yet fully learned from biological memory playback and priority selection mechanisms, resulting in the need to improve the efficiency and accuracy of behavior decision-making in complex environments.

Method used

A robotic human brain behavior decision model that simulates the working mechanism of time cells and memory-first playback in the hippocampus is designed, and a DQN-first memory playback algorithm is used to simulate the selective playback mechanism of the human brain, and combine the time cells to encode time series information.

Benefits of technology

By simulating the hippocampal-enthaler cortex memory mechanism, the robotic brain memory priority replay model can accelerate the learning process of the agent, improve the accuracy and efficiency of decision-making, and enable the robot to quickly adapt to unfamiliar environments and complete tasks efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940397A_ABST
    Figure CN119940397A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of robot intelligent control, and relates to a robot behavior decision method and device for simulating a hippocampus memory priority playback function. The method comprises the following steps: simulating behaviors of time cells in a hippocampus on continuous time and unique time coding characteristics in a memory process, further coding information completed by coding position cells and grid cells in a hippocampus-olfactory cortex structure, and carving time information on the information; then, a memory priority playback mechanism is designed, a working mechanism that a human brain performs selective playback according to the importance of memory is simulated by using a DQN memory priority playback algorithm, a robot brain-like memory priority playback model for simulating a hippocampus-endophytic cortex memory mechanism is designed, and equipment is used for realizing the robot behavior decision-making method. The algorithm provided by the invention has good convergence and generalization, so that the robot can quickly adapt to an unfamiliar environment and complete corresponding tasks more efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computers, and in particular to the field of robot intelligent control technology, and more specifically to a robot behavior decision method and device for simulating a hippocampal memory priority playback function. Background Art

[0002] In daily human activities, when entering an unfamiliar environment, even if we have never been there before, we can determine our own orientation and location through information in the surrounding environment, such as landmarks or objects, combined with past memories. This process relies on the unique ability of the brain to abstract common structures from the sensorimotor details of the experience and flexibly apply them to new and relevant situations. This abstract ability is described differently in different research fields. For example, it is called schema in human behavior and memory research, learning set in the context of animal reward-guided behavior, transfer learning and meta-learning in the field of machine learning, etc. This process of summarizing past memories or experiences and abstracting corresponding skills or knowledge is defined as memory playback and reorganization.

[0003] In the field of biological research, domestic and foreign researchers have found through a large number of experimental observations that higher organisms, including humans, will activate specific brain areas to simulate (replay) past experiences when they are resting or working. For example, Wilson first observed the phenomenon of fully mature memory replay during sleep in 2002. Many researchers interpret memory replay in the hippocampus as a memory consolidation process, that is, transferring short-term memories stored in the hippocampus to the prefrontal cortex to form long-term memories. Maingret's research further confirmed the causal role of the interaction between the hippocampus and the prefrontal cortex during sleep in memory consolidation. Michon's research pointed out that hippocampal replay has the function of selectively enhancing memory, especially for memories of high-reward locations in familiar scenes. In addition, Foster et al. found that place cells in the hippocampus of rats can reversely replay previously experienced behavioral sequences, which provides direct evidence for the occurrence of memory replay in the wakeful state, and further reflects its important role in decision-making and memory consolidation. Eichenbaum also discussed in depth the discovery and role of time cells in the hippocampus, including how they work with place cells to encode complex event sequences and temporal information, and their potential role in memory construction and playback.

[0004] Existing studies have shown that time cells in the hippocampus can encode organized temporal information of memories at multiple time scales, and can work in conjunction with place cells and grid cells in the hippocampus-entorhinal cortex to encode extremely complex sequence information. At the same time, during the memory playback process of the hippocampus-entorhinal cortex, the brain has the characteristic of preferentially replaying memories with higher importance. However, in the field of robotics, existing behavioral decision-making models have not yet fully borrowed from these biological memory playback and priority selection mechanisms, resulting in the need to improve the efficiency and accuracy of the robot's behavioral decision-making when facing complex environments and tasks. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of existing robot algorithms and solve the shortcomings of the robot's autonomous learning and decision-making ability when facing complex and changing environments. Inspired by the brain's memory replay mechanism, a robot brain-like behavior decision-making model and method that simulates the working mechanism of time cells in the hippocampus and memory priority replay are designed. The robot brain-like behavior decision-making model is a DQN memory priority replay (DQN Priority Experiential Replay based on Time Cell, DQN-PERTC) model based on time series.

[0006] Based on one of the purposes of the present invention, the robot behavior decision method proposed in the present application for simulating the hippocampal memory priority playback function comprises the following steps:

[0007] S1. Simulate the behavior of time cells in the hippocampus in continuous time and the unique time coding characteristics in the memory process, further encode the information encoded by place cells and grid cells in the hippocampus-entorhinal cortex structure, and engrave time information on it;

[0008] S2. Introduce the memory priority replay mechanism, use the DQN priority memory replay algorithm to simulate the working mechanism of the human brain for selective replay based on the importance of memory, and establish a robot brain-like memory priority replay model that simulates the memory mechanism of the hippocampus-entorhinal cortex. This model can both accelerate the learning process of the intelligent agent and improve the accuracy of its decision-making.

[0009] The working process of the robot brain-like memory priority playback model includes a learning and training phase and a task execution phase;

[0010] Among them, in the learning and training stage, the robot's work is divided into two steps: one is to collect and store information from the outside world, and the other is to organize the collected information into memory and use the memory to complete training;

[0011] After the learning and training phase, the robot enters the task execution phase, where it completes exploration tasks in a new environment based on the trained network.

[0012] The learning and training phase includes the following steps:

[0013] S31. The robot explores the environment in a reinforcement learning manner and continuously stores the explored information into the memory pool. In this process, the place cells and grid cells in the hippocampus-entorhinal cortex structure encode environmental information and self-motion information, and the time cells encode time series information, which are gradually converted into memory and stored in the memory pool;

[0014] S32. When the memory pool is full or the robot completes the exploration task, the robot enters the learning and training mode with memory priority playback.

[0015] When collecting and storing external information, the robot first needs to initialize the basic input parameters, including the maximum number of iterations, the current number of rounds, the number of learning times, the learning rate, the reward discount factor, the priority index and the deviation correction coefficient, the network update interval, the memory pool capacity, and the number of samples drawn each time. Based on the above parameters, the value network Qvalue and the target network Qtarget are created, and the memory pool is initialized, with the initial cumulative weight set to 0 and the initial priority set to 1.

[0016] In step S31, the robot's intelligent agent obtains the current state in the environment, and determines the action to be taken based on the existing memory and action strategy, enters the next state in the environment based on this action, and obtains the reward of environmental feedback, thereby obtaining a state transition sequence. The state transition sequence is encoded and stored in the memory pool as the history of the intelligent agent in the process of environmental interaction. After executing a state transfer, the Q function of the robot's brain-like memory priority playback model adopts a single-step update strategy to update the Q value of the current state-action pair, and updates the value network Qvalue and the target network Qtarget according to the generated Q value.

[0017] In the task execution phase, when the robot explores the new environment, the agent first obtains the current state S current , then calculate the current state S current With the stored history state S memory Similarity, set the similarity index C slt The calculation formula is as follows:

[0018]

[0019] If C slt≥ρ, ρ is the set similarity threshold, indicating that the current state has never appeared in the previous task process. The network needs to create a new cognitive node and store the group of state information into the memory pool. Otherwise, it means that the current state is experienced by the agent. The system will extract the decision plan corresponding to the current state to guide the agent to take the next action.

[0020] The system framework of the robot brain-like memory priority playback model includes:

[0021] Mobile robots can obtain environmental and motion information of their environment;

[0022] The coding region includes grid cells, time cells and place cells, and the grid cells, time cells and place cells work together to encode the environmental information and motion information;

[0023] Memory storage area, used to store encoded information;

[0024] The memory playback area is used to retrieve information in the memory storage area. When relevant memories are retrieved, a designated decision behavior is performed according to the relevant memories. When relevant memories are not retrieved, a random decision behavior is performed.

[0025] Among them, the grid cells, time cells and place cells in the coding area can simulate the working mechanism of the corresponding cells in the biological hippocampus-entorhinal cortex to realize the encoding of complex sequence information, and the memory playback area can simulate the priority selection mechanism of important memories during the biological memory playback process to improve the efficiency and accuracy of the mobile robot's behavioral decision-making.

[0026] According to another aspect of the present application, a computer-readable medium is also provided, on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to enable the processor to implement the above-mentioned robot behavior decision method.

[0027] The present application also proposes a robot behavior decision-making device that simulates the hippocampal memory priority playback function, and the robot behavior decision-making device includes:

[0028] one or more processors;

[0029] A computer readable medium for storing one or more computer readable instructions,

[0030] When the one or more computer-readable instructions are executed by the one or more processors, the one or more processors implement the above-mentioned robot behavior decision-making method.

[0031] The working principle of the present invention is that the present application combines the working mechanism of memory encoding, memory playback and guiding the behavior decision of organisms through memory in the brain of the hippocampus-entorhinal cortex, and establishes a robot brain-like memory priority playback model that simulates the memory mechanism of the hippocampus-entorhinal cortex. The working principle of the model can be described as follows: During the training process, the robot explores the environment, collects environmental information and its own motion information, and then transmits the information to the encoding area. After being encoded by the cellular computing model, it is stored in the memory storage area to form a memory pool for the robot's task execution process. When the memory pool stores a certain threshold or other conditions (such as the robot resting or encountering a major event), the robot will replay the memory, adjust the stored memory or seek solutions for related events. When the robot enters an unfamiliar environment, it first obtains the current environmental information. If the current information exists in the memory pool, the corresponding decision-making plan is enabled. Otherwise, the robot randomly selects a set of actions to execute.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] (1) The present invention designs a memory priority replay mechanism, using the DQN priority replay algorithm to simulate the working mechanism of the human brain to selectively replay memories based on their importance. The model will determine the priority of the memory based on its contribution to strategy improvement and the importance of the memory itself, and will be given a higher weight during replay, which not only ensures that the model focuses on these key data during training, but also accelerates the learning and training process.

[0034] (2) The present invention simulates the behavior of time cells in the hippocampus in continuous time and the unique time encoding characteristics in the memory process, and designs a time cell mechanism. The time cell mechanism not only focuses on important memories, but also takes into account the long-term impact of memories in time series. Through this mechanism, the intelligent agent can more accurately identify and preferentially replay those sparse memories that are important in the time series. By allocating more learning resources to these key memories, it is ensured that they receive sufficient attention during the learning process, so that the intelligent agent can extract valuable information from these rare memories that are easily overlooked, thereby improving its ability to adapt in new environments.

[0035] In summary, the robot brain-like memory priority playback model proposed in the present invention that simulates the memory playback mechanism of the hippocampus-entorhinal cortex has good convergence and generalization, enabling the robot to quickly adapt to unfamiliar environments and complete corresponding tasks more efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0037] Figure 1 A schematic diagram of the system framework of a robot brain-like memory priority playback model simulating the hippocampus-entorhinal cortex memory mechanism proposed in an embodiment of the present invention is shown;

[0038] Figure 2 A schematic diagram of the network structure of the brain-like memory priority playback model of the hippocampus-entorhinal cortex in an embodiment of the present invention is shown;

[0039] Figure 3 The simulation experiment environment in the embodiment of the present invention is shown;

[0040] Figure 4 The optimal path diagram of the agent in a static experimental environment is shown when four models are respectively used in the embodiment of the present invention;

[0041] Figure 5 The static experimental environment in the embodiment of the present invention is shown, wherein: Figure 5 (a) is a static physical experimental environment. Figure 5 (b) is the description of the static experimental environment in Rviz;

[0042] Figure 6 A comparison chart of the number of movement steps of robots using four models in 30 running processes in an embodiment of the present invention is shown;

[0043] Figure 7 The figure shows the path conditions presented when the four algorithms in the embodiments of the present invention successfully reach the target point at one time;

[0044] Figure 8 The dynamic experimental environment in the embodiment of the present invention is shown, wherein: Figure 8 (a) is a dynamic physical experimental environment. Figure 8 (b) Description of the dynamic experimental environment in Rviz;

[0045] Fig. 9 A certain successful path of the robot using four models in an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0046] The following will be combined with the attached Figures 1 to 9 The technical solution of the present invention is described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] The robot behavior decision-making method proposed in this application to simulate the hippocampal memory priority playback function includes the following steps:

[0048] S1. Simulate the behavior of time cells in the hippocampus in continuous time and the unique time coding characteristics in the memory process, further encode the information encoded by place cells and grid cells in the hippocampus-entorhinal cortex structure, and engrave time information on it;

[0049] S2. Introduce the memory priority replay mechanism and use the DQN priority memory replay algorithm to simulate the working mechanism of the human brain that selectively replays memories based on their importance.

[0050] The combination of the above two works established a robotic brain-like memory priority replay model that simulates the memory mechanism of the hippocampus-entorhinal cortex. This model can not only accelerate the learning process of the intelligent agent, but also improve the accuracy of its decision-making.

[0051] Based on the above neurobiological mechanism, this embodiment is designed as follows Figure 1 The system framework of the mobile robot behavior decision model in an unknown environment is shown. The system framework of the robot brain-like memory priority playback model includes:

[0052] Mobile robots can obtain environmental and motion information of their environment;

[0053] The coding region includes grid cells, time cells and place cells, and the grid cells, time cells and place cells work together to encode the environmental information and motion information;

[0054] Memory storage area, used to store encoded information;

[0055] The memory playback area is used to retrieve information in the memory storage area. When relevant memories are retrieved, a designated decision behavior is performed according to the relevant memories. When relevant memories are not retrieved, a random decision behavior is performed.

[0056] Among them, the grid cells, time cells and place cells in the coding area can simulate the working mechanism of the corresponding cells in the biological hippocampus-entorhinal cortex to realize the encoding of complex sequence information, and the memory playback area can simulate the priority selection mechanism of important memories during the biological memory playback process to improve the efficiency and accuracy of the mobile robot's behavioral decision-making.

[0057] The working principle of the above system framework can be described as follows: During the training process, the robot explores the environment, collects environmental information and its own motion information, and then transmits the information to the encoding area. After being encoded by the cellular computing model, it is stored in the memory storage area to form a memory pool for the robot's task execution process. When the memory pool stores a certain threshold or other conditions (such as the robot resting or encountering a major event), the robot will replay the memory, adjust the stored memory, or seek solutions for related events. When the robot enters an unfamiliar environment, it first obtains the current environmental information. If the current information exists in the memory pool, the corresponding decision plan is enabled. Otherwise, the robot randomly selects a set of actions to execute.

[0058] Based on the above system framework, combined with the working mechanism of memory encoding, memory playback and guiding the behavior decision of organisms through memory in the hippocampus-entorhinal cortex in the brain, the robot brain-like memory priority playback model that simulates the memory mechanism of the hippocampus-entorhinal cortex established in this application is as follows: Figure 2 The working process of this model is mainly divided into two stages: learning and training stage and task execution stage.

[0059] (1) Learning and training phase

[0060] In the learning and training phase, the robot's work is divided into two steps: one is the collection and storage of external information, and the other is to organize the collected information into memory (memory playback process), and use memory to complete training. In the learning phase, first, the robot explores the environment in a reinforced learning manner, and continuously stores the explored information in the memory pool. In this process, the place cells and grid cells in the hippocampus-entorhinal cortex structure encode environmental information and self-motion information, and the time cells encode time series information, which is gradually converted into memory and stored in the memory pool. Then, when the memory pool is full or the robot completes the exploration task, the robot enters the learning and training mode of memory priority playback.

[0061] In the process of external information collection and storage, the robot first needs to initialize the basic input parameters, including the maximum number of iterations T, the current number of rounds n, the number of learning times l, the learning rate η, the reward discount factor γ, the priority index α and the deviation correction coefficient β, the network update interval K, the memory pool capacity N, the number of samples drawn each time k, and create a value network Q based on the above parameters value and the target network Q target . Initialize the memory pool at the same time Set the initial cumulative weight Δ to 0 and the initial priority p1 to 1.

[0062] Afterwards, the robot uses reinforcement learning to explore the environment and obtain the current state s in the environment. t , and determine the action a to be taken based on the existing memory and action strategyt , according to this action, enter the next state s in the environment t+1 , and get the reward r of environmental feedback t , and then get the state transition sequence {s t ,a t ,s t+1}. After that, the state transition sequence is encoded and stored in the memory pool as the history of the agent's interaction with the environment. After executing a state transition, the model's Q function uses a single-step update strategy to update the Q value of the current state-action pair, and updates the value network Q according to the generated Q value. value With the target network Q target , as shown in formulas (1) and (2):

[0063]

[0064] Among them: S and A represent the state space and action space of the agent in the actual scene when it is awake; S' represents the new state space that the agent transfers to after executing action space A; R is the immediate reward space obtained by the agent from the environment after executing action space A; P represents the probability of the agent's state transition in the environment, which is generated by the historical statistics of the environment interaction when the agent is running. γ is a discount factor, which is used to represent the impact of delayed rewards on the total expected reward. After that, the reinforcement signal δ is calculated, and the robot selects an action a t Reward t The difference between the reward obtained when the action is actually performed, that is, the reward prediction error, is used to represent the reinforcement signal δ, that is, formula (3):

[0065] δ=rQ(s,a) (3)

[0066] The robot updates the action value and selects the action based on the reward prediction error. Therefore, the reinforcement signal δ based on the reward prediction error t It can be described as formula (4):

[0067] δ t =R+γQ target (S t ,argmax a Q(S t ,a))-Q(S t-1 ,A t-1 ) (4)

[0068] In the subsequent stages, when the priority memory playback method is combined with the deep reinforcement learning algorithm, formula (4) will become the standard for measuring sample priority. The larger the δ, the higher the sample priority, and the greater the sampling probability of the sample in the subsequent learning process.

[0069] During the robot's exploration of the environment, the relevant information (S, A, R, γ, S', t) is stored in the memory pool in time sequence. and give the initialization priority p t+1 =max i<t+1 p i If the stored information does not exceed the capacity of the memory pool and the maximum number of iterations of the training task has not been reached, the robot does not enter the memory playback phase. Otherwise, the robot starts the memory playback process to learn.

[0070] During the memory playback process, k samples are extracted each time for playback, and the priority of each sample is calculated during the sampling process. The present invention introduces a random sampling method between pure greedy priority and uniform random sampling. First, it is necessary to ensure that the sampling probability is proportional to the priority of the sample, that is, the probability of a sample with a high priority being sampled is greater. Secondly, it is ensured that even if the priority of a sample is very low, it still has a chance to be drawn and will not be completely ignored. The sampling probability is defined here as:

[0071]

[0072] where p i is the priority of the i-th sample in the memory pool. α is the priority index, and α=0 corresponds to the uniform situation. λ∈(0,1] is the attenuation factor used to control the priority. In the algorithm adopted by the present invention, the sampling priority is related to the reward value and the temporal difference (TD) error of each state-action pair. If the time series is relatively early, the reward value has a greater impact on its sampling probability during the memory playback process of the agent; otherwise, the temporal error has a greater impact on its sampling probability, thereby measuring the importance of learning memory at each different decision moment. Assume that the learning memory at the decision moment j is a training data in the memory pool, and its priority is defined as formula (6):

[0073] p j ←|δ j |+ε (6)

[0074] Where ε is a small positive number, which is used to avoid the TD error of learning memory to be 0. Since memory priority playback will introduce deviations, dynamic weights need to be introduced for balance. The dynamic weight ω corresponding to the decision time j is j The definitions are as follows:

[0075]

[0076] Where β is the bias correction coefficient of importance sampling, which aims to correct the bias of priority replay sampling during learning. Priority replay tends to extract samples with large errors multiple times, which may lead to bias in the learning process, so the bias correction coefficient β is used for correction. As training progresses, β will gradually increase and eventually tend to 1. After that, the gradient information is accumulated, the accumulated weights are updated and their changes are calculated to prepare for subsequent parameter updates:

[0077]

[0078] in, Represents the gradient operator about the parameter θ. After that, the network weights are updated according to the parameters, and the target network is continuously updated according to the new network weights. After the memory replay is completed, the agent selects a new action space, sorts the new action space and state space according to the time series information, and re-stores the updated state-action pairs into the memory pool. Based on the above description, a DQN priority memory replay algorithm based on time series is proposed to simulate the working method of the robot during learning and training. Its pseudo code is as follows:

[0079]

[0080] (2) Task execution phase

[0081] After the learning and training phase, the robot enters the task execution phase, where it completes the exploration task in the new environment based on the trained network. When exploring a new environment, the agent first obtains the current state S current , then calculate the current state S current With the stored history state S memory The present invention sets the similarity index C slt as follows:

[0082]

[0083] If C slt ≥ρ, ρ is the set similarity threshold, indicating that the current state has never appeared in the previous task process, and the network needs to create a new cognitive node and store the group of state information in the memory pool. On the contrary, it means that the current state is experienced by the agent, and the system will extract the decision plan corresponding to the current state to guide the agent to take the next action.

[0084] The working principle of the present invention is that the present application combines the working mechanism of memory encoding, memory playback and guiding the behavior decision of organisms through memory in the brain of the hippocampus-entorhinal cortex, and establishes a robot brain-like memory priority playback model that simulates the memory mechanism of the hippocampus-entorhinal cortex. The working principle of the model can be described as follows: During the training process, the robot explores the environment, collects environmental information and its own motion information, and then transmits the information to the encoding area. After being encoded by the cellular computing model, it is stored in the memory storage area to form a memory pool for the robot's task execution process. When the memory pool stores a certain threshold or other conditions (such as the robot resting or encountering a major event), the robot will replay the memory, adjust the stored memory or seek solutions for related events. When the robot enters an unfamiliar environment, it first obtains the current environmental information. If the current information exists in the memory pool, the corresponding decision-making plan is enabled. Otherwise, the robot randomly selects a set of actions to execute.

[0085] The key of the present invention is to propose a robot brain-like behavior decision model that simulates the working mechanism of time cells in the hippocampus and memory priority playback, and adopts the DQN priority memory playback algorithm to simulate the working mechanism of hippocampal-entorhinal cortex priority memory playback. At the same time, combined with the encoding method of time cells for time series information, a DQN priority memory playback algorithm based on time series is proposed. Experiments show that the algorithm has good convergence and generalization, allowing the robot to quickly adapt to unfamiliar environments and complete corresponding tasks more efficiently.

[0086] In order to enable researchers and application personnel in this technical field to better understand the scheme of the present invention, the simulation experiment results of the scheme will be analyzed below, which is also a verification of the specific application scenario of the present invention.

[0087] Simulation experiment verification

[0088] 1.1 Simulation experiment design and result analysis

[0089] (1) Environment settings

[0090] Set up a 500×500 coordinate environment, randomly place 35 obstacles (red solid circles) in the environment, set the agent's initial position (yellow solid circle) to (50,50), and the target point (green solid circle) to (450,450). In addition, set special L-shaped obstacles near the starting position and target point, such as Figure 3As shown. There are no obstacles within the safety range of the robot's starting position. Since the positions of obstacles are randomly set, the robot does not have prior knowledge of the environment, which is equivalent to the robot being in an unknown environment. The robot brain-like memory priority playback model that simulates the hippocampus-entorhinal cortex memory mechanism of the present invention is denoted as DQN-PERTC. In order to verify the performance of the DQN-PERTC algorithm, it is compared with the deep reinforcement learning algorithms DQN, DDQN and SAC algorithms to analyze the learning and decision-making capabilities of the four algorithms in an unknown environment.

[0091] (2) Parameter settings

[0092] In the DQN-PERTC algorithm, the state of the input network is a two-dimensional matrix obtained by mapping the grid map; a two-layer linear network with bias parameters is used, each layer of the network has 128 neurons, and the weight parameters of the neurons are sampled from a normal distribution with a mean of 0 and a variance of 0.5, and the bias correction parameter β is set to a constant of 0.3; the memory pool stores 1000 memories, and 50 memories are extracted from a uniform distribution for learning each time the network is trained, and the root mean square loss function and gradient descent algorithm (learning rate 0.01) are used for learning; the value network parameters are passed to the target network every 20 runs. The parameters of the DQN, DDQN and SAC algorithm networks and memory pools are the same as those of the DQN-PERTC algorithm, and the other parameters unique to this method are adjusted to the optimal as much as possible.

[0093] In addition, the reward function in the comparative experiment is set as follows: the agent reaches the end point and receives a reward value of r = 100; if the agent collides with an obstacle or a wall, the reward value is r = -20; if the distance from the next position of the agent to the end point is less than the distance from the current position to the end point, the reward value is r = 10, the purpose is to encourage the agent to approach the end point. The placement of the reward function in the environment is consistent for all algorithms. Using a reward value or penalty value with a larger absolute value is conducive to better updating the Q function or neural network.

[0094] (3) Simulation Experiment Results and Analysis

[0095] Figure 4 The figure shows the optimal path diagram of the agent in a static environment when using four models. The agent using the algorithm of the present invention only takes 77 steps to reach the target point, while the agents using the other three algorithms take 83, 79 and 79 steps respectively. In comparison, the behavior decision of the algorithm of the present invention is better.

[0096] In order to further verify the performance of the algorithm of the present invention, four more experiments were conducted. Each group randomly generated 100 highly complex environments. The starting point and target point of the agent in each group of environments were at the same position. The agent was tested using the four algorithms respectively, and the number of times the agent successfully reached the target point in each group of experiments was recorded. The results are shown in Table 1.

[0097] Table 1 Random environment test results

[0098] Group 1 Group 2 Group 3 Group 4 DQN-PERTC 90 88 89 89 DQN 81 79 80 79 DDQ 84 85 82 83 SAC 83 84 82 82

[0099] From Table 1, we can see that the DQN-PERTC algorithm has the most successful times in the four groups of experiments. Figure 4 The comparative experimental results show that the decision path and stability of the intelligent agent using the DQN-PERTC algorithm are better, indicating that the model has stronger environmental adaptability and strong robustness.

[0100] The reasons are as follows: First, the DQN-PERTC algorithm designs a priority playback strategy based on time series, which can give priority to identifying and processing those memories that contribute the most to strategy improvement during the playback process. The model will determine the priority of the memory by the contribution of the memory to the strategy improvement and the importance of the memory itself. At the same time, it is given a higher weight during playback, which not only ensures that the model focuses on these key data during training, but also accelerates the learning and training process. Secondly, the DQN-PERTC algorithm not only pays attention to the timing error and reward value through the time cell mechanism, but also considers the time background of the memory, that is, the time point when these memories occur and their position in the time series. If the time series is relatively early, the reward value has a greater impact on its sampling probability during the playback process; conversely, the timing error has a greater impact on its sampling probability. In this way, for some events that do not occur frequently but are very important in the environment, the DQN-PERTC algorithm can also allocate sufficient learning resources to these events to ensure that they can get enough attention during the learning process, so that the intelligent agent can extract valuable information from these key memories that are easily overlooked, and then show stronger adaptability when facing new and unknown environments.

[0101] 1.2 Static physical experiment design and result analysis

[0102] (1) Experimental environment

[0103] This experimental platform consists of an Autolabor Pro1 robot, The PC is composed of Ryzen33200G processor and DDR48GB, and the program running environment is Python3.6. Place the AutolaborPro1 robot in the following Figure 5 The navigation experiment is carried out in the physical environment shown in the figure. The experimental environment is a rectangular environment of 6m*4.5m. The static environment contains 5 static obstacles of different sizes and two L-shaped complex obstacles.

[0104] The robot's starting point is set to the upper right corner of the map, and the target is set at (5,3.5), which is located in the lower left corner of the environment. Because the radar scanning point is inconsistent with the robot's own collision distance, the radar is located at the robot's head. Considering the robot's own volume and safety distance, the robot's own volume must be considered when judging whether it collides with an obstacle. Therefore, all distances in this experiment refer to the distance between the robot body and the obstacle or target point, not the distance measured by the radar. When the distance between the robot and the target point is less than 0.2m, the robot stops running and the experiment ends.

[0105] (2) Static test results and analysis

[0106] exist Figure 5 In the static actual environment shown in the figure, the four algorithms DQN-PERTC (algorithm of the model of the present invention), DQN, DDQN, and SAC were tested in the actual environment. Each algorithm was run 30 times in the static physical environment, and the number of movement steps used by the robot to reach the target position and the success rate of reaching the target were recorded. The success rate, number of movement steps, and other results of each algorithm in the 30 experiments are recorded in Table 2. And the results of each of the 30 experiments are recorded in Figure 6 middle.

[0107] Table 2 Comparative experimental records

[0108]

[0109]

[0110] From the average number of steps in Table 2, we can see that the minimum number of movement steps and the average number of movement steps of the robot using the DQN-PERTC algorithm are both the smallest, that is, the robot can find a better decision path. When using the DQN-PERTC algorithm, the average number of successful navigation steps of the robot reaches 37.79, and the variance of the number of steps is only 0.74, indicating that it has good stability and performs better than the DQN, DDQN and SAC algorithms. Figure 7 The paths taken by the four algorithms when they successfully reach the target point are recorded.

[0111] From the variance analysis in Table 2, the DQN-PERTC algorithm has the best performance. DDQN and SAC are not much different, but its performance is much better than DQN. Although the starting point and target point of the robot are the same, DQN has greater randomness than other algorithms, resulting in the worst performance. From the number of successful navigation and the success rate, the DQN-PERTC algorithm has a higher success rate because it is better than several other algorithms in terms of stability. From the comparative experimental results, it can be seen that the decision path and stability of the DQN-PERTC algorithm are better, indicating that it has stronger learning ability.

[0112] The reason is that the design concept of the DQN-PERTC algorithm aims to significantly improve the learning efficiency and adaptability of the agent through an innovative priority replay strategy and time cell mechanism. First, the algorithm introduces a priority replay strategy based on time series. The key to this strategy is to accurately identify and process those memories that contribute the most to policy optimization. During the replay process, the model prioritizes memories with higher reward values ​​or larger timing errors for processing. This not only ensures that the model can focus on learning key data during training, but also accelerates the learning process by giving these important memories a higher replay weight. Secondly, the DQN-PERTC algorithm further improves the replay strategy by introducing the time cell mechanism. The time cell mechanism not only focuses on timing errors and reward values, but also takes into account the long-term impact of memory in time series. Through this mechanism, the agent can more accurately identify and prioritize replaying sparse memories that are important in time series. By allocating more learning resources to these key memories, the time cell mechanism ensures that they receive sufficient attention during the learning process, allowing the agent to extract valuable information from these rare memories that are easily overlooked, thereby improving its ability to adapt in new environments. The combination of the above two factors enables the agent to accurately replay memories, allowing it to not only focus on memories that contribute more, but also select different priority strategies based on the time series. This enables the agent to not only have higher learning efficiency, but also have a better ability to adapt to changing environments.

[0113] 1.3 Dynamic physical experiment design and result analysis

[0114] (1) Dynamic experimental environment

[0115] In addition to the five static obstacles of different sizes, the dynamic environment also contains two dynamic obstacles, which are composed of an experimenter and another AutolaborPro1 robot. The two move randomly in the environment to determine the algorithm's ability to avoid obstacles in a dynamic environment.

[0116] (2) Experimental results and analysis

[0117] exist Figure 8 In the dynamic real environment shown in the figure, the DQN-PERTC algorithm (algorithm of the present invention) and DQN, DDQN and SAC were tested in the real environment. The robot was run 30 times in the dynamic real environment, and the number of steps taken by the robot to reach the target position and the success rate of reaching the target were recorded. The results are shown in Table 3.

[0118] Table 3 The number of successes of the robot using four algorithms in a dynamic environment

[0119] DQN-PERTC DQN DDQ SAC Number of successes 29 22 24 24 Success rate 96.6% 73.3% 80% 80%

[0120] As shown in Table 3, in the dynamic test environment, the DQN-PERTC algorithm has a higher success rate in terms of the number of successful navigations and the success rate, because it is superior to the other algorithms in terms of adaptability.

[0121] Fig. 9 The following figure shows a successful path of a robot using four algorithms. Fig. 9 As shown in the figure, the red and yellow dots in the image are obstacle information returned by the scan / topic, and the black area is the obstacle RVIZ image information drawn by the map / topic. From the robot's motion path and dynamic obstacle information recorded by RVIZ, it can be seen that when the robot encounters an obstacle, it has obvious avoidance behavior. This shows that in an environment with dynamic obstacles, the robot can effectively avoid static obstacles and dynamic obstacles and successfully navigate to the target point. Combined with Table 3 and Fig. 9 It can be seen that the decision path and stability of the DQN-PERTC algorithm are better, indicating that the robot has stronger learning ability. The reason is that the DQN-PERTC algorithm has better adaptability to dynamically changing environments with the introduction of time cells. Experiments in dynamic environments further verify the effectiveness and reliability of the algorithm of the present invention.

[0122] It should be noted that the present application can be implemented in software and / or a combination of software and hardware, for example, can be implemented using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present application (including relevant data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to perform each step or function.

[0123] In addition, a part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. The program instruction for calling the method of the present application may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium, and / or stored in a working memory of a computer device that runs according to the program instruction. Here, according to an embodiment of the present application, a device is included, the device including a memory for storing computer program instructions and a processor for executing program instructions, wherein, when the computer program instruction is executed by the processor, the device is triggered to run the method and / or technical solution based on the aforementioned multiple embodiments according to the present application.

[0124] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic features of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present application is limited by the attached claims rather than the above description, so it is intended to include all changes that fall within the meaning and scope of the equivalent elements of the claims in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

Claims

1. A robot behavior decision-making method simulating the priority playback function of hippocampal memory, characterized in that: The steps include: S1. Simulate the behavior of time cells in the hippocampus in continuous time and the unique time coding characteristics in the memory process, further encode the information encoded by place cells and grid cells in the hippocampus-entorhinal cortex structure, and engrave time information on it; S2. Introduce the memory priority replay mechanism, use the DQN priority memory replay algorithm to simulate the working mechanism of the human brain that selectively replays memories according to their importance, and establish a robotic brain-like memory priority replay model that simulates the hippocampus-entorhinal cortex memory mechanism.

2. The robot behavior decision-making method according to claim 1, characterized in that: In step S2, the working process of the robot brain-like memory priority playback model includes a learning and training phase and a task execution phase; Among them, in the learning and training stage, the robot's work is divided into two steps: one is to collect and store information from the outside world, and the other is to organize the collected information into memory and use the memory to complete training; After the learning and training phase, the robot enters the task execution phase, where it completes exploration tasks in a new environment based on the trained network.

3. The robot behavior decision-making method according to claim 2, characterized in that: The learning and training phase includes the following steps: S31. The robot explores the environment in a reinforcement learning manner and continuously stores the explored information into the memory pool. In this process, the place cells and grid cells in the hippocampus-entorhinal cortex structure encode environmental information and self-motion information, and the time cells encode time series information, which are gradually converted into memory and stored in the memory pool; S32. When the memory pool is full or the robot completes the exploration task, the robot enters the learning and training mode with memory priority playback.

4. The robot behavior decision-making method according to claim 3, characterized in that: When collecting and storing external information, the robot first needs to initialize the basic input parameters, including the maximum number of iterations, the current number of rounds, the number of learning times, the learning rate, the reward discount factor, the priority index and the deviation correction coefficient, the network update interval, the memory pool capacity, and the number of samples drawn each time. Based on the above parameters, the value network Qvalue and the target network Qtarget are created, and the memory pool is initialized, with the initial cumulative weight set to 0 and the initial priority set to 1.

5. The robot behavior decision-making method according to claim 4, characterized in that: In step S31, the robot's intelligent agent obtains the current state in the environment, and determines the action to be taken based on the existing memory and action strategy, enters the next state in the environment based on this action, and obtains the reward of environmental feedback, thereby obtaining a state transition sequence. The state transition sequence is encoded and stored in the memory pool as the history of the intelligent agent in the process of environmental interaction. After executing a state transfer, the Q function of the robot's brain-like memory priority playback model adopts a single-step update strategy to update the Q value of the current state-action pair, and updates the value network Qvalue and the target network Qtarget according to the generated Q value.

6. The robot behavior decision-making method according to claim 5, characterized in that: In the task execution phase, when the robot explores the new environment, the agent first obtains the current state S current , then calculate the current state S current With the stored history state S memory Similarity, set the similarity index C slt The calculation formula is as follows: If C slt ≥ρ, ρ is the set similarity threshold, indicating that the current state has never appeared in the previous task process. The network needs to create a new cognitive node and store the group of state information into the memory pool. Otherwise, it means that the current state is experienced by the agent. The system will extract the decision plan corresponding to the current state to guide the agent to take the next action.

7. The robot behavior decision-making method according to claim 1, characterized in that: The system framework of the robot brain-like memory priority playback model includes: Mobile robots can obtain environmental and motion information of their environment; The coding region includes grid cells, time cells and place cells, and the grid cells, time cells and place cells work together to encode the environmental information and motion information; Memory storage area, used to store encoded information; The memory playback area is used to retrieve information in the memory storage area. When relevant memories are retrieved, a designated decision behavior is performed according to the relevant memories. When relevant memories are not retrieved, a random decision behavior is performed. Among them, the grid cells, time cells and place cells in the coding area can simulate the working mechanism of the corresponding cells in the biological hippocampus-entorhinal cortex to realize the encoding of complex sequence information, and the memory playback area can simulate the priority selection mechanism of important memories during the biological memory playback process to improve the efficiency and accuracy of the mobile robot's behavioral decision-making.

8. A robot behavior decision-making device simulating the priority playback function of hippocampal memory, characterized in that: The robot behavior decision-making device includes: one or more processors; A computer readable medium for storing one or more computer readable instructions, When the one or more computer-readable instructions are executed by the one or more processors, the one or more processors implement the robot behavior decision-making method according to any one of claims 1 to 7.