Robot behavior decision-making method and device for simulating brain learning and memory mechanism

By combining the developmental network DN2, reinforcement learning and self-organized map SOM, the functions of the cerebellum, basal ganglia and hippocampus learning system are simulated, and the SORLDN2 model is constructed, which solves the problems of low efficiency and no consideration of hippocampus learning system in existing brain-like models, and realizes rapid learning and efficient decision-making of agents in complex environments.

CN119927902AActive Publication Date: 2025-05-06ZHENGZHOU UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510063270.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Existing brain-like models are inefficient in complex learning tasks, lack of transverse connections of hidden layer neurons, and do not consider hippocampal learning system, resulting in slower robot learning and decision-making.

Method used

By combining the developmental network DN2, reinforcement learning and self-organized map SOM, the functions of the cerebellum, basal ganglia and hippocampus learning systems are simulated, and a self-organized reinforcement learning model SORLDN2 based on the developmental network DN2 is constructed, and decision weights are balanced using timing differential errors and synergistically evaluated action value.

Benefits of technology

It significantly improves the accuracy and efficiency of agent behavior decision-making, improves resource utilization, convergence speed and decision-making quality, and can quickly learn and obtain optimized decision-making solutions in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119927902A_ABST
    Figure CN119927902A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of robot intelligent control, and relates to a robot behavior decision-making method and device for simulating a brain learning and memory mechanism. According to the method, a development network DN2, reinforcement learning and a self-organizing map SOM are combined, functions of a cerebellum learning system, a basal ganglion and a hippocampus learning system are simulated, and a self-organizing reinforcement learning model based on the development network DN2 is obtained; the equipment comprises one or more processors and a computer readable medium in which one or more computer readable instructions are stored, and is used for realizing the robot behavior decision-making method. According to the method, the functions of the cerebellum learning system, the basal ganglion and the hippocampus learning system are properly utilized, and the fusion algorithm combines reinforcement learning and supervised learning methods, so that the robustness of the self-organizing reinforcement learning model based on the development network DN2 is enhanced, the capability of the algorithm for obtaining an effective obstacle avoidance strategy is improved, and the method is suitable for being applied to the self-organizing reinforcement learning of the self-organizing reinforcement learning of the self-organizing reinforcement learning of the self-organizing reinforcement learning. And the intelligent agent can be helped to quickly learn an efficient solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computers, and in particular to the field of robot intelligent control technology, and more specifically to a robot behavior decision-making method and device that simulates the learning and memory mechanism of the brain. Background Art

[0002] With the development of science and technology, building reliable brain-like models for application in the field of intelligent robots so that they can efficiently complete complex tasks has become an important research direction in the fields of artificial intelligence and neuroscience.

[0003] In the exploration of the cognitive functions of the brain, the key role of the basal ganglia and cerebellum in motor control and decision-making has been confirmed by many neuroscience research studies. Researchers have carried out a lot of work on robot learning that simulates basal ganglia and cerebellum models. In 1999, Doya et al. proposed a neural network that combines different learning modules in the cerebellum, basal ganglia, and cerebral cortex, giving the cerebellum supervised learning, basal ganglia reinforcement learning, and cerebral cortex unsupervised learning functions. Since then, related research has continued to advance, such as the guided action-dependent heuristic dynamic programming learning mechanism developed by Ruan et al. in 2012, which uses the Actor-Critic model to simulate the functions and cooperation mechanisms of the two; in 2024, Zhang et al. proposed a biologically inspired hybrid model for musculoskeletal robots, exploring the balance between exploration and utilization in reinforcement learning.

[0004] In addition to motion control, researchers have also explored the role of the cerebellum and basal ganglia in intelligent agent behavior decision-making. For example, some researchers combined association-based learning and reward-based learning, and used the Actor-Critic framework of reinforcement learning to simulate the functions of the two to complete the robot's goal-oriented decision-making task; Wang et al. used the Modulatory Developmental Network (MDN) and Q-learning based on radial basis function networks to simulate the functions of the cerebellum and basal ganglia, respectively, demonstrating the potential of neural regulation models in mobile robot behavior decision-making, but there are also certain problems.

[0005] At the same time, the important role of brain regions such as the hippocampus and prefrontal cortex in memory encoding, retrieval and decision-making has gradually become clear. Related research results, such as the hippocampus-striatum system constructed by Yuan et al. and the biologically inspired hippocampal sequence memory system constructed by Casanueva-Morato et al., have provided important references for brain-inspired artificial intelligence research, and the new complementary learning system theory has revealed the benefits of the complementary characteristics between the hippocampus and the neocortex in achieving complex behaviors.

[0006] As an unsupervised learning method, the pattern separation property of Self-Organizing Map (SOM) is widely used to improve model performance. For example, Notsu et al. combined it with differential learning to generate state space for reinforcement learning, Su et al. combined it with deep reinforcement learning technology to solve the problem of detecting anomalies in multivariate time series data, and Modhej et al. used it to simulate the pattern separation function of the hippocampus to solve the problem of deep learning in handwriting recognition.

[0007] However, the current brain-like models still have many shortcomings. Taking the research of Wang et al. as an example, the regulatory developmental network of the cerebellum simulation part is similar to the developmental network-1 (DN1), which only uses the low-level representation part (that is, the 100,001 two types of neurons in the Y layer), which is inefficient in complex learning tasks and the hidden layer neurons lack lateral connections; in the reinforcement learning part that simulates the function of the basal ganglia, the intelligent agent needs to retrieve the entire state space to perform actions, and the computing resources are consumed in a complex environment, which slows down the robot's learning and decision-making speed; in addition, the model does not consider the hippocampus learning system. Therefore, it is urgent to establish a more complete and efficient brain-like model to meet the needs of practical applications, which is also an important background and starting point for the research and development of this patent. Summary of the invention

[0008] The purpose of this application is to provide a robot behavior decision-making method and equipment that simulates the learning and memory mechanism of the brain. By cleverly integrating the functions of the cerebellum, basal ganglia and hippocampus learning systems, an efficient decision-making model is provided for the intelligent agent, which can help the intelligent agent to quickly learn and obtain optimized decision-making plans in complex environments, significantly improve the accuracy and efficiency of the intelligent agent's behavior decision-making, and achieve significant progress in resource utilization, convergence speed and decision quality compared to the existing technology. It has broad application prospects and market value.

[0009] Based on one of the above purposes, the present application proposes a robot behavior decision-making method that simulates the learning and memory mechanism of the brain. The robot behavior decision-making method combines the developmental network DN2, reinforcement learning, and the self-organizing map SOM to simulate the functions of the cerebellum learning system, the basal ganglia, and the hippocampus learning system, and obtains a self-organizing reinforcement learning model based on the developmental network DN2 (abbreviated as "SORLDN2");

[0010] In most methods of simulating hippocampal function, the hippocampal learning system cannot be constructed alone, and often needs to be combined with another method to realize its function. The self-organizing reinforcement learning model based on the developmental network DN2 in this application first uses the developmental network DN2 as the cerebellum learning system to realize the supervised learning function of the cerebellum, and then combines the self-organizing map and reinforcement learning, and uses the SORL (Self-Organized Reinforcement Learning) model to describe the basal ganglia-hippocampus learning system, denoted as SORL, and then uses the temporal differential error from the developmental network DN2 to train SORL, so that the self-organizing reinforcement learning model based on the developmental network DN2 can store the state that the developmental network DN2 is difficult to accurately evaluate, and then balance the decision weights between the cerebellum learning system and the basal ganglia-hippocampus learning system by simulating the brain's familiarity with the state input, and then collaborate in evaluating the value of actions, so that the intelligent agent can make full use of the respective advantages of the cerebellum and the basal ganglia-hippocampus learning system to select actions.

[0011] Furthermore, Zhu et al. used DN1 that mimics the cerebellum to achieve good results in navigation decision-making tasks, and then Zheng and Wu proposed an upgraded developmental network DN2. The developmental network DN2 has three regions, including a perceptual input layer X from the external environment, a motor layer Z that acts on the external environment, and a hidden layer Y that organizes the interaction between the perceptual input layer X and the motor layer Z;

[0012] The input x(t) from the sensor input layer X, the z(t) from the motion layer Z, and the response y(t) from the hidden layer Y are represented as vectors, respectively. The adaptive part N = (V, G), where V is a matrix of synaptic weights, G is a vector of neuron activation ages, and N y and N z They represent the adaptive parts of the Y neurons in the hidden layer Y and the Z neurons in the motion layer Z respectively.

[0013] The main workflow of the above developmental network DN2 can be described as follows:

[0014] (1) Initialization. At time t = 0, input the environmental perception input x and the supervisory signal z, and the response value y(0) = 0. Set the synaptic weights V of all Y neurons in the hidden layer Y y and the synaptic weight V of the Z neuron in the motor layer Z z are random values. Since no neurons are activated, their neuron age is equal to zero.

[0015] (2) Response value and adaptive part calculation N. At time t = 1, ..., repeat the following two steps:

[0016] (a) Calculate the response values ​​of all Y neurons using the region function f y Update the adaptive part N y , as shown below:

[0017] (y(t),N′ y ) = f y (p y ,N y ) (1)

[0018] Among them, p y Represents a tuple of [x(t-1), y(t-1), z(t-1)], which is used as the input vector to update the response value y(t).

[0019] (b) Instructing the robot to learn. The motor layer Z has two working modes: instruction mode and response mode. The instruction mode teaches the robot to learn, and the response mode calculates the response value of the Z neuron and uses the area function f z Update the adaptive part N z , as shown below:

[0020] (z(t),N′ z ) = f z (p z ,N z ) (2)

[0021] Among them, p z =y(t-1) is the input vector for updating the response value z(t).

[0022] The calculation process of the area function is as follows:

[0023] By calculating the pre-response value r' of each neuron i , to determine the function f in formula (1) y And the function f in formula (2) z The process is as follows:

[0024] (1) The winning neuron is selected through the Top-k competition mechanism. Only the top k neurons with the largest response values ​​will be selected for activation, which simulates the lateral inhibition of neurons. Specifically, the pre-response value r' of neuron i in the hidden layer Y is i for:

[0025] r' i =r' b,i +r' t,i +r' l,i (3)

[0026] r' b,i =x(t-1) / ||x(t-1)||·v b,i / ||vb,i || (4)

[0027]

[0028] r' l,i =y(t-1) / ||y(t-1)||·v l,i / ||v l,i || (6)

[0029] In the formula, r' b,i , and r' l,i They represent the pre-response values ​​of the bottom-up connection, top-down connection, and lateral connection of a neuron, respectively, and v b,i , and v l,i Represent the weight vectors of these three connections respectively. If neuron i is one of the larger Top-k neurons, the response value of neuron i is set to y i =1, activation age n i Add 1: n i ←n i +1. If not, then y i = 0, the neuron age remains unchanged. The calculation method of the neuron response of the motor layer Z is similar to the above description.

[0030] (2) Hebbian rule updates synaptic weights. For the winning neuron i, its synaptic weights are updated according to the Hebbian learning rule as follows:

[0031] v b,i ←β b,1 *v b,i +β b,2 *y b,i *x(t-1)(7)

[0032]

[0033] v l,i ←β l,1 *v l,i +β l,2 *y l,i *y(t-1)(9)

[0034] Where, the activation age n i The determined β1 and β2 are the retention rate and learning rate of the neuron, β1+β2≡1, where

[0035] β1=(n i -1) / n i ,β2=1 / n i (10)

[0036] And β b1 , β b2 , β t1 , β t2 , β l1 , β l2 They respectively represent the retention rate and learning rate of a neuron from bottom to top, from top to bottom, and from horizontal connections.

[0037] (3) Split. If the pre-response value of the first winning neuron is lower than the threshold ρ(t) corresponding to the Y neuron type, the neuron in the initial stage is transferred to the active stage and then activated. ρ(t) is defined as

[0038] ρ(t)=μ(t)(θ-∈) (11)

[0039] Where μ(t) is a function used to adjust the speed of increasing the number of neurons. θ is a predefined constant. When the Y neuron type is 100, θ is 1, and when the Y neuron type is 011, θ is 2. ∈ is a constant.

[0040] (4) Synaptic maintenance: Keep stable connections where the deviation between weights and inputs is less than a threshold, and prune other connections.

[0041] Furthermore, the self-organizing reinforcement learning model based on the developmental network DN2 utilizes two types of Y neurons of the developmental network DN2 to assist the robot in completing navigation decision-making tasks, wherein the two types of Y neurons include lower-level Y neurons and higher-level Y neurons, the lower-level Y neurons are responsible for identifying different local features in the perception input layer X, and the higher-level Y neurons are connected to the lower-level Y neurons instead of being directly connected to the perception input layer X.

[0042] Furthermore, the basal ganglia-hippocampus learning system uses the temporal difference error generated by the cerebellum learning system to update its state representation to achieve rapid learning of the intelligent agent. The temporal difference error is a real-time adjustment factor of the learning rate and the standard deviation of the neighborhood function in the basal ganglia-hippocampus learning system.

[0043] Reinforcement learning belongs to the category of semi-supervised learning techniques. When an agent performs an action, it will receive a corresponding reward or punishment. The agent obtains the optimal solution to the task by maximizing the immediate reward of each state while evaluating actions that may bring greater rewards in the future.

[0044] Furthermore, the basal ganglia-hippocampus learning system allows individual memories to be stored in a unique, pattern-separated manner and promotes rapid updating of the system through a simple error-driven learning mechanism.

[0045] Initially, state sensory inputs are provided to the hippocampus, which stimulate the hippocampus in a pattern-separated manner, with neurons arranged in a grid to represent the level of excitation. t With weight w u The similarity between them is judged to activate neuron u t , calculated as follows:

[0046] u t =argmin u ||w u -s t || 2 (12)

[0047] In the formula, w u Represents the neuron in the input state s t and activate neurons u t For each neuron in the grid, there is a corresponding action value function Q SOM (u,a) realizes learning and memory, and sets the action value Q SOM The update method of (u,a) is:

[0048] Q SOM (u t ,a t )←Q SOM (u t ,a t )+αη t [y t+1 -Q SOM (u t ,a t )](13)

[0049] Where α is the learning rate. The weighting parameter η ensures that the action value will only get a large update when the closest neuron u is similar to the state value, which is very consistent with the way humans learn, that is, quickly learn similar things and slowly adapt to unfamiliar things. The setting rules for η are as follows:

[0050]

[0051] Among them, the parameter τ η Used to scale w u and t In addition, considering the cooperation mechanism between the cerebellum and the basal ganglia-hippocampus learning system, the calculation of the state-action value adopts the following weighted average rule:

[0052] Q(s,a)=ηQ SOM (u t ,a t )+(1-η)DN2(st ,a) (15)

[0053] It can be inferred that if the value of η is smaller, the contribution of DN2 is greater, that is, u is unfamiliar with s, otherwise the contribution of SORL is greater. The above weighted average rule (15) ensures a learning method in which the early state-behavior mainly depends on the developmental network DN2, and after learning, the contribution of SORL in the later stage is greater. The target value y set in formula (13) is:

[0054] y=r+γmax a' Q(s,a) (16)

[0055] Where r is the reward and γ is the discount factor. For learning in SORL, the target value y and the temporal difference error generated by DN2 are used to influence the update of the model. The weight of neuron i is updated using the following rule:

[0056]

[0057] Where λ is the learning rate. The exponential scaling parameter δ is:

[0058] δ=exp(|y-DN2(s,a)| / τ δ )-1 (18)

[0059] Here, the neighborhood function Represented by a Gaussian function

[0060]

[0061] Among them, the parameter τ δ Used to scale the timing difference error, u t and u i They represent the mapping positions of the winning neuron and the output neuron i respectively, and the neighborhood range of the neuron is determined by the parameter σ c , δ and σ.

[0062] Finally, the robot’s action is chosen based on the weighted average of the SORL and developmental network DN2 values.

[0063] The above method combines the developmental network DN2, reinforcement learning and self-organizing map SOM to simulate the functions of the cerebellum, basal ganglia and hippocampus learning systems.

[0064] Furthermore, the basal ganglia are responsible for implementing reinforcement learning functions, and the hippocampus is responsible for memorizing and encoding the knowledge learned by the intelligent agent.

[0065] Furthermore, the self-organizing reinforcement learning model based on the developmental network DN2 also includes:

[0066] The environment, which provides status and reward information;

[0067] The sensory cortex, which is responsible for representing input states;

[0068] the thalamus, responsible for action selection;

[0069] The motor cortex is responsible for movement output.

[0070] According to another aspect of the present application, a computer-readable medium is also provided, on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to enable the processor to implement the above-mentioned robot behavior decision method.

[0071] According to another aspect of the present application, a robot behavior decision-making device simulating the learning and memory mechanism of the brain is also provided, and the robot behavior decision-making device comprises:

[0072] One or more processors; a computer-readable medium for storing one or more computer-readable instructions, when the one or more computer-readable instructions are executed by the one or more processors, the one or more processors implement the above-mentioned robot behavior decision method.

[0073] The working principle of the present invention is that the present application proposes a self-organizing reinforcement learning model SORLDN2 based on the developmental network DN2, uses the developmental network DN2 as the cerebellum learning system, and describes the basal ganglia-hippocampus learning system through self-organizing reinforcement learning. In particular, the cerebellum learning system provides exploratory guidance for the basal ganglia-hippocampus learning system, alleviating the problem of slow convergence of reinforcement learning. The basal ganglia-hippocampus learning system uses the temporal difference error generated by the cerebellum learning system to update its state representation. Specifically, the temporal difference error is a real-time adjustment factor for the learning rate and the standard deviation of the neighborhood function in the basal ganglia-hippocampus learning system. This dynamic adjustment process enables the basal ganglia-hippocampus learning system to give priority to and retain the state memory that the cerebellum learning system predicts is not very accurate, and then use these memories to enhance overall decision-making ability and learning.

[0074] In addition, for the hippocampus, since experience is stored in a separate mode, the tabular method of representing the action-value function is more consistent with the hippocampus, allowing the model to adopt a higher learning rate. However, as the state or action space increases, the tabular method requires more experience to fully explore the action value of each state, and therefore requires more computing resources to store these values. In this regard, the present application uses the advantages of the self-organizing map SOM to help the system reduce the complexity of the state space, so that reinforcement learning can run more efficiently with limited computing resources. SORL can generalize to similar states, grouping similar states together instead of treating each state as independent, making the reinforcement learning process more efficient by reducing the number of updates required to cover the state space.

[0075] Compared with the prior art, the present invention has the following beneficial effects:

[0076] (1) Using the developmental network DN2 to simulate supervised learning of the cerebellum can better learn complex tasks.

[0077] (2) Combining self-organizing maps and reinforcement learning can achieve rapid learning of intelligent agents.

[0078] (3) Using the temporal difference errors from the developmental network DN2 to train self-organizing reinforcement learning (SORL), the self-organizing reinforcement learning model based on the developmental network DN2 can store states that are difficult for the developmental network DN2 to accurately evaluate. This innovative combination of self-organizing reinforcement learning and the developmental network DN2 highlights how the temporal difference of the cerebellar system signals the hippocampus-basal ganglia system, telling the hippocampus when and which memories to store.

[0079] (4) Simulate the brain's familiarity with state inputs to balance the decision weights between the cerebellum learning system and the basal ganglia-hippocampus learning system. All systems collaborate in evaluating the value of actions, allowing the agent to take advantage of the respective strengths of the cerebellum and basal ganglia-hippocampus learning systems to select actions.

[0080] In summary, it can be seen that the present invention properly utilizes the functions of the cerebellum learning system, the basal ganglia and the hippocampus learning system. This fusion algorithm combines reinforcement learning and supervised learning methods, which not only enhances the robustness of the self-organizing reinforcement learning model based on the developmental network DN2, but also improves the algorithm's ability to obtain effective obstacle avoidance strategies, which can help the intelligent agent quickly learn efficient solutions, thereby improving the intelligent agent's behavioral decision-making ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0082] Figure 1 The overall architecture of the self-organizing reinforcement learning model based on the developmental network DN2 proposed in the embodiment of the present invention is shown, wherein: Figure 1 (a) brain regions and pathways, Figure 1 (b) The structure of the self-organizing reinforcement learning model based on the developmental network DN2;

[0083] Figure 2 The learning process of the developmental network DN2 in the embodiment of the present invention is shown;

[0084] Figure 3 A schematic diagram of a simulation environment in an embodiment of the present invention is shown;

[0085] Figure 4 The navigation paths of the agent in different success periods in the embodiment of the present invention are shown;

[0086] Figure 5 A diagram showing the change in the number of intelligent agent navigation decision steps in an embodiment of the present invention is shown;

[0087] Figure 6 A comparison chart showing the rewards obtained by SORL and SORLDN2 during the navigation decision task in an embodiment of the present invention is shown;

[0088] Figure 7 A comparison chart showing the number of steps achieved by SORL, DN2 and SORLDN2 in the navigation decision task in an embodiment of the present invention is shown;

[0089] Figure 8 A comparison chart showing the number of movement steps obtained by the agent when using six algorithms respectively in five experimental environments in an embodiment of the present invention is shown;

[0090] Fig. 9 A box plot showing the number of motion steps obtained by six algorithms under five environments in an embodiment of the present invention is shown;

[0091] Fig.10 The process architecture of the actual navigation decision of the robot in the embodiment of the present invention is shown;

[0092] Fig.11 A simple experimental environment in an embodiment of the present invention is shown;

[0093] Fig.12 The process of completing the navigation decision-making task of the robot in the embodiment of the present invention is shown;

[0094] Fig.13 The dynamic environment-1 in the embodiment of the present invention is shown;

[0095] Fig.14 The performance of the robot avoiding dynamic obstacles in the first dynamic experiment in the embodiment of the present invention is shown;

[0096] Fig.15 The figure shows the performance of the robot avoiding dynamic obstacles in the second dynamic experiment of the embodiment of the present invention;

[0097] Fig.16 The figure shows the path diagram of the robot in a simple environment according to an embodiment of the present invention;

[0098] Fig.17 The experimental environment 2 in the embodiment of the present invention is shown;

[0099] Fig.18 A navigation path diagram of a robot in a complex environment in an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0100] The following will be combined with the attached Figures 1 to 18 The technical solution of the present invention is described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0101] This application proposes a robot behavior decision-making method that simulates the learning and memory mechanism of the brain. The robot behavior decision-making method combines the developmental network DN2, reinforcement learning, and the self-organizing map SOM to simulate the functions of the cerebellum learning system, the basal ganglia, and the hippocampus learning system, and obtains a self-organizing reinforcement learning model based on the developmental network DN2 (abbreviated as "SORLDN2").

[0102] In most methods of simulating hippocampal function, the hippocampal learning system cannot be constructed alone, and often needs to be combined with another method to realize its function. The self-organizing reinforcement learning model based on the developmental network DN2 in this application first uses the developmental network DN2 as the cerebellum learning system to realize the supervised learning function of the cerebellum, and then combines the self-organizing map and reinforcement learning, and uses the SORL (Self-Organized Reinforcement Learning) model to describe the basal ganglia-hippocampus learning system, denoted as SORL, and then uses the temporal differential error from the developmental network DN2 to train SORL, so that the self-organizing reinforcement learning model based on the developmental network DN2 can store the state that the developmental network DN2 is difficult to accurately evaluate, and then balance the decision weights between the cerebellum learning system and the basal ganglia-hippocampus learning system by simulating the brain's familiarity with the state input, and then collaborate in evaluating the value of actions, so that the intelligent agent can use the respective advantages of the cerebellum and the basal ganglia-hippocampus learning system to select actions.

[0103] Figure 1The overall scheme of the self-organizing reinforcement learning model based on developmental network DN2 (SORLDN2) is shown in FIG. Figure 1 (a) brain regions and pathways, Figure 1 (b) is the structure of the self-organizing reinforcement learning model based on the developmental network DN2. For the cerebellum, the DN2 algorithm is used to implement the characteristics of supervised learning, which can avoid the decision-making behavior of the SORLDN2 model leading to small rewards during the agent's exploration process. The basal ganglia are responsible for implementing the reinforcement learning function. The hippocampus memorizes and encodes the knowledge learned by the agent. SORL simulates the mutual coordination function of the basal ganglia and the hippocampus. SORL not only quickly learns knowledge from DN2 ( Figure 1 (b) but also learns from other experiences where DN2 shows limitations ( Figure 1 (b) Memory B). In addition, the sensory cortex is responsible for representing input states, the thalamus is responsible for action selection, and the motor cortex is responsible for action output.

[0104] Specifically, Zhu et al. used DN1, which mimics the cerebellum, to achieve good results in navigation decision-making tasks. Then Zheng and Wu proposed an upgraded developmental network DN2. The developmental network DN2 has three regions, including a perceptual input layer X from the external environment, a motor layer Z that acts on the external environment, and a hidden layer Y that organizes the interaction between the perceptual input layer X and the motor layer Z;

[0105] The input x(t) from the sensor input layer X, the z(t) from the motion layer Z, and the response y(t) from the hidden layer Y are represented as vectors, respectively. The adaptive part N = (V, G), where V is a matrix of synaptic weights, G is a vector of neuron activation ages, and N y and N z They represent the adaptive parts of the Y neurons in the hidden layer Y and the Z neurons in the motion layer Z respectively.

[0106] Figure 2 The learning process of DN2 is shown in Figure 1. The input layer region X is the sensory input of the agent. In the hidden layer region Y, there are two types of internal neurons, equipped with different connection mechanisms, which adaptively learn features from different regions. Activated neurons are marked in red or green, while inactive neurons are represented in gray. Blue connections represent top-down inputs from the motor layer region Z, orange connections represent lateral interactions within the hidden layer region Y, and green connections correspond to bottom-up inputs from the input layer region X.

[0107] In this embodiment, the self-organizing reinforcement learning model based on the developmental network DN2 uses two types of Y neurons of the developmental network DN2 to assist the robot in completing navigation decision-making tasks, wherein the two types of Y neurons include lower-level Y neurons and higher-level Y neurons, the lower-level Y neurons are responsible for identifying different local features in the perception input layer X, and the higher-level Y neurons are connected to the lower-level Y neurons instead of being directly connected to the perception input layer X.

[0108] The following Algorithm 1 gives the pseudo code of the learning process of the developmental network DN2:

[0109]

[0110] More specifically, the basal ganglia-hippocampus learning system uses the temporal difference error generated by the cerebellum learning system to update its state representation to achieve rapid learning of the intelligent agent. The temporal difference error is a real-time regulating factor of the learning rate and the standard deviation of the neighborhood function in the basal ganglia-hippocampus learning system.

[0111] Reinforcement learning belongs to the category of semi-supervised learning techniques. When an agent performs an action, it will receive a corresponding reward or punishment. The agent obtains the optimal solution to the task by maximizing the immediate reward of each state while evaluating behaviors that may bring greater rewards in the future. The basal ganglia-hippocampus learning system allows individual memories to be stored in a unique, pattern-separated manner, and promotes rapid updates of the system through a simple error-driven learning mechanism.

[0112] The above method combines the developmental network DN2 (hereinafter referred to as "DN2"), reinforcement learning, and self-organizing map SOM (hereinafter referred to as "SOM") to simulate the functions of the cerebellum, basal ganglia, and hippocampus learning systems. In order to clearly understand this method, Algorithm 2 describes the workflow of self-organizing reinforcement learning and DN2.

[0113]

[0114]

[0115] From the above description, it can be seen that the proposed algorithm has several obvious characteristics: First, the calculation of the Q value is affected by DN2 and SORL, and their relative contributions are determined by the parameter η. This ensures that when the current state is consistent with the state stored in SORL, the predicted Q value of SORL contributes more. This process is similar to retrieving episodic memory and using the correlation between memory and state to make decisions. In addition, since SORL uses the temporal difference error generated by DN2 to update its memory storage, it can store states that DN2 can accurately evaluate. In theory, this design allows SORL to store both experience in DN2 and knowledge outside DN2. In addition, SORL uses a higher learning rate, potentially improving the utilization of data.

[0116] Specifically, the basal ganglia are responsible for implementing the reinforcement learning function, and the hippocampus is responsible for memorizing the knowledge learned by the encoding agent. The self-organizing reinforcement learning model based on the developmental network DN2 also includes: environment, sensory cortex, thalamus and motor cortex, the environment is used to provide state and reward information, the sensory cortex is responsible for representing the input state, the thalamus is responsible for action selection, and the motor cortex is responsible for action output.

[0117] The working principle of the present invention is that the present application proposes a self-organizing reinforcement learning model SORLDN2 based on the developmental network DN2, uses the developmental network DN2 as the cerebellum learning system, and describes the basal ganglia-hippocampus learning system through self-organizing reinforcement learning. In particular, the cerebellum learning system provides exploratory guidance for the basal ganglia-hippocampus learning system, alleviating the problem of slow convergence of reinforcement learning. The basal ganglia-hippocampus learning system uses the temporal difference error generated by the cerebellum learning system to update its state representation. Specifically, the temporal difference error is a real-time adjustment factor for the learning rate and the standard deviation of the neighborhood function in the basal ganglia-hippocampus learning system. This dynamic adjustment process enables the basal ganglia-hippocampus learning system to give priority to and retain the state memory that the cerebellum learning system predicts is not very accurate, and then use these memories to enhance overall decision-making ability and learning.

[0118] In addition, for the hippocampus, since experience is stored in a separate mode, the tabular method of representing the action-value function is more consistent with the hippocampus, allowing the model to adopt a higher learning rate. However, as the state or action space increases, the tabular method requires more experience to fully explore the action value of each state, and therefore requires more computing resources to store these values. In this regard, the present application uses the advantages of the self-organizing map SOM to help the system reduce the complexity of the state space, so that reinforcement learning can run more efficiently with limited computing resources. SORL can generalize to similar states, grouping similar states together instead of treating each state as independent, making the reinforcement learning process more efficient by reducing the number of updates required to cover the state space.

[0119] In order to verify the effectiveness of the self-organizing reinforcement learning model based on the developmental network DN2 in the robot behavior decision-making problem, simulation experiments and physical experiments were carried out. The parameters set in the experiment are: discount factor γ = 0.99, neuron neighborhood range parameter σ c =0.1, σ=0.1, α=0.9, λ=0.01. The input layer X is set to X r = {cosθ g ,sinθ g ,cosθ o ,sinθ o ,d g / (d g +d o ),d o / (d g +d o )}, where θ g , d g Represent the angle and distance between the robot and the target, θ o , d o Represents the angle and distance between the robot and the obstacle respectively. The output layer consists of 36 action numbers a∈(1,36), and the robot motion is defined as state next =state+step*(cos(a*π / 18),sin(a*π / 18)), where step is the robot step length. The robot needs to move from the starting point to the target point. When it reaches the target, it will get a reward of 100. If the robot collides with an obstacle or a wall, it will get a reward of -50.

[0120] The specific reward r is set as follows:

[0121]

[0122] in, and Respectively represent the distance between the robot and the target and obstacle at time t. arrive , d safe and d arrive is the set distance constant. Less than d arrive The mission is successful when Less than d safe When the task fails, d warning is the distance at which a warning of an impending collision is issued.

[0123] The reward for each movement of the robot is in the form of: r = r1 + r2 + r3 + r4.

[0124] 1 Simulation Experiment

[0125] Figure 3 The following is a schematic diagram of the simulation environment. The blue dot represents the starting position of the agent, the green pentagon represents the target position, and the black square represents the obstacle. The environment size is set to 500×500 units, the robot movement step is set to 10 units, the coordinates of the starting point and the target point are (20,50) and (470,450) respectively, and the parameters for evaluating the reward in the simulation are set to: d arrive =10,d safe =10,d warning =25.

[0126] Figure 4 The navigation paths of the agent at different success periods are shown in Figure 1. It can be seen intuitively that as the number of interactions between the agent and the environment increases, the robot path length gradually decreases. In particular, Figure 4 (a) to (f) indicate that the agent reaches the goal with 100, 87, 85, 81, 80, and 74 decision steps when completing the 1st, 10th, 20th, 30th, 40th, and 50th tasks, respectively.

[0127] In order to verify the integrity of the robot's learning process, Figure 5 The chart of the changes in the number of navigation steps of 400 experiments is given. In the figure, the x-axis represents the number of experiments, and 400 experiments were conducted, and the y-axis represents the number of movement steps of the agent. The results are represented by green circles to indicate that the agent successfully reaches the target point, and red bifurcation points indicate failure. The number of navigation decision steps of the agent finally remains at 74 steps. Figure 5 In the above example, fewer steps mean better decision-making ability of the agent. As mentioned before, the number of steps is gradually reduced. This is mainly due to the contribution of the basal ganglia, which uses reinforcement learning to continuously improve learning ability through interaction with the environment.

[0128] 1.1 Ablation Experiment

[0129] In order to verify the synergistic contribution of SORL and DN2, an ablation experiment was conducted. The parameters of DN2 and SORL are the same as SORLDN2. The results of the three algorithms are shown in Figure 6 , Figure 7 shown.

[0130] Figure 6 The rewards obtained by SORL and SORLDN2 during the navigation decision task. SORL and SORLDN2 were run 30 times, and each curve was smoothed with the average of 30 data. The solid line represents the average value, and the shaded area represents the standard deviation of the reward. The green curve is the SORLDN2 model proposed in this example.

[0131] Figure 7The number of steps achieved by SORL, DN2 and SORLDN2 in the navigation decision task. The shaded bar graph represents task failure, and the solid bar graph represents task success. The green bar graph represents the SORLDN2 model proposed in this embodiment.

[0132] exist Figure 6 In the figure, the reward curve of DN2 is not plotted, because the learning of DN2 is independent of the reward function, so only SORL and SORLDN2 involve reward comparison. Importantly, in order to reflect the effectiveness of combining SORL and DN2, experimental results in five environments are given.

[0133] from Figure 6 The experimental results show that the reward of SORLDN2 is significantly better than that of SORL. This is mainly due to the guidance of DN2. DN2 has prior knowledge of navigation task decisions, which can provide preliminary decision guidance and help reinforcement learning reduce unnecessary exploration paths. However, this knowledge is not perfect enough, so using SORL's exploration can help SORLDN2 further improve its performance. Figure 7 From the experimental results, we can see that SORLDN2 has the smallest number of decision steps in each environment, which shows the superiority of SORLDN2. Combining reinforcement learning and supervised learning methods not only enhances the robustness of the SORLDN2 model, but also improves the ability of the algorithm to obtain effective obstacle avoidance strategies. The SORLDN2 method optimizes the ability of agents with DN2 to accumulate experience and enables the behavior strategies learned by the agents to contribute to high-quality decisions in future tasks. Therefore, it can be concluded that the cooperative learning of SORL and DN2 can improve the performance of agents in navigation tasks.

[0134] 1.2 Comparative Experiment

[0135] In order to verify the advantages of the proposed method, the experimental performance of the proposed method (SORLDN2), RLDN1, DN2, DN1, Deep Q-Network (DQN) and Double DQN (DDQN) are compared in five environments. The parameter settings of the six models are shown in Table 1.

[0136] Table 1 Parameter settings of each algorithm in navigation decision task

[0137]

[0138] For the sake of fairness, DN1 used in RLDN1 was transplanted to the algorithm of this embodiment for comparison. DN2 in the work of Wu et al. uses two Y-layer neuron types (111, 100) to complete visual navigation decisions. Unlike the experimental environment of this embodiment, the second intermediate layer of this embodiment does not require state input information in the environment. Therefore, two Y-layer neuron types (011, 100) are used. In addition, SORLDN2 is not suitable for comparison with DQN and DDQN because the developmental network used contains a small number of training sets and learns very quickly, while DQN and DDQN require a lot of exploration to train a good model. SORLDN2 can perform well with a small number of interactions with the environment. However, DQN and DDQN require a large number of interactions to achieve excellent performance. For this issue, in the comparative experiment, only the test results are displayed and analyzed. The test results of these algorithms are as follows. Figure 8 , Fig. 9 shown.

[0139] Figure 8 The figure shows the number of movement steps obtained by the agent using six algorithms in five experimental environments in the form of a line graph. The x-axis represents the environment, and the y-axis represents the number of movement steps of the algorithm in each environment. The results of the six algorithms are displayed by using different color markers. For successful experiments, the number of steps is plotted; for failed experiments, the number of steps is not plotted.

[0140] exist Figure 8 In the figure, DN1 failed to complete the task in the second experimental environment and always performed more decision-making actions in other environments, indicating that its learning ability was poor. DN2 performed slightly better than DN1 because it added a new y-layer neuron type to perform more refined learning. However, compared with SORLDN1 and SORLDN2, DN1 and DN2 performed poorly. The reason is the lack of a large number of perfect training sets. The excellent performance of both depends on a large amount of professional experience for training, however, a large amount of professional experience is difficult to obtain. Due to the contribution of SORL, SORLDN2 and SORLDN1 both improved the performance of the agents based on DN2 and DN1.

[0141] Fig. 9 Box plots of the number of motion steps obtained by six algorithms in five environments. The figure shows the overall trend in the form of a box plot. The x-axis represents the algorithm and the y-axis represents the number of steps in successful experiments. Each box contains all the steps of an algorithm in successful experiments. It can be seen that DN2 always performs better than DN1 in the five environments. SORLDN1 and SORLDN2 outperform DN1 and DN2 respectively. DDQN shows better performance compared to DQN. Overall, SORLDN2 performs the best.

[0142] 2. Physical Experiment

[0143] In order to verify the actual application effect of SOMRLDN2, a real environment experiment was conducted using the Robot Operation Systems (ROS). The experimental environment size is about 6.83m×3.84m. AutolaborPro.1 (AP1) is used as the experimental platform.

[0144] Fig.10 The process architecture for the actual navigation decision of the robot. AP1 obtains its own position and obstacle information by subscribing to the " / odom" and " / scan" topics, realizes the conversion of the coordinate system through " / tf", and executes the speed command by subscribing to the " / cmd vel" topic. After obtaining the input state, AP1 implements action selection through the proposed SORLDN2 algorithm. During the action execution process, AP1 is set to start moving when it faces the target heading angle, that is, the action selection direction. The linear velocity of AP1 is set to 0.2m / s, and the angular velocity is set to 0.2rad / s.

[0145] 2.1 Simple environmental experiment

[0146] Fig.11 This is a simple experimental environment. The environment shown contains three regular obstacles, the robot's starting position is set to (0,0), and the target position is set to (4.0,0.8).

[0147] Fig.12 This figure shows the process of the robot completing the navigation decision-making task in the experiment, and demonstrates the navigation trajectory of the robot in four stages in the actual environment. Fig.12 (a) is the first stage, Fig.12 (b) is the second stage, Fig.12 (c) is the third stage, Fig.12 (d) Stage 4. From the navigation paths of stages 2 and 3, we can see that robot AP1 avoids obstacles very well. From the results of stage 4, we can see that AP1 successfully reaches the target area and executes a very short path. The results show that AP1 using the SORLDN2 algorithm has excellent decision-making ability in the navigation decision task.

[0148] 2.2 Dynamic environment experiment

[0149] In order to verify the adaptability of the SORLDN2 algorithm, two experiments were conducted in a dynamic environment. In the first experiment, people were used as dynamic obstacles. Fig.13 This is the dynamic environment-1, which sets three obstacle areas within the range of AP1's intended walking path. When AP1 moves to the target point, a person actively blocks the robot's path. Fig.14This is the effect picture of the robot avoiding dynamic obstacles in the experiment. Fig.14 (a) is the first stage, Fig.14 (b) is the second stage, Fig.14 (c) is the third stage, Fig.14 (d) Stage 4.

[0150] Fig.14 The four pictures in the first row show the navigation process in the real world, and the four pictures in the second row show the corresponding trajectory through the visualization tool Rviz in ROS. These four groups of pictures show the navigation decision process in four stages. In order to facilitate the observation of the overall obstacle avoidance of AP1, after AP1 completes the task, the dynamic obstacle returns to the original position area, and the obstacle avoidance effect is observed by combining the overall path of the robot and the radar scanning information. Fig.14 As shown in the four pictures in the third row, only four sets of experimental process diagrams are shown here.

[0151] In the second dynamic experiment, a "puppet" AP1 controlled by a handle and a person were used as two dynamic obstacles in the experiment. The "puppet" robot was responsible for blocking the movement of AP1 in the middle road, and the person was responsible for blocking the road around the target point.

[0152] Fig.15 Visualization of the test results, showing the robot's performance in avoiding dynamic obstacles in the second dynamic experiment. Fig.15 (a) is the first stage, Fig.15 (b) Phase II, Fig.15 (c) Phase III, Fig.15 (d) Stage 4. It can be seen that AP1 with SORLDN2 can autonomously navigate to the target without any collision.

[0153] 2.3 Comparative Experiment

[0154] In the comparative experiment, tests were conducted in two unknown environments, thirty experiments were carried out respectively, and four indicators were collected to compare the performance of the six algorithms.

[0155] Fig.16 is the path diagram of the robot in a simple environment, Fig.16 (a) is the DN1 algorithm, Fig.16 (b) is the DN2 algorithm, Fig.16 (c) is the DQN algorithm, Fig.16 (d) is the DDQN algorithm, Fig.16 (e) is the SORLDN1 algorithm, Fig.16 (f) is the SORLDN2 algorithm.

[0156] The numerical results are given in Table 2. The indicators include the number of successful goals, the minimum number of steps in a successful task, the maximum number of steps in a successful task, the average number of steps in a successful task, and the variance of the number of steps in a successful task. It can be seen that DN2, SORLDN1, and SORLDN2 show good results, with an average number of steps of about 25-26. However, DN1, DQN, and DDQN remain at about 27-28 steps. It can be found that these algorithms achieved a 100% success rate and showed good capabilities in simple environments. Compared with DN1, DQN, and DDQN, DN2, SORLDN1, and SORLDN2 have greater advantages in decision-making ability.

[0157] Table 2 Results of 30 experiments on 6 algorithms in a simple environment

[0158]

[0159] Previous experiments have verified the effectiveness and superiority of SORLDN2 in a simple environment. Fig.17 A comparative experiment is carried out in the experimental environment 2 shown in the figure. In the experimental environment 2, the starting position of AP1 is set to (0, 0) and the target position is set to (4.0, 2.4).

[0160] Fig.18 It is the navigation path map of the robot in a complex environment. Fig.18 (a) is the DN1 algorithm. Fig.18 (b) DN2 algorithm, Fig.18 (c) is the DQN algorithm. Fig.18 (d) is the DDQN algorithm, Fig.18 (e) is the SORLDN1 algorithm, Fig.18 (f) is the SORLDN2 algorithm. The specific experimental results are shown in Table 3.

[0161] Table 3 Results of 30 experiments on 6 algorithms in complex environments

[0162]

[0163] From the experimental results, the performance of these algorithms in complex environments has declined compared to simple environments. Due to the increase in the number of obstacles, the success rate of DN1, DQN and DDQN has decreased, and the average number of steps in successful tasks has remained between 30-32. The average number of steps of DN2 and SORLDN1 has remained between 29-30. The average number of steps of SORLDN2 is the smallest, which is maintained between 28-29, and the variance of the number of steps of SORLDN2 is also relatively small. Therefore, it is proved again that SORLDN2 has better decision-making ability.

[0164] It should be noted that the present application can be implemented in software and / or a combination of software and hardware, for example, can be implemented using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present application (including relevant data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to perform each step or function.

[0165] In addition, a part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. The program instruction for calling the method of the present application may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium, and / or stored in a working memory of a computer device that runs according to the program instruction. Here, according to an embodiment of the present application, a device is included, the device including a memory for storing computer program instructions and a processor for executing program instructions, wherein, when the computer program instruction is executed by the processor, the device is triggered to run the method and / or technical solution based on the aforementioned multiple embodiments according to the present application.

[0166] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic features of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present application is defined by the attached claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by one unit or device through software or hardware.

Claims

1. A robot behavior decision-making method that simulates the learning and memory mechanism of the brain, characterized by: The robot behavior decision-making method combines the developmental network DN2, reinforcement learning and the self-organizing map SOM to simulate the functions of the cerebellum learning system, the basal ganglia and the hippocampus learning system, and obtains a self-organizing reinforcement learning model based on the developmental network DN2; The self-organizing reinforcement learning model based on the developmental network DN2 first uses the developmental network DN2 as the cerebellum learning system to realize the supervised learning function of the cerebellum, and then combines self-organizing maps and reinforcement learning, and uses the SORL model to describe the basal ganglia-hippocampus learning system, denoted as SORL, and then uses the temporal difference error from the developmental network DN2 to train SORL, so that the self-organizing reinforcement learning model based on the developmental network DN2 can store states that are difficult for the developmental network DN2 to accurately evaluate, and then balance the decision weights between the cerebellum learning system and the basal ganglia-hippocampus learning system by simulating the brain's familiarity with the state input, and then collaborate in evaluating the value of actions, so that the intelligent agent can use the respective advantages of the cerebellum and basal ganglia-hippocampus learning systems to choose actions.

2. The robot behavior decision-making method according to claim 1, characterized in that: The developmental network DN2 has three regions, including a perception input layer X from the external environment, a motor layer Z that acts on the external environment, and a hidden layer Y that organizes the interaction between the perception input layer X and the motor layer Z; The input x(t) from the sensor input layer X, the z(t) from the motion layer Z, and the response y(t) from the hidden layer Y are represented as vectors, respectively. The adaptive part N = (V, G), where V is a matrix of synaptic weights, G is a vector of neuron activation ages, and N y and N z They represent the adaptive parts of the Y neurons in the hidden layer Y and the Z neurons in the motion layer Z respectively.

3. The robot behavior decision-making method according to claim 2, characterized in that: The self-organizing reinforcement learning model based on the developmental network DN2 utilizes two types of Y neurons of the developmental network DN2 to assist the robot in completing navigation decision-making tasks. The two types of Y neurons include lower-level Y neurons and higher-level Y neurons. The lower-level Y neurons are responsible for identifying different local features in the perception input layer X, and the higher-level Y neurons are connected to the lower-level Y neurons instead of being directly connected to the perception input layer X.

4. The robot behavior decision-making method according to claim 3, characterized in that: The basal ganglia-hippocampus learning system uses the temporal difference error generated by the cerebellum learning system to update its state representation to achieve rapid learning of the intelligent agent. The temporal difference error is a real-time regulating factor of the learning rate and the standard deviation of the neighborhood function in the basal ganglia-hippocampus learning system.

5. The robot behavior decision-making method according to claim 4, characterized in that: The basal ganglia-hippocampus learning system allows individual memories to be stored in a unique, pattern-separated manner and promotes rapid updating of the system through a simple error-driven learning mechanism.

6. The robot behavior decision-making method according to claim 1, characterized in that: The basal ganglia are responsible for implementing the reinforcement learning function, and the hippocampus is responsible for memorizing and encoding the knowledge learned by the intelligent agent.

7. The robot behavior decision-making method according to claim 1, characterized in that: The self-organizing reinforcement learning model based on the developmental network DN2 also includes: The environment, which provides status and reward information; The sensory cortex, which is responsible for representing input states; the thalamus, responsible for action selection; The motor cortex is responsible for movement output.

8. A robot behavior decision-making device that simulates the learning and memory mechanism of the brain, characterized in that: The robot behavior decision-making device includes: one or more processors; A computer readable medium for storing one or more computer readable instructions, When the one or more computer-readable instructions are executed by the one or more processors, the one or more processors implement the robot behavior decision-making method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intrinsic motivation based self-cognition system for motion balance robot and control method

    CN104992059A

  • MRI (Magnetic Resonance Imaging) hippocampus segmentation method and system based on hypergraph numerical neural membrane system

    CN114359555A

  • Incremental self-organization-based cerebellum learning model and application method thereof

    CN117422124A

  • Server for automatic generation of esg regulatory response reports using generative artificial intelligence and automatic generation method using the same

    KR1020250128822A