A brain-like perception-learning-decision-making system and method
By designing a brain-like perception-learning-decision system, and using the collaborative work of multiple modules, the existing simulated brain system's shortcomings in scalability and task diversity are solved, independent learning and decision-making of environmental information are achieved, and coordinated scheduling is achieved between multiple independent brain regions.
Patent Information
- Application Number
- CN202310757825.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-06-26
AI Technical Summary
The existing simulated brain systems have shortcomings in scalability and task diversity, making it difficult to achieve coordinated scheduling between multiple independent brain regions.
A brain-like perception-learning-decision system is designed, including perceptual cortex module, gated module, reward and decision-making module, working memory module, cognitive map construction module and user interaction module. Through the collaborative work of these modules, independent learning and decision-making of environmental information can be realized, and coordinated scheduling is achieved between multiple independent brain regions.
It realizes independent learning decision-making behavior of environmental information received by the human brain, enhances the scalability and multi-task processing capabilities of the system, and can exist multiple brain regions with independent functions in parallel within a system, realizing coordinated scheduling between multiple independent brain regions.
Smart Images

Figure CN116795942B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human brain simulation, and in particular to a brain-like perception-learning-decision-making system and method. Background Art
[0002] Spiking neural network (SNN) is a new generation of neural network with higher biological credibility designed to combine the advantages of neuroscience and machine learning. It is fundamentally different from the currently popular neural network and machine learning methods in the transmission and processing of internal signals, that is, it uses discrete pulses instead of floating point numbers for learning. Pulse signal is a discrete event, generally represented by 0 and 1, similar to the action potential in biological neurons. Generally speaking, the input sequence and output sequence of SNN are also pulse sequences. Compared with other artificial neural networks, SNN has great prospects in computing performance, processing complex time series data, low energy consumption and biological credibility, and has therefore attracted widespread attention. SNN has therefore become a new generation of artificial intelligence that is expected to achieve general artificial intelligence and a major method that can help neuroscientists build brain models and explore brain mechanisms.
[0003] Currently, researchers have two main methods for building multi-brain region models to understand and simulate human brain behavior and decision-making functions: bottom-up and top-down. Top-down methods usually focus on detailed modeling of various attributes of the brain. Bottom-up methods focus on reproducing the specific behaviors of biological objects observed in experiments, but ignore the diversity of brain decisions and behaviors.
[0004] Some work combines both top-down and bottom-up approaches to model the functions of multiple brain regions, and achieves the coordinated operation of multiple brain regions through complex simulated neuronal dynamics and predefined neural circuit rules. Although this type of work roughly constructs a simulated brain system, the constructed neuronal connections are too complex, and as a simulated brain system, the tasks it can perform and its scalability are limited.
[0005] The existing simulated brain systems have the following main problems: 1) poor scalability, and can only complete specific tasks, or the expansion to new tasks is very difficult or complicated; 2) it cannot balance the coordination of multiple brain regions or the realization of brain region functions, or emphasizes the modeling of the real organizational structure in the brain but lacks functional realization, or emphasizes certain specific functions of the brain but fails to form coordinated scheduling among many independent brain regions. Summary of the invention
[0006] The purpose of the present invention is to provide a brain-like perception-learning-decision-making system and method, which can autonomously perform learning and decision-making behaviors on the environmental information received by the human brain and realize coordinated scheduling among multiple independent brain regions.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A brain-like perception-learning-decision system, the system comprising: a perception cortex module, a gating module, a reward and decision module, a working memory module, a cognitive map construction module and a user interaction module;
[0009] The perceptual cortex module is used to input at least one perceptual information, integrate all the input perceptual information into a perceptual pulse sequence by using an attention mechanism or a Bayesian decision, and transmit the perceptual pulse sequence to the gating module;
[0010] The gating module is used to splice the latest control signal with the perception pulse sequence, and automatically encode and output the control signal at the next moment according to the spliced signal, further control whether the perception pulse sequence is output according to the control signal at the next moment, and transmit the control signal at the next moment to the cognitive map construction module and / or the reward and decision module;
[0011] The reward and decision module is used to input the perception pulse sequence after receiving the control signal at the next moment, and obtain a corresponding reward signal according to the perception pulse sequence, and select a corresponding operation mode according to the control signal, and then generate a decision signal according to the reward signal using the corresponding operation mode;
[0012] The cognitive map construction module is used to input the visual information and odometer information in the perception cortex module after receiving the control signal at the next moment, and continuously construct the cognitive map according to the visual information and odometer information;
[0013] The user interaction module is used to display the specific content represented by the decision signal and / or to display the constructed cognitive map in an image manner.
[0014] Optionally, the perceptual cortex module includes: a plurality of neural networks and an attention-based modality fusion layer;
[0015] Each neural network is used to extract features and identify targets from the sensory information input by a channel, and obtain the recognition result of a channel;
[0016] The attention-based modality fusion layer is used to perform modality fusion on the recognition results output by all neural networks to synthesize a perception pulse sequence.
[0017] Optionally, when the perception information is visual information, the neural network is a convolutional neural network based on LIF neurons; the visual information includes dynamic visual information and static visual information;
[0018] When the sensory information is auditory information, the neural network is a recurrent spiking neural network based on CLIF neurons.
[0019] Optionally, the gating module includes a pulse neural network trained according to control signal issuance rules.
[0020] Optionally, the system further comprises: a working memory module and a motorized output module;
[0021] The working memory module is connected to the reward and decision module; the working memory module is used to receive the reward signal and the corresponding decision signal of each brain region, and store the decision record after forming it;
[0022] The motorized output module is connected to the gate control module; when the gate control module transmits the control signal of the next moment to the reward and decision module, the gate control module simultaneously transmits the control signal of the next moment to the motorized output module;
[0023] The motorized output module is also connected to the user interaction module and the reward and decision module respectively; the motorized output module is used to output the decision signal generated by the reward and decision module to the user interaction module.
[0024] Optionally, the reward and decision module includes: a reward submodule and a decision submodule;
[0025] The reward submodule is connected to the decision submodule, and the reward submodule is used to obtain a reward signal corresponding to the perception signal in the perception pulse sequence, and transmit the reward signal to the decision submodule;
[0026] The decision submodule is connected to the gating module and the working memory module respectively. The decision submodule is used to select a corresponding operation mode according to a control signal. When the selected operation mode is an offline decision mode, a preset decision signal is directly outputted. When the selected operation mode is an online decision mode, after receiving a reward signal, a plurality of decision records are read from the working memory module, and a decision model is updated according to the plurality of decision records, and then a decision signal is generated by using the updated decision model according to the reward signal.
[0027] Optionally, the cognitive map construction module includes: a grid cell submodule, a place cell submodule, a visual cell submodule and an experience map construction submodule;
[0028] The grid cells in the grid cell submodule are modeled using a continuous attractor network. The grid cells are driven and activated by odometer information. The neural activity of the grid cells changes as the agent moves, and the connection with the place cell submodule drives the activity of the place cells to represent the current location information of the agent.
[0029] The visual cells in the visual cell submodule are driven and activated by visual information, representing the visual feature information of the current environment of the agent, and are transmitted to the experience map construction submodule together with the position information;
[0030] The experience map construction submodule is used to update the experience map based on the visual feature information and the position information, and to correct the accumulated errors by combining the visual information with the updated experience map, so as to continuously generate the cognitive map and output it to the user interaction module.
[0031] Optionally, the experience map is updated according to the visual feature information and the position information, and the accumulated errors are corrected by combining the visual information with the updated experience map, so as to continuously generate the cognitive map, specifically including:
[0032] The visual feature information and the position information are used as trajectory information and compared with the experience map stored in the experience map construction submodule;
[0033] If the trajectory information overlaps with the experience map, the grid cell and place cell firing at the current point are reset to the grid cell and place cell firing of the matched experience;
[0034] Iteratively update the global pose of each experience point along the experience map based on the coincidence information;
[0035] In the process of updating the global pose of the experience point in each iteration, all experience points and connections in the experience map are traversed, and the global pose of the two connected experience points is corrected according to the direction and distance of each connection, so that the global pose difference between the two experience points converges to the direction of the connection between the two experience points, generating a cognitive map.
[0036] Optionally, the user interaction module is used to provide a graphical interaction interface, input signals on the graphical interaction interface, and obtain visual execution results; generate corresponding signals according to a given instruction sequence, and send them to the perceptual cortex module; execute different instructions based on the analysis of the instruction sequence, and display the information and activation status in the perceptual cortex module, gating module, reward and decision module, working memory module, and cognitive map construction module.
[0037] A brain-like perception-learning-decision-making method, comprising:
[0038] Attention mechanism or Bayesian decision is used to integrate at least one acquired sensory information into a sensory pulse sequence;
[0039] splicing the latest control signal with the sensing pulse sequence, and automatically encoding and outputting the control signal at the next moment according to the spliced signal;
[0040] When the control signal at the next moment controls the output of the perception pulse sequence, a corresponding reward signal is obtained according to the perception pulse sequence, and a corresponding operation mode is selected according to the control signal, and then a decision signal is generated according to the reward signal using the corresponding operation mode;
[0041] When the control signal at the next moment inhibits the output of the perception pulse sequence, receiving visual information and odometer information, and continuously constructing a cognitive map according to the visual information and odometer information;
[0042] Comprehensively output the decision signals and / or cognitive maps at all times.
[0043] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0044] The present invention discloses a brain-like perception-learning-decision system and method, wherein the perception cortex module integrates at least one perception information, the gate control module automatically encodes and outputs a control signal according to the integrated perception pulse sequence and the latest control signal, the control signal can be used to select the task path of the reward and decision module and the cognitive map construction module, the reward and decision module can generate a decision signal when selected, the cognitive map construction module can continuously construct a cognitive map when selected, and the user interaction module displays the specific content represented by the decision signal, and / or displays the constructed cognitive map in an image. The automatic encoding of the gate control module realizes the autonomous learning and decision-making behavior of the environmental information received by the human brain, and the gate control module enables multiple brain regions with independent functions to exist in parallel in one system, realizing the coordinated scheduling between multiple independent brain regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0046] Figure 1 A schematic diagram of the structure of a brain-like perception-learning-decision-making system provided by an embodiment of the present invention;
[0047] Figure 2 An example diagram of a perceptual cortex module provided by an embodiment of the present invention;
[0048] Figure 3 An example diagram of the path correction process of the cognitive map construction module provided in an embodiment of the present invention;
[0049] Figure 4 An example diagram of a visualization interface provided by an embodiment of the present invention;
[0050] Figure 5 An example diagram of the module interaction mode when executing cognitive map construction provided by an embodiment of the present invention;
[0051] Figure 6 An example diagram of module interaction when performing perception information recognition provided by an embodiment of the present invention;
[0052] Figure 7 An example diagram of module interaction when performing online reinforcement learning provided by an embodiment of the present invention;
[0053] Figure 8 A flowchart of a brain-like perception-learning-decision-making method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] An embodiment of the present invention provides a brain-like perception-learning-decision system, including: a perception cortex module, a gating module, a reward and decision module, a working memory module, a cognitive map construction module and a user interaction module.
[0057] The perceptual cortex module is used to input at least one perceptual information, integrate all the input perceptual information into a perceptual pulse sequence using an attention mechanism or Bayesian decision, and transmit the perceptual pulse sequence to the gating module.
[0058] The gating module is used to splice the latest control signal with the perception pulse sequence, and automatically encode and output the control signal at the next moment according to the spliced signal, further control whether the perception pulse sequence is output according to the control signal at the next moment, and transmit the control signal at the next moment to the cognitive map construction module and / or the reward and decision module.
[0059] The reward and decision module is used to input the perception pulse sequence after receiving the control signal at the next moment, and obtain the corresponding reward signal according to the perception pulse sequence, and select the corresponding operation mode according to the control signal, and then generate a decision signal according to the reward signal by adopting the corresponding operation mode.
[0060] The cognitive map construction module is used to input the visual information and odometer information in the perception cortex module after receiving the control signal at the next moment, and continuously construct the cognitive map according to the visual information and odometer information.
[0061] The user interaction module is used to display the specific content represented by the decision signal and / or to display the constructed cognitive map in an image manner.
[0062] See also Figure 1 , each module in the system is elaborated in detail.
[0063] (1) Perceptual cortical module
[0064] The input of the perceptual cortex module: quantitative information of one or more channels: such as visual information (pictures) + auditory information (audio), output: perceptual pulse sequence. The perceptual cortex module includes at least one input perceptual channel, and the module is equipped with a pulse neural network that can integrate information from multiple channels. The pulse neural network can first process (feature extraction + target recognition) the input information of each channel and perform modal fusion. The pulse neural network can integrate external multi-channel perceptual information and generate an integrated perceptual pulse signal based on the feature level or decision level fusion of the attention mechanism or Bayesian decision-making, which is used for subsequent modules to use and process.
[0065] The perceptual cortex module perceives external multimodal information and integrates it into a unified representation through a multimodal pulse neural network. The perceived information can be one or more of the multi-channel information. That is, the perceptual cortex module includes: multiple neural networks and an attention-based modal fusion layer. Each neural network is used to perform feature extraction and target recognition on the perceptual information input by a channel to obtain the recognition result of a channel. The attention-based modal fusion layer is used to perform modal fusion on the recognition results output by all neural networks to synthesize a perceptual pulse sequence.
[0066] Figure 2 It is a feasible network structure that can simultaneously integrate one or more of static visual, auditory, and dynamic visual information. Figure 2 The above integration is to first normalize the information of three (or only one or two) channels, and then preliminarily encode the information of the channel through their respective networks (i.e. Figure 2 The first half of the signal is then integrated into a unified signal ( Figure 2the second half of the ). Figure 2 The graphics with bold lines and no text inside (including bold rectangles and stacked parallelograms) represent feature maps during the execution of the neural network, while the graphics with arrows, no bold outlines, and text inside represent the middle layers of the neural network or some vector operations (used to extract features from vectors), where FC stands for fully connected layer, Pooling stands for pooling layer, Conv stands for convolutional layer, and RC stands for recurrent connection layer, which are the middle layers of the neural network. CONCAT and ATTENTION represent concatenation and attention weighted sum operations, respectively.
[0067] exist Figure 2 In the network, the convolutional neural network (CNN) based on LIF neurons is used to process visual information (including dynamic vision and static vision), while the recurrent spiking neural network based on CLIF neurons is used to process auditory information. The dynamic definition of LIF neurons is as follows:
[0068]
[0069]
[0070]
[0071]
[0072] in represents the decay process of membrane potential, g(x) represents the pulse equation, that is, when the membrane potential exceeds the threshold V th When , the neuron will emit a pulse. l(n) represents the number of neurons in the nth layer, w ij represents the synaptic weight from the jth neuron before the synapse to the ith neuron after the synapse. j ∈{0, 1} represents the output of the jth neuron. When the value is 1, it means a pulse is generated, and when the value is 0, it means nothing happens. represents the membrane potential value of the ith neuron in the nth layer at time step t+1, It represents the current value obtained by the ith neuron in the nth layer from the input of the neurons in the previous layer. Represents a constant offset current value (generally negligible in pulse neural networks).
[0073] The dynamics of a C-LIF neuron are defined as follows:
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081] where θ represents the threshold potential. k, superscript n, and subscript i represent the i-th neuron in the n-th layer at time step k, and L(n) represents the number of neurons in the n-th layer. represents the output signal of the neuron, Represents the pulse signal of the neurons in the previous layer The weighted value of . It is an abstract representation of the "refractory period" produced by biological neurons after emitting a pulse. It is an abstraction of the output of a neuron combined with the impulses received from neurons in the upper layer. represents the final membrane potential of the neuron. m , β s is the time constant, V 0 is the normalization parameter.
[0082] The network that encodes different perceptual input information will eventually be integrated into an attention-based modal fusion layer to fuse modal information and obtain the final encoding result, completing the encoding of multi-channel perceptual information.
[0083] (2) Gate control module
[0084] The gate control module includes a pulse neural network trained according to the control signal release rules (different from the multimodal pulse neural network structure in the perception cortex module). The input of the gate control module: the perception pulse sequence output by the perception cortex module and the control signal output by the gate control module at the previous moment. The control signal can be used to control the selection of the data path. The output of the gate control module: includes the pulse sequence in the perception cortex module and the control signal (used for the selection of the task path).
[0085] The gating module is equivalent to the center of the entire brain-like perception-decision-making-simulation system. It is mainly responsible for generating control signals and transmitting the perception information transmitted by the perception cortex module. It is used to control the input pulses of all modules except the perception cortex module, thereby controlling the operation of each module of the entire system and the transmission of data flow between different modules, so that tasks can be carried out in an orderly manner.
[0086] In actual implementation, this module will be equipped with a trained model, continuously receive sensory signals and splice them with the pulse control signals in the system at that time. After network processing, new pulse control signals are generated and transmitted to downstream modules. The network here refers to the "spiking neural network". The processing of pulse signals (input vectors) by this type of neural network can be intuitively regarded as a certain nonlinear transformation of the input signal. The processing result here is the input sensory signal and the pulse control signal in the system, and a set of new pulse control signals are output. Specifically, the spiking neural network is composed of several layers of neurons. The layers of neurons are connected by synapses. The synapses have weights inside. The input signals from the upper layer will be weighted and summed and transferred to the next layer. Each layer of neurons will receive the signals from the synapses and gradually accumulate the membrane potential until the threshold potential is reached, and then pulses will be emitted, realizing a nonlinear process.
[0087] The gate control module continuously receives the perceptual pulse signal in the perceptual pulse sequence transmitted by the perceptual cortex module, and splices it with the pulse control signal issued by itself in the system at the same time as the input of the built-in pulse neural network. The splicing method can be direct splicing of signal vectors or addition of signal vectors. The built-in pulse neural network can automatically encode and output new control signals according to the rules learned in advance after receiving the input, which is used to control the activities of other modules in real time, and process the perceptual signal (control the downward propagation of the perceptual pulse signal), and transmit it to other subsequent modules; the output control signal will also immediately affect the current output of the module (blocking the continued propagation of the pulse sequence). In some situations, some pulse sequences will continue to be transmitted backwards, and some pulse sequences will not be transmitted (the specific rules are set in advance and the pulse neural network will be executed according to this rule). For example, when the signal transmitted by the perceptual cortex module represents the signal "9", it means that the next transmitted signal should represent the type of instruction. At this time, no matter what the next signal represents, it will not be transmitted backwards.
[0088] (3) Reward and decision-making module
[0089] The reward and decision module receives the sensory pulse signal (pulse sequence) and control signal (task path selection) processed by the gate module. This module has a reward submodule and a decision submodule. The decision submodule has a built-in pulse neural network model for reinforcement learning.
[0090] The reward submodule obtains the corresponding reward signal (output) (i.e., a rough matching relationship) based on the perception signal (input). The reward signal can be a positive or negative signal and is passed to the decision submodule to help the decision submodule learn a reasonable decision-making method.
[0091] The decision submodule (completely different from the structure of the spiking neural network in the perceptual cortex module, input: reward signal and control signal, output: decision signal (a vector used to characterize which action needs to be taken)) is used to output the decision signal. The decision submodule has two operating modes depending on the control signal: online decision and offline decision. In the case of offline decision, the decision submodule can directly continue to output the decision signal without updating its own decision strategy based on the decision record and reward signal (that is, the parameters of the network do not need to be updated). In the online decision mode, when receiving the reward signal, the submodule will store the reward signal and the last decision information in the working memory module (only one corresponding decision record is stored at a time). After receiving the reward signal, the decision submodule reads a certain number of decision records from the working memory module, learns the decision strategy that tends to obtain the best reward signal based on the decision conditions, actions, and decision results recorded in the decision record (updates the spiking neural network model), and generates a new decision signal after the learning is completed. For example, consider a simple decision-making task. The relationship between a given number and its reward value is that the numbers "0", "1", and "2" correspond to a reward value of 0, and the number "3" corresponds to a reward value of 3. Then under the same conditions, the decision subsystem will be more inclined to make a decision "3" after multiple decisions.
[0092] (4) Working memory module
[0093] The working memory module is a module abstracted from the long and short-term memory of the hippocampus. It has a certain upper capacity limit and will perform random memory access according to given indicators. It can also solidify short-term memory into long-term memory through consolidation. As an information storage module, the working memory module is used to receive information (input) from other brain regions and store and awaken memory (output). The storage capacity is fixed. When the capacity is full, it will overwrite non-critical information according to importance. When some modules such as the reward and decision module need to awaken memory from it, a certain amount of stored memory will be randomly read and handed over to the reward and decision module.
[0094] (5) Mobile output module
[0095] Mobile module: The input is the decision signal generated by the reward and decision module, and the output is to the user interaction module (in the user interaction module, it is displayed as the specific content represented by the decision signal). At the same time, additional models can also be introduced for processing (depending on the input requirements of the external output device, such as extension, etc. The current default situation is no processing.) and then the decision signal is output to any possible external output device.
[0096] The motor module continuously receives the decision signals generated in the reward and decision module, and outputs the decision signals to possible external interfaces, which may also affect the input of perceptual information.
[0097] (6) Cognitive map construction module
[0098] The input signal of the cognitive map construction module is the control signal from the gate module, as well as the independent visual information and odometer information input from the outside world. The module contains a grid cell submodule, a place cell submodule, a visual cell submodule and an experience map construction submodule. The grid cell is modeled with a simplified CAN (Continuous Attractor Network). The activity of the cell is driven and activated by the odometer information (angle, distance, etc.). Its neural activity will change with the movement of the agent to achieve the path integration of its own movement, and drive the activity of the place cells through the connection with the place cells to represent the current position information of the agent; the visual cells are driven and activated by the visual information, representing the visual feature information of the current environment of the agent, and updating the experience map together with the position information. In addition, the accumulated errors need to be corrected by combining the visual information with the experience map, so as to continuously build the cognitive map and output it to the user interaction module (the specific output content is a series of real points that can describe a cognitive map, and the user interaction module will output them in the form of images);
[0099] The specific process of correcting the accumulated error is as follows: in each iteration of updating the path integral information, all the experience points and their connections in the experience map are traversed, and the global pose of the two connected experience points is corrected according to the direction and distance of each connection, so that the global pose difference between the two experience points converges in the direction of the connection between the two experience points, and the error convergence size depends on the length of the connection. Since the error is introduced at the latest experience point, this process needs to be iterated to ensure that the error can be dispersed to all locations in the map.
[0100] The method for constructing a cognitive map based on image sequences and odometer information in the cognitive map construction module includes the following steps:
[0101] 1) Receive control signals and control whether to run or not according to the control signals.
[0102] 2) Receive external information, obtain visual information from the external visual input interface, receive the input of the odometer, and submit this information to the grid cell submodule, the position cell submodule and the visual cell submodule.
[0103] 3) Simulate neural activity. The grid cell submodule of the cognitive map construction module is modeled using a simplified version of the CAN (Continuous Attractor Network) model. The cell activity is driven and activated by the odometer input in step 2), and its neural activity changes as the agent moves to achieve path integration of its own movement. Since the representation of a single grid cell may be ambiguous, multiple grid cells are used for joint driving, and the activity of the place cells is driven by the connection with the place cells, thereby representing the current location information of the agent.
[0104] 4) Obtaining position information based on the result of step 3) and correcting it with the input visual information and the generated empirical map.
[0105] 5) Correct the path. If the current experience is found to overlap with the past visual experience in step 4), the information of all path integration along the way is iteratively updated according to the overlap content to correct the error generated in the path integration process. The correction process will be set several times according to experience to prevent the correction intensity from being too high or the generated cognitive map from not converging, such as Figure 2 An example of a path correction process.
[0106] The results of step 4) and step 5) are combined to construct a cognitive map, and step 2) is executed repeatedly until the transmitted control signal changes.
[0107] like Figure 3 The example of module interaction when executing cognitive map construction is shown, and the direction of the arrow represents time. The horizontal and vertical axes of each figure represent the position (x, y) of the global map. The agent starts from point (0, 0), the solid line represents the trajectory that has been moved, and the dot represents the current position of the agent. The first three figures show that as time goes by, the path integral error continues to accumulate, resulting in the global coordinates in the experience map not returning to (0, 0) when the agent has actually returned to the current position; the fourth figure shows that the agent matches the visual information in the previous experience through visual information and corrects the current position; the fifth figure shows that the path integral information of the trajectory is iteratively updated, and the position of all experience points in the experience map is corrected; Figures 6-7 show that during the movement of the agent, since it can successfully match the previous experience every time, the trajectory almost overlaps with the previous movement trajectory.
[0108] (7) User Interaction Module
[0109] User interaction module: used to help users regulate the perception data flow (i.e., it can control the external signal input, number of channels, etc. to the perception cortex module. The input is the instructions (symbols) given by the user, and various channel information corresponding to the symbols are output to the perception cortex module (i.e., pictures, sounds, etc.)), stimulate the activation of the perception cortex module and then control the system, and show the user the output of each module in the system and its various state parameters at a macro level (i.e., a graphical display of the impact of control signals on them).
[0110] The user interaction module includes the following functions: 1) Providing a graphical interactive interface, the user controls the input signal and obtains a visual execution result; 2) Generates corresponding signals according to the instruction sequence given by the user, and the perception cortex module and the gating module start to execute the instructions according to the given signal, and coordinately sends control pulses to stimulate each other module; 3) Executes different instructions based on the analysis of the instruction sequence, and displays the relevant information and activation status of each module and its sub-modules.
[0111] The visual interface of the brain-like perception-learning-decision-making system of the embodiment of the present invention is as follows: Figure 4 shown. Figure 4 In the figure, Orbitofrontal represents frontal cortex, Prefrontal represents prefrontal lobe, MotorCortex represents motor cortex, Entorhinal represents entorhinal cortex, DA represents reward, Hippocampus represents hippocampus, and Output represents output.
[0112] Exemplarily, the four tasks that can be achieved by the present invention are:
[0113] Separate instructions (90): Cognitive map construction - receiving the image sequence and odometer information from the external input interface, stimulating the simulated grid cells and position cells in the cognitive map construction module, constructing the cognitive map, and outputting the reconstructed path information in the user interaction module. When cognitive map construction is in progress, the cognitive map construction process will be terminated when the instruction (90) is received again, and other instructions will not interfere with the cognitive map construction process that has been started.
[0114] Instruction (91): Perceptual information recognition - input a random symbol sound / image / neuromorphic image, and the system outputs the perceived symbol content. This process can continue until switching to another task (that is, when the perceptual cortical module perceives "9" again).
[0115] Instruction (92): Offline reinforcement learning - based on a given decision network, continue to perform pre-set decision-making behaviors (such as offline balancing pole games, etc.). When (92) is received again, the reinforcement learning task will end. When (93) is received, it will switch to the online reinforcement learning task, and other instructions will not interfere with it.
[0116] Instruction (93): Online reinforcement learning - module (3) will start working only when the instruction (92) is input. Therefore, the instruction required to execute the task from the beginning is (92)(93). The task will stop when switching tasks (that is, when the perceptual cortex module perceives "9" again, (92)(93) needs to be input again to enter the task again). In this task, the system will decide a number with the largest expected value among several numbers and receive a reward to adjust the decision-making mode. The judgment of expected value will continue to change with the feedback of reward value.
[0117] By inputting a specific picture / sound sequence before each task, the execution of different tasks can be controlled. In addition, the system is not task-oriented and is a continuous and uninterrupted system. Therefore, not only can the same modules be used between different tasks, but they can also be performed simultaneously without conflicting modules, and new tasks can be performed by expanding new modules. Through the iteration of the built-in network model, the system can perform more complex decisions and gradually approach the multi-brain parallelism of the human brain.
[0118] Reference Figure 5 You can see the diagram of the relevant brain areas in the process of cognitive map construction. Figure 5 The part above the double arrows is the change of the control signal encoded by the gating module that receives the relevant sensory signal. It can be seen that the signals encoded by the gating module are initially inhibitory signals. When the gating module receives the information number "9" represented by a pulse from the sensory cortex module, the gating module activates itself, but at this time the input of the cognitive map construction module is still inhibited by the control signal, so it is unable to receive input to build a cognitive map. Subsequently, when the gating module receives the information represented by a pulse from the sensory cortex module, which is the number "0", the gating module activates the control signal sent to the cognitive map construction module and inhibits itself, so that the cognitive map construction module begins to build a cognitive map: the self-movement information (angle / distance) with a slight error is input into the grid cells to drive the activity of the simulated grid cells, and at the same time, the connection strength between the simulated place cells and the grid cells is adjusted through the Hebbian learning rule, and the simulated place cells are connected to drive the path integration.
[0119] At the same time, continuous visual input is received and stored in the cognitive map along with the activity of the place cells, which is used to correct the accumulated errors in the path integral when the path closure is detected visually. Until the cortical perception module perceives the same signal sequence again (or perceives the perception information that needs to trigger the conflict module).
[0120] Figure 6The performance of each module of the system in the perception recognition task is shown. From top to bottom, four perception signals "9", "1", "4", and "1" are input (and they actually form different channels), and the signal transmission and reception of related brain areas are shown. It can be seen that the perception signal module integrates the information of all relevant modalities and outputs it, and does not require corresponding input for each perceptible modality. And the task can be executed until it is stopped by a new instruction.
[0121] Figure 7 The performance of each module of the system during the online reinforcement learning task is shown. The sensory signal inputs of the previously opened pathways are ignored. The initial model decision is random, and then after each decision, the sensory signal is received again. The reward generation submodule generates the corresponding reward value, and the decision submodule stores the decision record in the working memory module, and then wakes up several decision records from the working memory module to learn the decision strategy, and eventually it will slowly learn to make decisions with greater potential rewards.
[0122] In both cases, the cognitive map construction task can still be performed simultaneously without interfering brain regions.
[0123] In general, the brain-like perception-learning-decision system is a simulated brain decision model with the perception cortex module, gating module, reward and decision module, maneuvering module, and cognitive map construction module as the core. It uses the perception cortex module to receive the surrounding perception signals and transmits them to the gating module for control encoding to inhibit / activate several functionally independent brain areas, realizing a single model and multiple tasks.
[0124] As a popular subject, neural networks have achieved corresponding results in various fields, but progress in simulating human brain behavior has been slow. Due to the complexity and correspondence between human brain behavior and human behavior, and the structural differences of different functional networks, the constructed brain simulation systems often have problems such as poor scalability, complex execution logic, and low biological rationality.
[0125] The beneficial effect of the present invention is to construct a feasible brain-like simulation system and a corresponding framework, which not only realizes brain-like perception, decision-making, and learning, and can perform multiple tasks simultaneously, but also retains the characteristics of different neural circuits in multiple brain regions to realize specific functions.
[0126] The present invention can expand new tasks only by retraining the execution order of the new tasks for the network inside the gating module, thereby solving the first problem existing in the existing simulated brain system. Due to the existence of the gating module, multiple brain regions with independent functions can exist in parallel in one system, and the gating module can perform different functions according to the state of the perception module and the system, thereby realizing the coordination of the functions of each brain region and the multi-brain region system, thereby solving the second problem existing in the existing simulated brain system.
[0127] The embodiment of the present invention also provides a brain-like perception-learning-decision-making method, such as Figure 8 As shown, including:
[0128] Step 1: Use attention mechanism or Bayesian decision to integrate at least one acquired sensory information into a sensory pulse sequence.
[0129] Step 2: splice the latest control signal with the sensing pulse sequence, and automatically encode and output the control signal at the next moment according to the spliced signal.
[0130] Step 3: When the control signal at the next moment controls the output of the perception pulse sequence, a corresponding reward signal is obtained according to the perception pulse sequence, and a corresponding operating mode is selected according to the control signal, and then a decision signal is generated by adopting the corresponding operating mode according to the reward signal.
[0131] Step 4: When the control signal at the next moment suppresses the output of the perception pulse sequence, visual information and odometer information are received, and a cognitive map is continuously constructed based on the visual information and odometer information.
[0132] Step 5: Comprehensively output the decision signals and / or cognitive maps at all times.
[0133] As for the method disclosed in the embodiment, since it corresponds to the system disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the system part.
[0134] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0135] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A brain-like perception-learning-decision-making system, It is characterized in that The system includes: a perception cortex module, a gating module, a reward and decision module, a working memory module, a cognitive map construction module and a user interaction module; The perceptual cortex module is used to input at least one perceptual information, integrate all the input perceptual information into a perceptual pulse sequence by using an attention mechanism or a Bayesian decision, and transmit the perceptual pulse sequence to the gating module; The gating module is used to splice the latest control signal with the perception pulse sequence, and automatically encode and output the control signal at the next moment according to the spliced signal, further control whether the perception pulse sequence is output according to the control signal at the next moment, and transmit the control signal at the next moment to the cognitive map construction module and / or the reward and decision module; The reward and decision module is used to input the perception pulse sequence after receiving the control signal at the next moment, and obtain a corresponding reward signal according to the perception pulse sequence, and select a corresponding operation mode according to the control signal, and then generate a decision signal according to the reward signal using the corresponding operation mode; The cognitive map construction module is used to input the visual information and odometer information in the perception cortex module after receiving the control signal at the next moment, and continuously construct the cognitive map according to the visual information and odometer information; The user interaction module is used to display the specific content represented by the decision signal and / or to display the constructed cognitive map in an image form; The cognitive map construction module includes: a grid cell submodule, a place cell submodule, a visual cell submodule and an experience map construction submodule; The grid cells in the grid cell submodule are modeled using a continuous attractor network. The grid cells are driven and activated by odometer information. The neural activity of the grid cells changes as the agent moves, and the connection with the place cell submodule drives the activity of the place cells to represent the current location information of the agent. The visual cells in the visual cell submodule are driven and activated by visual information, representing the visual feature information of the current environment of the agent, and are transmitted to the experience map construction submodule together with the position information; The experience map construction submodule is used to update the experience map based on the visual feature information and the position information, and to correct the accumulated errors by combining the visual information with the updated experience map, so as to continuously generate the cognitive map and output it to the user interaction module.
2. The brain-like perception-learning-decision-making system according to claim 1, It is characterized in that The perceptual cortex module includes: multiple neural networks and an attention-based modality fusion layer; Each neural network is used to extract features and identify targets from the sensory information input by a channel, and obtain the recognition result of a channel; The attention-based modality fusion layer is used to perform modality fusion on the recognition results output by all neural networks to synthesize a perception pulse sequence.
3. The brain-like perception-learning-decision-making system according to claim 2, It is characterized in that When the perception information is visual information, the neural network is a convolutional neural network based on LIF neurons; the visual information includes dynamic visual information and static visual information; When the sensory information is auditory information, the neural network is a recurrent spiking neural network based on CLIF neurons.
4. The brain-like perception-learning-decision-making system according to claim 1, It is characterized in that The gate control module includes a pulse neural network trained according to the control signal issuance rule.
5. The brain-like perception-learning-decision-making system according to claim 1, It is characterized in that The system further comprises: a working memory module and a motorized output module; The working memory module is connected to the reward and decision module; the working memory module is used to receive the reward signal and the corresponding decision signal of each brain region, and store the decision record after forming it; The motorized output module is connected to the gate control module; when the gate control module transmits the control signal of the next moment to the reward and decision module, the gate control module simultaneously transmits the control signal of the next moment to the motorized output module; The motorized output module is also connected to the user interaction module and the reward and decision module respectively; the motorized output module is used to output the decision signal generated by the reward and decision module to the user interaction module.
6. The brain-like perception-learning-decision-making system according to claim 5, It is characterized in that The reward and decision module includes: a reward submodule and a decision submodule; The reward submodule is connected to the decision submodule, and the reward submodule is used to obtain a reward signal corresponding to the perception signal in the perception pulse sequence, and transmit the reward signal to the decision submodule; The decision submodule is connected to the gating module and the working memory module respectively. The decision submodule is used to select a corresponding operation mode according to a control signal. When the selected operation mode is an offline decision mode, a preset decision signal is directly outputted. When the selected operation mode is an online decision mode, after receiving a reward signal, a plurality of decision records are read from the working memory module, and a decision model is updated according to the plurality of decision records, and then a decision signal is generated by using the updated decision model according to the reward signal.
7. The brain-like perception-learning-decision-making system according to claim 1, It is characterized in that The experience map is updated based on the visual feature information and the location information, and the accumulated errors are corrected by combining the visual information with the updated experience map, so as to continuously generate the cognitive map, including: The visual feature information and the position information are used as trajectory information and compared with the experience map stored in the experience map construction submodule; If the trajectory information overlaps with the experience map, the grid cell and place cell firing at the current point are reset to the grid cell and place cell firing of the matched experience; Iteratively update the global pose of each experience point along the experience map based on the coincidence information; In the process of updating the global pose of the experience point in each iteration, all experience points and connections in the experience map are traversed, and the global pose of the two connected experience points is corrected according to the direction and distance of each connection, so that the global pose difference between the two experience points converges to the direction of the connection between the two experience points, generating a cognitive map.
8. The brain-like perception-learning-decision-making system according to claim 1, It is characterized in that The user interaction module is used to provide a graphical interaction interface, input signals on the graphical interaction interface, and obtain visual execution results; generate corresponding signals according to a given instruction sequence and send them to the perception cortex module; execute different instructions according to the analysis of the instruction sequence, and display the information and activation status in the perception cortex module, gating module, reward and decision module, working memory module, and cognitive map construction module.
9. A brain-like perception-learning-decision-making method, It is characterized in that include: Attention mechanism or Bayesian decision is used to integrate at least one acquired sensory information into a sensory pulse sequence; splicing the latest control signal with the sensing pulse sequence, and automatically encoding and outputting the control signal at the next moment according to the spliced signal; When the control signal at the next moment controls the output of the perception pulse sequence, a corresponding reward signal is obtained according to the perception pulse sequence, and a corresponding operation mode is selected according to the control signal, and then a decision signal is generated according to the reward signal using the corresponding operation mode; When the control signal at the next moment inhibits the output of the perception pulse sequence, receiving visual information and odometer information, and continuously constructing a cognitive map according to the visual information and odometer information; Comprehensively output the decision signals and / or cognitive maps at all times; The construction of cognitive maps specifically includes: grid cells are modeled using a continuous attractor network, the grid cells are driven and activated by odometer information, the neural activity of the grid cells changes as the agent moves, and is connected to drive the activity of position cells to represent the current position information of the agent; visual cells are driven and activated by visual information, representing the visual feature information of the current environment of the agent; the experience map is updated based on the visual feature information and position information, and the accumulated errors are corrected through the visual information combined with the updated experience map, thereby continuously generating cognitive maps.
Citation Information
Patent Citations
Multi-brain area cooperative autonomous decision making method based on multi-modal fusion
CN108197698A
Improved hippocampus-forehead cortex network space cognition method
CN114186675A