Robot cognitive decision-making method and system
By adopting cognitive decision-making methods of multimodal information fusion and up-down attention mechanisms in robots, the problem of difficulty in imitating advanced cognitive functions of humans in the prior art is solved, and efficient decision-making and adaptability improvement of robots in complex environments is achieved.
Patent Information
- Application Number
- CN202510211646.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to mimic humans simultaneously using perception, knowledge, internal states and goals to complete advanced cognitive functions and decision-making tasks, especially in complex dynamic environments to deal with noisy, weak or vague stimuli, and cannot cope with real-time challenges and goal changes.
A robot cognitive decision-making method is adopted to generate multimodal stimulation aggregation signals by receiving the sound information and visual information corresponding to multiple task objectives in real time, and combine the bottom-up and top-down attention mechanisms to form cognitive decision-making information with a comprehensive impact through delay processing and reward reinforcement mechanisms.
It realizes that robots make efficient cognitive decisions in complex environments, integrate multimodal information and internal knowledge, can respond to real-time challenges and goal changes, and improves the accuracy and adaptability of decisions.
Smart Images

Figure CN120023809A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to robot cognitive control, and more specifically, relates to a robot cognitive decision-making method and system. Background Art
[0002] As intelligent manufacturing, smart services, unmanned systems and other fields are booming, robots are rapidly moving from structured closed scenarios to dynamic open environments. In this process, cognitive decision-making capabilities have become the core constraint for robots to achieve autonomy and intelligence. Efficient cognitive decision-making not only requires robots to respond to environmental changes in real time, quickly analyze multimodal sensor signals such as vision, voice, and touch, and promptly capture dynamic information such as sudden obstacles and command updates; it also requires robots to be able to integrate experience and knowledge, and use historical behavior data and task goals to generate adaptive strategies, such as optimizing paths and predicting human intentions.
[0003] The current mainstream cognitive decision-making technologies mainly include rule-driven systems, reinforcement learning, and multimodal fusion decision-making. Rule-driven systems such as finite state machines and decision trees have the advantages of transparent logic and predictable results, but they rely heavily on manually preset rules and are helpless in the face of undefined scenarios, such as when home robots encounter new furniture layouts. Although reinforcement learning supports autonomous learning in dynamic environments, it ignores the integration of multimodal signals and often focuses only on the information input of a single modality, ignoring the rich environmental information contained in other modal signals. This makes the robot's decision-making basis one-sided and the trial-and-error cost extremely high when facing complex and changeable real scenes with diverse information. There is an averaging trap in the multimodal fusion solution. When there is a cross-modal conflict, the conflicting signals are simply processed, which violates the biological decision priority. For example, when there is a contradiction between "voice command to move forward" and "visual detection obstacles", there is a lack of effective arbitration logic. These problems seriously limit the decision-making ability of robots in complex scenarios and urgently need new ideas and methods to solve them. Summary of the invention
[0004] In response to the above defects or improvement needs of the prior art, the present invention provides a robot cognitive decision-making method and system, which aims to solve the problems existing in the prior art that it is unable to imitate humans to simultaneously use perception, knowledge, internal state and goals to complete advanced cognitive functions and decision-making tasks, it is difficult to handle noisy, weak or ambiguous stimuli in complex dynamic environments, and it is unable to cope with real-time challenges and target changes.
[0005] To achieve the above object, according to one aspect of the present invention, a robot cognitive decision-making method is provided, comprising:
[0006] Receive the sound information and visual information corresponding to n task objectives in real time, where n ≥ 1; fuse the sound information and visual information corresponding to each task objective to obtain a multimodal stimulus convergence signal, and the fusion duration corresponding to different task objectives is used to reflect the difference in the capture speed of the attention based on incentive drive for each task objective; set the start time of each convergence signal as the start time of the attention based on incentive drive for the corresponding task objective, input the convergence signal into the time-delay processing unit to determine the self-sustaining duration of the attention based on incentive drive for the corresponding task objective, and the self-sustaining durations corresponding to different task objectives are used to reflect the difference in the attention maintenance durations for each task objective, so as to obtain the bottom-up incentive-driven attention signal for each task objective.
[0007] Receive the decision reward information and context information in real time; use the currently received decision reward information to cumulatively reward the current cognitive decision information for executing a certain task objective, and when the cumulative reward value reaches the threshold, form a continuous top-down reward-reinforced attention signal for the certain task objective; judge whether the currently received context information has changed, and if it has changed, reset the cumulative reward values corresponding to all task objectives.
[0008] Sum the bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal for each task objective, and map the largest value among the summation results corresponding to each task objective to obtain the current cognitive decision information for executing a certain task objective, so as to realize the cognitive decision of the robot.
[0009] Furthermore, the implementation method for fusing the sound information and visual information corresponding to each task objective to obtain a multimodal stimulus convergence signal is as follows:
[0010] Generate analog signals of the received sound information and visual information by sensors arranged in the robot, and sequentially perform noise filtering, synchronous integration and encoding on each analog signal to obtain two spike signals; perform a logical operation between the two spike signals corresponding to each task objective to obtain a multimodal stimulus convergence signal corresponding to the task objective.
[0011] Furthermore, the time-delay processing unit is a time-delay processing circuit based on memristors.
[0012] Furthermore, the implementation method for forming a continuous top-down reward-reinforced attention signal for a single task objective is as follows:
[0013] An electrical signal representing the currently received decision reward information is formed. When no decision reward information is input, the electrical signal representing the corresponding information is set to zero. An AND operation is performed on the electrical signal representing the currently received decision reward information and the electrical signal representing the current cognitive decision information, and the AND operation results are accumulated and integrated. If the integrated signal reaches a threshold value, an uninterrupted and continuous top-down reward reinforcement attention signal is formed for the task goal corresponding to the current cognitive decision information.
[0014] According to another aspect of the present invention, there is provided a robot cognitive decision-making system, comprising: an audio-visual receiving component, an incentive-driven attention control module, a context buffer module, and a central executive control module;
[0015] The audio-visual receiving component is used to receive the sound information and visual information corresponding to n task targets in real time, where n≥1;
[0016] The incentive-driven attention control module is used to fuse the sound information and visual information corresponding to each task target to obtain a multimodal stimulus convergence signal, and the fusion duration corresponding to different task targets is used to reflect the difference between the capture speeds of the incentive-driven attention of each task target; the start time of each convergence signal is set as the start time of the incentive-driven attention to the corresponding task target, and the convergence signal is input into the delay processing unit to determine the self-maintaining duration of the incentive-driven attention to the corresponding task target, and the self-maintaining duration corresponding to different task targets is used to reflect the difference between the attention maintenance durations of each task target, so as to obtain a bottom-up incentive-driven attention signal for each task target;
[0017] The context buffer module includes a reward reinforcement unit and a context control unit, wherein the reward reinforcement unit is used to receive decision reward information and context information in real time; the decision reward information currently received is used to accumulate rewards for the current cognitive decision information used to execute a certain task target, and when the accumulated reward value reaches a threshold, a continuous top-down reward reinforcement attention signal for the certain task target is formed; the context control unit is used to determine whether the currently received context information has changed, and if so, reset the accumulated reward values corresponding to all task targets;
[0018] The central executive control module is used to sum the bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal of each task target, and map the largest summation result corresponding to n task targets to obtain current cognitive decision information.
[0019] Furthermore, the audio-visual receiving component includes a speech loop and a visual space board, which are respectively used to receive analog signals of sound information and visual information based on sensors, and perform noise filtering, synchronous integration and encoding on each analog signal in turn to obtain a spike signal.
[0020] Furthermore, when the attention control module fuses the sound information and visual information corresponding to each task target to obtain a multimodal stimulation convergence signal, the implementation method is: performing a logical operation between the two peak signals corresponding to each task target to obtain a multimodal stimulation convergence signal corresponding to the task target.
[0021] Furthermore, the delay processing unit is a delay processing circuit based on a memristor.
[0022] Furthermore, when the reward reinforcement unit forms a continuous top-down reward reinforcement attention signal for a single task target, the implementation method is as follows:
[0023] An electrical signal representing the currently received decision reward information is formed. When no decision reward information is input, the electrical signal representing the corresponding information is set to zero. An AND operation is performed on the electrical signal representing the currently received decision reward information and the electrical signal representing the current cognitive decision information, and the AND operation results are accumulated and integrated. If the integrated signal reaches a threshold value, an uninterrupted and continuous top-down reward reinforcement attention signal is formed for the task goal corresponding to the current cognitive decision information.
[0024] According to another aspect of the present invention, a robot is provided for implementing the steps of a robot cognitive decision-making method as described above when executing a task objective.
[0025] In general, compared with the prior art, the technical solution conceived by the present invention has the following beneficial effects:
[0026] 1. The present invention proposes a robot cognitive decision-making method. For each task goal decision selection, it is necessary to include a bottom-up attention pathway (audio-visual information reception and incentive-driven attention signal generation) and a top-down attention pathway (receiving reward reinforcement information, situational information reception and cognitive decision information of current decision feedback, controlling the formation of reward-reinforced attention signals). The final cognitive decision will compare the combined effects of top-down and bottom-up attention, that is, summing the bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal of each task goal, and mapping the largest sum of the corresponding results of n task goals to obtain the current cognitive decision information for executing one of the task goals. The method of the present invention integrates the working memory framework in cognitive neuroscience and the two attention mechanisms of top-down and bottom-up. It can not only integrate and encode external audio-visual modal information, but also retrieve relevant information from internal knowledge and experience, providing an interface connecting perception, internal state and action, thereby realizing goal-oriented behavior and decision-making.
[0027] 2. The present invention further proposes to use analog signals of external environmental information for processing. The intensity of the analog signal reflects the saliency of the external environmental stimulus, and can distinguish and process different audio-visual input modes and different stimulus saliencies. The form of the spike signal can effectively realize the differentiation of the bottom-up attention capture speed and self-maintenance time under different sensory information input modes and different stimulus saliencies, simulating the brain's processing and integration of different sensory information, and the adjustment of attention resources under different cognitive needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a flowchart of a robot cognitive decision-making method provided by an embodiment of the present invention;
[0029] Figure 2 is a schematic diagram of a robot cognitive decision-making model for a single task objective provided by an embodiment of the present invention;
[0030] Figure 3 It is a schematic diagram of a robot cognitive decision-making model for two task objectives provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0032] Embodiment 1
[0033] A robot cognitive decision-making method, such as Figure 1 Said, including:
[0034] Receive the sound information and visual information corresponding to n task targets in real time, where n≥1; fuse the sound information and visual information corresponding to each task target to obtain a multimodal stimulus convergence signal, and the fusion duration corresponding to different task targets is used to reflect the difference between the capture speeds of the attention driven by the incentive for each task target; set the start time of each convergence signal as the start time of the attention driven by the incentive for the corresponding task target, input the convergence signal into the delay processing unit to determine the self-maintaining duration of the attention driven by the incentive for the corresponding task target, and the self-maintaining duration corresponding to different task targets is used to reflect the difference between the attention maintenance durations for each task target, so as to obtain the bottom-up incentive-driven attention signal for each task target;
[0035] Receive decision reward information and context information in real time; use the currently received decision reward information to cumulatively reward the current cognitive decision information for executing a certain task goal. When the cumulative reward value reaches a threshold, form an attention signal for continuous top-down reward reinforcement for the certain task goal; determine whether the currently received context information has changed. If it has changed, reset the cumulative reward values corresponding to all task goals.
[0036] Sum the bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal for each task goal, and map the largest value among the summation results corresponding to each task goal to obtain the current cognitive decision information for executing one of the task goals, thereby realizing the cognitive decision of the robot.
[0037] Aiming at the problems existing in the existing cognitive decision-making, the method of this embodiment proposes a solution starting from the cognitive decision-making mechanism of the human brain. From the perspective of cognitive neuroscience, human decision-making depends on continuously screening out task-related information from external multimodal stimuli and internal memory representations and other states. The external environment and internal regulation are the two most basic and crucial information channels in the biological brain. Therefore, the cognitive decision-making design that only considers one information channel does not conform to the situation where humans use perception, knowledge, and goals simultaneously to complete higher cognitive functions.
[0038] Furthermore, the method of this embodiment provides the bionic robot with multimodal information processing and internal state guidance capabilities based on working memory and attention mechanisms, thereby making up for the limitations of the existing memory and decision-making frameworks in enhancing cognitive learning and cognitive decision-making functions. Among them, working memory, as an important subfield of short-term memory, can temporarily load knowledge or experience in long-term memory, process and encode multimodal signals in sensory memory, and connect sensory memory and long-term memory. More and more evidence helps to establish the connection between working memory and the ability to switch between different tasks, maintain attention, and inhibit irrelevant data. It is a key cognitive skill, mainly related to adaptive strategies, task switching, decision-making, and reasoning, enabling organisms to make appropriate decisions quickly when adapting to the environment.
[0039] In addition, attention and working memory show similar patterns in the cerebral cortex, indicating that they utilize common neural mechanisms. The completion of cognitive control depends on the realization of two forms of attention: top-down, goal-driven attention, and bottom-up, stimulus-driven attention. These two forms of attention are determined by the individual's current behavioral goals and shaped by learning priorities based on personal experience and evolutionary adaptation, and are always in competition or cooperation: one supports the ability to concentrate attention to achieve instantaneous behavioral goals, and the other is driven by the surrounding world.
[0040] Traditional cognitive and decision-making models lack effective integration of working memory and attention mechanisms, resulting in the model being unable to comprehensively consider factors such as perception, knowledge, internal state, and goals that truly affect the final behavioral decision in the human cognitive process. This makes robots have poor adaptability, low accuracy, and limited flexibility in advanced cognitive functions and decision-making tasks, and it is difficult for them to have real-time learning capabilities, especially in complex environments, and there is still a big gap between the models and real human judgment.
[0041] The method of this embodiment provides a cognitive decision-making model based on working memory and attention mechanism, integrating the working memory framework in cognitive neuroscience and the top-down and bottom-up attention mechanism to achieve the integrated encoding of external audiovisual modal information and the retrieval of internal knowledge and experience. It solves the problems existing in the prior art, such as the inability to imitate humans to use perception, knowledge, internal state and goals at the same time to complete advanced cognitive functions and decision-making tasks, the difficulty in handling noisy, weak or ambiguous stimuli in complex dynamic environments, and the inability to cope with real-time challenges and target changes.
[0042] Specifically, for each task goal decision selection, it is necessary to include a bottom-up attention pathway (audio-visual information reception and incentive-driven attention signal generation) and a top-down attention pathway (receiving reward reinforcement information, receiving contextual information, and cognitive decision information of current decision feedback, controlling the formation of reward-reinforced attention signals). The final cognitive decision will compare the combined effects of top-down and bottom-up attention, that is, summing the bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal for each task goal, mapping the largest of the summed results corresponding to n task goals to obtain the current cognitive decision information, preferably in combination with the comparison of the quality of the encoded information. Assuming that the potential behavior result of the robot is binary: action 1 or action 2, this embodiment uses the interaction between the stimulus-driven bottom-up attention and the reward-reinforced top-down attention to achieve the influence on the robot's cognitive decision.
[0043] To be more specific, for the bottom-up attention pathway, the sensor installed in the robot first generates analog signals of the received sound information and visual information, and each analog signal is noise filtered, synchronously integrated and encoded in turn to obtain two spike signals, so as to achieve the differentiation of the bottom-up attention capture speed and self-maintenance time under different sensory information input modes and different stimulus saliences, and simulate the brain's processing and integration of different sensory information, as well as the adjustment of attention resources under different cognitive needs: when the external environment stimulus has only visual or auditory unimodal signals, the brain processes sensory information slowly, and the bottom-up attention capture is slow and the maintenance time is short; when there is significant stimulus input from both vision and hearing in the external environment, the visual and auditory information promote each other, and attention is captured faster and maintained longer; when visual and auditory signals exist at the same time and the salience is reduced, attention capture becomes slower and the maintenance time becomes shorter. Unless relevant significant information is applied continuously, the attention will be maintained for a short time and eventually disappear. Since the bottom-up incentive-driven attention signals and top-down reward-enhanced attention signals of each task goal will be mapped to obtain cognitive decisions, different incentive sizes and types will affect subsequent cognitive decisions. The method of this embodiment can handle different audio-visual input modes and different stimulus saliences, and consider the interaction of audio-visual stimuli, using the information synergy effect between modalities to enhance the attention and memory processes of relevant brain areas.
[0044] In addition, for the top-down attention pathway, the method of this embodiment accumulates rewards for the current cognitive decision information used to execute a certain task goal by targeting the currently received decision reward information. When the accumulated reward value reaches the threshold, a continuous top-down reward-reinforced attention signal for the certain task goal is formed, thereby giving rewards multiple times in a certain cognitive decision action, creating an association between stored rewards and behaviors, and this internal storage bias will not disappear over time, achieving the accumulation of experience and extracting useful information from experience. The bias will only disappear when the received context control information changes. The accumulated reward values corresponding to each task goal are stored and accumulated independently, and an attention signal of reward reinforcement for the corresponding task goal is independently formed. Therefore, the method of this embodiment also has the ability to selectively pay attention and suppress irrelevant data, realizes the formation and reset of attention bias through decision feedback, reward reinforcement and context control, and significantly optimizes the performance of working memory to handle noisy, weak or blurred perceptual stimuli in complex dynamic environments, and cope with real-time challenges such as target changes and context switching.
[0045] For the final cognitive decision step, each task goal corresponds to two attention pathways. The final cognitive decision connects the bottom-up attention pathway and the top-down attention pathway at the same time, that is, it receives attention signals from the two pathways at the same time, and by summing the bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal of each task goal, the largest sum corresponding to the n task goals is mapped to obtain the current cognitive decision information, realizing the synergy, complementarity or competition relationship between the two attention pathways, thereby transforming from a random exploration process to a learnable goal-oriented behavior, outputting real-time cognitive decision information, and also serving as the output of the entire system. Among them, synergy means that when the external perceptual stimulation and the bias of internal storage both point to the same decision, the two attention pathways work together, the tendency of behavior execution is stronger, the attention is more focused, and the decision-making speed is faster; complementarity means that when the bottom-up external perceptual stimulation is noisy, weak or vague, since the stimulation signal will be filtered, if the audiovisual perceptual stimulation is irrelevant data, it will be filtered out, and the top-down attention signal will not be generated. At this time, the information of the two attention pathways is complementary, that is, the final cognitive decision is made through the top-down internal bias (that is, the top-down reward-reinforced attention signal); competition means that when the external perceptual stimulation and the bias of internal storage point to two different decision choices, the external perceptual stimulation essentially contains the environmental externalities that are crucial to the perception and execution of the organism, and is given a higher priority. In order to complete subsequent new decisions, realize continuous decision-making and continuous learning, this method can also complete the reset operation through context control.
[0046] Therefore, the method of this embodiment integrates the working memory framework in cognitive neuroscience and the two attention mechanisms of top-down and bottom-up, which can not only integrate and encode external audiovisual modal information, but also retrieve relevant information from internal knowledge and experience, providing an interface connecting perception, internal state and action, thereby realizing goal-oriented behavior and decision-making.
[0047] The method of this embodiment can be applied to robot cognitive control systems, which can comprehensively consider factors that truly affect the final behavioral decision in the human cognitive process, such as perception, knowledge, internal state and goals, so that the robot has stronger accuracy, real-time and learning ability in advanced cognitive functions and decision-making tasks such as intention analysis, path planning, random exploration, cognitive decision-making, etc., and is closer to real human judgment.
[0048] As a preferred implementation, the method of fusing the sound information and visual information corresponding to each task target to obtain a multimodal stimulus convergence signal is as follows:
[0049] The sensor installed in the robot generates analog signals of the received sound information and visual information, and each analog signal is noise-filtered, synchronously integrated and encoded in turn to obtain two peak signals; a logical operation is performed between the two peak signals corresponding to each task target to obtain a multimodal stimulus convergence signal corresponding to the task target.
[0050] The intensity of the simulated signal reflects the saliency of the external environment stimulus, and can distinguish and process different audiovisual input modes and different stimulus saliency. The form of the spike signal can effectively achieve the differentiation of the bottom-up attention capture speed and self-maintenance time under different sensory information input modes and different stimulus saliency, simulating the brain's processing and integration of different sensory information, as well as the adjustment of attention resources under different cognitive needs. Among them, filtering refers to filtering out input signals with low task relevance or intensity to avoid interference with subsequent processing and decision-making; integration refers to the use of a signal integration mode based on the same frequency to ensure that working memory can be used as an interface to connect multimodal inputs; encoding refers to the ability to generate spikes with millisecond accuracy similar to that of biological neurons, providing strong anti-interference capabilities.
[0051] As a preferred implementation, the delay processing unit is a delay processing circuit based on a memristor, which simplifies the circuit design and facilitates implementation.
[0052] As a preferred implementation, the method for forming a continuous top-down reward-reinforced attention signal for a single-task goal is:
[0053] An electrical signal representing the currently received decision reward information is formed. When no decision reward information is input, the electrical signal representing the corresponding information is set to zero. An AND operation is performed on the electrical signal representing the currently received decision reward information and the electrical signal representing the current cognitive decision information, and the AND operation results are accumulated and integrated. If the integrated signal reaches a threshold value, an uninterrupted and continuous top-down reward reinforcement attention signal is formed for the task goal corresponding to the current cognitive decision information.
[0054] Embodiment 2
[0055] A robot cognitive decision-making system, comprising: an audio-visual receiving component, an incentive-driven attention control module, a context buffer module, and a central executive control module;
[0056] The audio-visual receiving component is used to synchronously receive the sound information and visual information corresponding to n task targets in real time, where n ≥ 1;
[0057] The incentive-driven attention control module is used to fuse the sound information and visual information corresponding to each task target to obtain a multimodal stimulus convergence signal, and the fusion duration corresponding to different task targets is used to reflect the difference between the capture speeds of the incentive-driven attention of each task target; the start time of each convergence signal is set as the start time of the incentive-driven attention of the corresponding task target, and the convergence signal is input into the delay processing unit to determine the self-maintaining duration of the incentive-driven attention of the corresponding task target, and the self-maintaining duration corresponding to different task targets is used to reflect the difference between the attention maintenance durations of each task target, so as to obtain a bottom-up incentive-driven attention signal for each task target; the attention control module outputs the bottom-up stimulus-driven attention signal to the central executive control module;
[0058] The context buffer module includes a reward reinforcement unit and a context control unit. The reward reinforcement unit is used to receive decision reward information and context information in real time; the decision reward information currently received is used to accumulate rewards for the current cognitive decision information used to execute a certain task goal. When the accumulated reward value reaches a threshold, a continuous top-down reward reinforcement attention signal for the certain task goal is formed; the context control unit is used to determine whether the currently received context information has changed. If it has changed, the accumulated reward values corresponding to all task goals are reset;
[0059] The central executive control module is used to sum the bottom-up incentive-driven attention signals and the top-down reward-reinforced attention signals of each task goal, and map the largest sum of the results corresponding to each task goal to obtain the current cognitive decision information used to execute one of the task goals.
[0060] For a single target task, cognitive decision-making models based on working memory and attention mechanisms are as follows: Figure 2As shown, it includes an audio-visual receiving component (which may preferably include a speech loop and a visual-spatial board), an attention control module, a context buffer (i.e., a context buffer module), and a central executive control module; wherein the model can be further divided into: a bottom-up attention pathway (including a speech loop, a visual-spatial board, and an attention control module), a top-down attention pathway (including a context buffer and feedback information from a central executive control module), and a central executive control module; the speech loop and the visual-spatial board receive auditory and visual stimuli in the environment as inputs to the model and are connected to the attention control module; the central executive control module simultaneously receives signals from the context buffer in the top-down attention pathway and from the attention control module in the bottom-up attention pathway, and finally obtains a cognitive decision signal as the output of the model and feeds back to the context buffer. For the case where there are multiple task objectives at the same time, it can be regarded as the presence of multiple cognitive decision models in the robot, which cooperate, complement, or compete with each other, such as Figure 3 shown.
[0061] The attention control module can be divided into an attention capture unit and an attention self-maintenance unit connected in sequence, which can be used to: capture attention driven by bottom-up stimuli, and realize self-maintenance of attention after the disappearance of the stimulus, and realize the differentiation of attention capture speed and maintenance time under different sensory information input modes and different stimulus salience. The difference reflects how the brain processes and integrates information from different senses, and how to adjust attention resources when facing different cognitive needs. Among them, unless relevant salient information is continuously applied, attention will be maintained for a short time and eventually disappear.
[0062] As a preferred embodiment, the audio-visual receiving component includes a speech loop and a visual space board, which are respectively used to receive analog signals of sound information and visual information based on sensors, and each analog signal is noise filtered, synchronously integrated and encoded in turn to obtain a spike signal.
[0063] The speech loop and the visual-spatial board directly receive the analog signal from the sensor as the system input. The strength of the analog signal reflects the salience of the external environment stimulus. The speech loop and the visual-spatial board work in parallel. The speech loop processes language or sound information, and the visual-spatial board processes visual or spatial information, ensuring that auditory and visual-spatial data are effectively integrated and utilized. After filtering, integration and encoding in the speech loop and the visual-spatial board, the two output spike signals (digital signals) are converged in the attention control modules connected in sequence.
[0064] As a preferred implementation mode, when the attention control module fuses the sound information and visual information corresponding to each task target to obtain a multimodal stimulation convergence signal, the implementation method is: performing a logical operation between the two peak signals corresponding to each task target to obtain a multimodal stimulation convergence signal corresponding to the task target.
[0065] As a preferred implementation, the delay processing unit is a delay processing circuit based on a memristor, which simplifies the circuit design and facilitates implementation.
[0066] As a preferred implementation, when the reward reinforcement unit forms a continuous top-down reward reinforcement attention signal for a single task target, the implementation method is as follows:
[0067] An electrical signal representing the currently received decision reward information is formed. When no decision reward information is input, the electrical signal representing the corresponding information is set to zero. An AND operation is performed on the electrical signal representing the currently received decision reward information and the electrical signal representing the current cognitive decision information, and the AND operation results are accumulated and integrated. If the integrated signal reaches a threshold value, an uninterrupted and continuous top-down reward reinforcement attention signal is formed for the task goal corresponding to the current cognitive decision information.
[0068] In this embodiment, the context buffer module (also called the context buffer) includes a reward reinforcement unit and a context control unit connected in sequence. Among them, the reward reinforcement unit receives in real time the electrical signal reflecting the reward information and the cognitive decision information (electrical signal) currently fed back by the central executive control module. When there is no reward information, the electrical signal reflecting the reward information is zero; the two electrical signals are ANDed, and the AND operation results are accumulated and integrated. If the accumulated integration result reaches the threshold, a top-down reward-enhanced attention signal (also called an attention bias signal) is formed, and it is continuously output to the central executive control module. If the accumulated integration result does not reach the threshold, there is no top-down reward-enhanced attention signal output. The accumulated reward values corresponding to each task target are independently stored and accumulated, and the reward-enhanced attention signal for the corresponding task target is independently formed. By giving rewards multiple times in a cognitive decision action, the association between rewards and behaviors is created and stored. This internally stored bias will not disappear over time, realizing the accumulation of experience and extracting useful information from experience. The context control unit receives electrical signals reflecting context information and determines whether the context or task has changed. If so, the analog signal is used to reset the cumulative integral value, thereby resetting the top-down attention bias signal to 0, indicating that the unbiased attention signal is output to the central executive control module.
[0069] This embodiment will be applied to the robot cognitive control system. For each decision-making choice of a task goal, it is necessary to include a bottom-up attention pathway (including the speech loop, the visual space board, and the attention control module) and a top-down attention pathway (including the context buffer and feedback information from the central executive control module). The final cognitive decision occurs in the central executive module, and the combined effects of the quality of the encoded information, top-down, and bottom-up attention will be compared. Assuming that the potential behavioral results of the robot are binary: action 1 or action 2, the embodiment is applied to the robot cognitive control system, and the interaction between the stimulus-driven bottom-up attention and the reward-reinforced top-down attention realizes the cognitive decision control of the robot.
[0070] The description of the technical solutions related to this embodiment is applicable to the understanding of the technical solutions of the first embodiment. In addition, the technical solutions related to this embodiment are the same as those of the first embodiment and will not be described in detail here.
[0071] Embodiment 3
[0072] A robot implements the steps of a robot cognitive decision-making method as described above when executing a task objective.
[0073] The relevant technical solution is the same as the above embodiment and will not be described again here.
[0074] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A robot cognitive decision-making method, characterized in that: include: Receive the sound information and visual information corresponding to n task targets in real time, where n ≥ 1; The sound information and visual information corresponding to each task target are fused to obtain a multimodal stimulus convergence signal, and the fusion duration corresponding to different task targets is used to reflect the difference between the capture speeds of the attention driven by the incentive for each task target; the start time of each convergence signal is set as the start time of the attention driven by the incentive for the corresponding task target, and the convergence signal is input into the delay processing unit to determine the self-maintaining duration of the attention driven by the incentive for the corresponding task target, and the self-maintaining duration corresponding to different task targets is used to reflect the difference between the attention maintenance durations for each task target, so as to obtain the bottom-up incentive-driven attention signal for each task target; Receive decision reward information and situation information in real time; use the currently received decision reward information to accumulate rewards for the current cognitive decision information used to execute a certain task goal, and when the accumulated reward value reaches a threshold, form a continuous top-down reward reinforcement attention signal for the certain task goal; Determine whether the currently received situation information has changed. If so, reset the accumulated reward values corresponding to all task objectives. The bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal of each task goal are summed, and the largest sum corresponding to each task goal is mapped to obtain the current cognitive decision information used to execute one of the task goals, thereby realizing robot cognitive decision-making.
2. A robot cognitive decision-making method as claimed in claim 1, characterized in that: The implementation method of fusing the sound information and visual information corresponding to each task target to obtain a multimodal stimulus convergence signal is as follows: The sensor installed in the robot generates analog signals of the received sound information and visual information, and each analog signal is noise-filtered, synchronously integrated and encoded in turn to obtain two peak signals; a logical operation is performed between the two peak signals corresponding to each task target to obtain a multimodal stimulus convergence signal corresponding to the task target.
3. A robot cognitive decision-making method as claimed in claim 1, characterized in that: The delay processing unit is a delay processing circuit based on a memristor.
4. A robot cognitive decision-making method as claimed in claim 2, characterized in that: The way to form a continuous top-down reward-reinforced attention signal for a single-task goal is: Forming an electrical signal representing the currently received decision reward information, when no decision reward information is input, the electrical signal representing the corresponding information is set to zero; An AND operation is performed on the electrical signal representing the currently received decision reward information and the electrical signal representing the current cognitive decision information, and the AND operation results are accumulated and integrated. If the integrated signal reaches the threshold, an uninterrupted and continuous top-down reward reinforcement attention signal is formed for the task goal corresponding to the current cognitive decision information.
5. A robot cognitive decision-making system, characterized in that: include: Audiovisual reception component, incentive-driven attention control module, situational buffer module, central executive control module; The audio-visual receiving component is used to receive the sound information and visual information corresponding to n task targets in real time, where n≥1; The incentive-driven attention control module is used to fuse the sound information and visual information corresponding to each task target to obtain a multimodal stimulus convergence signal, and the fusion duration corresponding to different task targets is used to reflect the difference between the capture speeds of the incentive-driven attention of each task target; the start time of each convergence signal is set as the start time of the incentive-driven attention to the corresponding task target, and the convergence signal is input into the delay processing unit to determine the self-maintaining duration of the incentive-driven attention to the corresponding task target, and the self-maintaining duration corresponding to different task targets is used to reflect the difference between the attention maintenance durations of each task target, so as to obtain a bottom-up incentive-driven attention signal for each task target; The context buffer module includes a reward reinforcement unit and a context control unit, wherein the reward reinforcement unit is used to receive decision reward information and context information in real time; the decision reward information currently received is used to accumulate rewards for the current cognitive decision information used to execute a certain task target, and when the accumulated reward value reaches a threshold, a continuous top-down reward reinforcement attention signal for the certain task target is formed; the context control unit is used to determine whether the currently received context information has changed, and if so, reset the accumulated reward values corresponding to all task targets; The central executive control module is used to sum the bottom-up incentive-driven attention signal and the top-down reward-reinforced attention signal of each task goal, and map the largest summation result corresponding to each task goal to obtain the current cognitive decision information for executing one of the task goals.
6. The robot cognitive decision-making system according to claim 5, characterized in that: The audio-visual receiving component includes a speech loop and a visual space board, which are respectively used to receive analog signals of sound information and visual information based on sensors, and perform noise filtering, synchronous integration and encoding on each analog signal in turn to obtain a spike signal.
7. The robot cognitive decision-making system according to claim 6, characterized in that: When the attention control module fuses the sound information and visual information corresponding to each task target to obtain a multimodal stimulation convergence signal, the implementation method is: performing a logical operation between the two peak signals corresponding to each task target to obtain a multimodal stimulation convergence signal corresponding to the task target.
8. The robot cognitive decision-making system according to claim 5, characterized in that: The delay processing unit is a delay processing circuit based on a memristor.
9. The robot cognitive decision-making system according to claim 5, characterized in that: The reward reinforcement unit forms a continuous top-down reward reinforcement attention signal for a single task target in the following manner: Forming an electrical signal representing the currently received decision reward information, when no decision reward information is input, the electrical signal representing the corresponding information is set to zero; An AND operation is performed on the electrical signal representing the currently received decision reward information and the electrical signal representing the current cognitive decision information, and the AND operation results are accumulated and integrated. If the integrated signal reaches the threshold, an uninterrupted and continuous top-down reward reinforcement attention signal is formed for the task goal corresponding to the current cognitive decision information.
10. A robot, characterized in that: The steps of a robot cognitive decision-making method as described in any one of claims 1 to 4 are implemented when executing the task objectives.
Citation Information
Patent Citations
Intelligent spectrum cooperative sensing method based on reinforcement learning
CN108833040A
Intelligent decision-making method and device based on multi-modal data fusion and reinforcement learning
CN114860893A
Brain-like perception-learning-decision system and method
CN116795942A
Sensing neuron circuit based on memristor and application
CN117669676A
Multi-modal continuous learning method and device, equipment and storage medium
CN117875407A