Cognitive reconstruction method, system and equipment based on humanoid robot and medium

By collecting and analyzing multimodal information through humanoid robots, and combining cognitive reconstruction and progressive emotion mining thought chains, dynamic and personalized psychological guidance for users' emotions is achieved. This solves the problem of the inability to actively intervene in existing technologies and improves the scientificity and effectiveness of psychological counseling.

CN121479660APending Publication Date: 2026-02-06DIGITAL HUAXIA (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511622139.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies cannot achieve dynamic, personalized, and in-depth cognitive intervention of users' emotions, and lack in-depth psychological counseling processes based on psychological theories, especially when users are in a low mood, they cannot provide proactive and timely intervention.

Method used

The humanoid robot automatically collects multimodal information, including auditory and visual modal information, through sensors pre-installed on it. It analyzes the multimodal information to identify user emotions and events, trigger corresponding interaction processes, and uses cognitive reconstruction of thought chains or progressive emotion mining of thought chains for psychological guidance. It also combines large-scale language models to generate personalized dialogues.

Benefits of technology

It enables autonomous emotion intervention when users are feeling down, improves the accuracy of emotion recognition, ensures a complete closed loop from emotion recognition to cognitive reconstruction, provides personalized and scientific psychological counseling, and overcomes the shortcomings of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479660A_ABST
    Figure CN121479660A_ABST
Patent Text Reader

Abstract

The invention discloses a cognitive reconstruction method, system and equipment based on a humanoid robot and a medium, relates to the technical field of artificial intelligence, is applied to a cognitive reconstruction system pre-arranged in the humanoid robot, and comprises the following steps: automatically collecting multi-modal information of a target user, and analyzing the multi-modal information to obtain a multi-modal analysis result; the multi-modal analysis result comprises voice transcription text information, acoustic emotion information and visual emotion information; identifying the user emotion and a target event corresponding to the user emotion according to the multi-modal analysis result, and determining whether the identification is successful or not; if successful, initial cognition is determined through a cognition reconstruction thinking chain according to the user emotion and the target event, and new cognition is obtained through reconstruction according to a first target dialogue with the user; and if not, identifying the user emotion and the target event according to the second target dialogue with the user through the progressive emotion mining thinking chain, and skipping to the step of determining whether the identification is successful or not. And psychological counseling can be autonomously performed on the user to reconstruct user cognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a cognitive reconstruction method, system, device and medium based on humanoid robots. Background Technology

[0002] Currently, there is a huge demand for mental health services, but professional resources are scarce. Existing technologies mainly include: 1. Mental health applications (Apps): These offer functions such as meditation, emotion diaries, and psychological courses. Their drawbacks are that they are static, passive, and generalized, unable to provide dynamic, personalized, and in-depth cognitive intervention based on the user's specific current emotions. 2. Chatbots: These engage in dialogue based on predefined scripts or large models. Their disadvantages are twofold: firstly, they lack structured process guidance based on psychological theories (such as Cognitive Behavioral Therapy, CBT), making dialogue prone to divergence and making it difficult to guarantee therapeutic effects from emotion recognition to cognitive change; secondly, they require users to actively input text, making proactive and timely intervention impossible when users are emotionally distressed but unwilling to express themselves. 3. Humanoid robots: Currently, most are used for education, companionship, or demonstrations. Their interaction logic is simple, lacking a deeply integrated, psychologically based "brain" to execute complex psychological counseling processes.

[0003] In conclusion, how to proactively provide psychological guidance to users in order to reconstruct their cognition is an urgent problem that needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a cognitive reconstruction method, system, device, and medium based on a humanoid robot, which can autonomously provide psychological guidance to users to reconstruct their cognition. The specific solution is as follows:

[0005] In a first aspect, this application discloses a cognitive reconstruction method based on a humanoid robot, applied to a cognitive reconstruction system pre-built into the humanoid robot, comprising:

[0006] The humanoid robot automatically collects multimodal information of the target user through sensors pre-installed in it, and analyzes the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modality information and visual modality information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information;

[0007] Based on the results of the multimodal analysis, determine whether to trigger the interaction process. If so, identify the user's emotions and the target events corresponding to those emotions based on the results of the multimodal analysis, and determine whether the identification was successful.

[0008] If the recognition is successful, the cognitive reconstruction mind chain is triggered. The cognitive reconstruction mind chain determines the user's initial cognition based on the user's emotions and target events. The cognition is reconstructed based on the user's first target dialogue to obtain the new cognition and inform the user. The first target dialogue is a verification and analysis dialogue for the initial cognition.

[0009] If identification fails, a progressive emotion mining mindset is triggered. This mindset identifies the user's emotions and target events based on the second target dialogue with the user, and then proceeds to the step of determining whether the current identification was successful. The second target dialogue includes a dialogue to determine the user's emotions by asking about the user's state and a dialogue to connect emotions and events through contextual association hypotheses.

[0010] Optionally, the analysis of multimodal information to obtain multimodal analysis results includes:

[0011] The speech-transcribed text content is obtained by processing auditory modal information through an automatic speech recognition model;

[0012] Acoustic emotion information is obtained by analyzing auditory modal information through a speech emotion recognition model; the acoustic emotion information includes acoustic feature description, acoustic feature confidence, acoustic emotion description, and acoustic emotion confidence.

[0013] The speech-transcribed text content, acoustic emotion information, and visual modal information are processed simultaneously using a visual language model to obtain speech-transcribed text information, visual emotion information, and the acoustic emotion information. The speech-transcribed text information includes the semantic emotion state and semantic emotion confidence of the transcribed text. The visual emotion information includes the visual emotion state and the visual emotion confidence.

[0014] Optionally, the step of identifying user emotions and corresponding target events based on multimodal analysis results includes:

[0015] If the three emotions identified based on the speech-transcribed text information, acoustic emotion information, and visual emotion information are the same target emotion, then the target emotion is identified as the user's emotion.

[0016] If the three emotions determined based on the speech-transcribed text information, acoustic emotion information, and visual emotion information are not the same target emotion, and / or, the confidence level of any of the speech-transcribed text information, acoustic emotion information, and visual emotion information is greater than the confidence level threshold corresponding to any of the information, then no user emotion was identified.

[0017] Optionally, the user state includes the user's physiological feelings and the user's behavioral intentions.

[0018] Optionally, after reconstructing the cognition based on the first target dialogue with the user to obtain the new cognition and informing the user, the process further includes:

[0019] After users accept new knowledge, they perform embodied behaviors that reinforce the new knowledge.

[0020] Optionally, after identifying user emotions and target events through a progressive emotion mining thought chain based on the second-target dialogue with the user, the method further includes:

[0021] Users' emotions are scored based on their emotional intensity, and their embodied behaviors are adjusted based on the scores.

[0022] Optionally, determining whether to trigger the interaction process based on the multimodal analysis results includes:

[0023] Identify dangerous behavior signals based on multimodal analysis results;

[0024] If the dangerous behavior signal is detected, the crisis intervention protocol is triggered to intervene in the user's behavior, and the interaction process is triggered according to the dangerous behavior signal.

[0025] Secondly, this application discloses a cognitive reconstruction system applied to humanoid robots, comprising:

[0026] The information acquisition module is used to automatically acquire multimodal information of the target user through sensors pre-installed in the humanoid robot, and analyze the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modal information and visual modal information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information;

[0027] The recognition module is used to determine whether to trigger the interaction process based on the multimodal analysis results. If so, it identifies the user's emotions and the target events corresponding to the user's emotions based on the multimodal analysis results, and determines whether the recognition is successful.

[0028] The cognitive reconstruction module is used to trigger a cognitive reconstruction thought chain if recognition is successful. The cognitive reconstruction thought chain determines the user's initial cognition based on the user's emotions and target events. The cognition is reconstructed based on the user's first target dialogue to obtain a new cognition and then informed to the user. The first target dialogue is a verification and analysis dialogue for the initial cognition.

[0029] The emotion mining module is used to trigger a progressive emotion mining thought chain if identification fails. The progressive emotion mining thought chain identifies the user's emotions and target events based on the second target dialogue with the user, and then jumps to the step of determining whether the current identification is successful. The second target dialogue includes a dialogue to determine the user's emotions by asking the user for state information and a dialogue to connect emotions and events through contextual association assumptions.

[0030] Thirdly, this application discloses an electronic device, including:

[0031] Memory, used to store computer programs;

[0032] A processor is used to execute the computer program to implement the aforementioned cognitive reconstruction method based on a humanoid robot.

[0033] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned cognitive reconstruction method based on a humanoid robot.

[0034] As can be seen, this application automatically collects multimodal information of the target user through sensors pre-installed in the humanoid robot, and analyzes the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modality information and visual modality information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information; based on the multimodal analysis results, it determines whether to trigger the interaction process; if so, it identifies the user's emotion and the target event corresponding to the user's emotion based on the multimodal analysis results, and determines whether the identification is successful; if the identification is successful, it triggers the cognitive reconstruction thinking chain, which determines the user's initial cognition based on the user's emotion and the target event, reconstructs the cognition based on the first target dialogue with the user to obtain new cognition, and informs the user; the first target dialogue is a verification analysis dialogue of the initial cognition; if the identification fails, it triggers the progressive emotion mining thinking chain, which identifies the user's emotion and the target event based on the second target dialogue with the user, and jumps to the step of determining whether the identification is successful; the second target dialogue includes a dialogue to determine the user's emotion by asking the user's state and a dialogue to connect the emotion and the event through contextual association hypothesis. Therefore, it is evident that the humanistic robot of this application can autonomously intervene in emotions without user initiation and without missing the best intervention opportunity; this application integrates auditory modality information and visual modality information to identify emotional features and improve the accuracy of emotion recognition; this application establishes an emotion recognition process, a cognitive reconstruction thinking chain, and a progressive emotion mining thinking chain, which can realize a complete closed loop from emotion recognition to cognitive reconstruction, overcome the shortcomings of machine-generated dialogue divergence and lack of therapeutic effect, and improve scientificity and effectiveness. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0036] Figure 1This is a flowchart of a cognitive reconstruction method based on a humanoid robot disclosed in this application;

[0037] Figure 2 This is a flowchart of a specific cognitive reconstruction method based on a humanoid robot disclosed in this application;

[0038] Figure 3 This is a schematic diagram of the cognitive reconstruction device based on a humanoid robot disclosed in this application;

[0039] Figure 4 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Currently, there is a huge demand for mental health services, but professional resources are scarce. Existing technologies mainly include: 1. Mental health applications: These offer functions such as meditation, emotion diaries, and psychological courses. Their drawbacks are that they are static, passive, and generalized, unable to provide dynamic, personalized, and in-depth cognitive intervention based on the user's specific current emotions. 2. Text-based chatbots: These engage in dialogue based on predefined scripts or large models. Their disadvantages are twofold: firstly, they lack structured process guidance based on psychological theory, making dialogue prone to divergence and making it difficult to guarantee therapeutic effects from emotion recognition to cognitive change; secondly, they require users to actively input text, making proactive and timely intervention impossible when users are emotionally distressed but unwilling to express themselves. 3. Humanoid robots: Currently, most are used for education, companionship, or demonstrations. Their interaction logic is simple, lacking a deeply integrated, psychologically based "brain" to execute complex psychological counseling processes.

[0042] Therefore, this application proposes a cognitive reconstruction scheme based on a humanoid robot, which can autonomously provide psychological guidance to users in order to reconstruct their cognition.

[0043] This application discloses a cognitive reconstruction method based on a humanoid robot. See [link to relevant documentation]. Figure 1 As shown, the method includes:

[0044] Step S11: Automatically collect multimodal information of the target user through sensors pre-installed in the humanoid robot, and analyze the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modal information and visual modal information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information and visual emotion information.

[0045] In this embodiment, the sensors include visual sensors (such as cameras) and auditory sensors (such as microphone arrays) mounted on the humanoid robot. The auditory modal information includes speech, etc., and the visual modal information includes the user's facial expressions, posture, behavior, etc.

[0046] In this embodiment, the analysis of multimodal information to obtain multimodal analysis results includes: processing auditory modal information through an automatic speech recognition model to obtain speech-transcribed text content; analyzing auditory modal information through a speech emotion recognition model to obtain acoustic emotion information; the acoustic emotion information includes acoustic feature description, acoustic feature confidence, acoustic emotion description, and acoustic emotion confidence; and simultaneously processing speech-transcribed text content, acoustic emotion information, and visual modal information through a visual language model to obtain speech-transcribed text information, visual emotion information, and the acoustic emotion information; the speech-transcribed text information includes the semantic emotion state of the transcribed text and the semantic emotion confidence of the transcribed text; the visual emotion information includes the visual emotion state and the visual emotion confidence.

[0047] It should be noted that the process of obtaining speech-transcribed text content by processing auditory modal information through an automatic speech recognition model, and obtaining acoustic emotion information by analyzing auditory modal information through a speech emotion recognition model, is as follows: Inputting the original audio stream A_t, the original audio stream is first transcribed: the audio stream A_t is processed using an automatic speech recognition (ASR) model to generate transcribed text T_asr (speech-transcribed text information). Example output: "Sigh, I really don't know how to proceed with this project." Secondly, acoustic emotion analysis is performed: the acoustic features of the audio stream A_t (including but not limited to pitch, speech rate, energy, spectrum, etc.) are analyzed using a speech emotion recognition (SER) model to generate a structured emotion analysis result. This result is formatted as a plain text natural language description E_audio (acoustic emotion information). Example output: The speech features are: a distinct sigh, followed by a slow speech rate (confidence 0.92), a low pitch (confidence 0.88), and accompanied by a slight tremor (confidence 0.7). The overall acoustic features suggest an emotional tendency of frustration (0.75 confidence level) or anxiety (0.68 confidence level). Finally, the above analyses were combined: the transcribed text T_asr and the acoustic sentiment analysis description E_audio were combined to generate a unified, information-rich plain text audio summary S_audio; the fixed text format used was: User says: {T_asr}. Its speech sounds like {E_audio}. A final example of S_audio output: User says: Sigh, I really don't know how to proceed with this project. Its speech features: a distinct sigh, followed by a slow speech rate (0.92 confidence level), a low pitch (0.88 confidence level), and a slight tremor (0.7 confidence level). The overall acoustic features suggest an emotional tendency of frustration (0.75 confidence level) or anxiety (0.68 confidence level).

[0048] It should be noted that after obtaining the speech-transcribed text content, acoustic emotion information, and visual modality information through the Visual Language Model (VLM), speech-transcribed text information, visual emotion information, and the aforementioned acoustic emotion information will be obtained.

[0049] Step S12: Determine whether to trigger the interaction process based on the multimodal analysis results. If so, identify the user's emotions and the target events corresponding to the user's emotions based on the multimodal analysis results, and determine whether the identification is successful.

[0050] In this embodiment, determining whether to trigger the interaction process based on the multimodal analysis results includes: judging whether the user's emotion exceeds the emotion threshold based on the multimodal analysis results; if so, switching from the monitoring state to the interaction state; and subsequently determining the form of the interaction state based on whether the recognition is successful.

[0051] It should be noted that this application utilizes the multimodal perception capabilities of a humanoid robot to achieve proactive emotional intervention without user initiation. The system can recognize features such as facial expressions and tone of voice, and proactively provide psychological assistance when users are feeling down but unwilling to express it, seizing the optimal opportunity for intervention.

[0052] It should be noted that the user's emotions are considered to exceed the emotional threshold if the following situations occur:

[0053] 1. High-intensity signal of a single modality: The sentiment confidence of any modality exceeds its independent trigger threshold. For example, in visual sentiment analysis, the confidence of "sadness" exceeds 0.85 for 3 seconds; in acoustic sentiment analysis, the confidence of the "crying" feature exceeds 0.9.

[0054] 2. Multimodal signal synergy: Multiple modal signals occur simultaneously, forming a synergistic effect. For example, visual detection of "head drooping" (confidence > 0.7) and acoustic detection of "long sigh" (confidence > 0.6) occur simultaneously.

[0055] 3. Conflicting information across different modalities: For example, semantic and acoustic conflict signals: The ASR transcribed text content may be neutral or positive, but the acoustic features may indicate strong negative emotions. For instance, a user says "I'm fine" (T_asr), but acoustic analysis shows "voice tremor" (confidence > 0.8) and "low pitch" (confidence > 0.75). This conflict itself constitutes a strong trigger signal.

[0056] 4. Crisis Behavior Signals: Pre-defined behavioral patterns associated with high risk are detected. For example: severe body tremors accompanied by sobbing are detected; prolonged periods of stillness and silence (e.g., more than 30 seconds) are detected (visual and audio energy are both below the threshold).

[0057] It should be noted that the conditions for judging whether a user's emotions exceed the emotional threshold are not limited to the four situations mentioned above. Conditions for judgment can be added or removed according to the specific circumstances.

[0058] In this embodiment, determining whether to trigger the interaction process based on the multimodal analysis results includes: identifying dangerous behavior signals based on the multimodal analysis results; if the dangerous behavior signals are identified, triggering a crisis intervention protocol to intervene in the user's behavior, and triggering the interaction process based on the dangerous behavior signals.

[0059] It should be noted that the aforementioned dangerous behavior signals may include severe trembling accompanied by sobbing, prolonged stillness or silence, etc. Specific signals can be set according to the actual situation. Interventions to user behavior may include hugging the user, stopping the user's current behavior, etc. These interventions are distinct from embodied behaviors and aim to prevent irrational behavior before reconstructing the user's cognition. The crisis intervention protocol specifies which dangerous behavior signals correspond to which intervention behaviors.

[0060] It should be noted that the initial emotion report when identifying emotions is as follows:

[0061] {

[0062] "analysis": {

[0063] "emotional_state": "[Dominant sentiment based on multimodal information analysis]",

[0064] "confidence": [0.0-1.0],

[0065] "visual_emotion_confidence": [0.0-1.0],

[0066] "acoustic_emotion_confidence": [0.0-1.0],

[0067] "trigger_reason": "[Briefly describe the reason that triggered this analysis, such as: visual 'sadness' confidence > 0.85 for 3 seconds, or acoustic 'crying' confidence > 0.9, or visual 'head down' + acoustic 'long sigh' synergy, or semantic-acoustic conflict (the user says 'I'm fine' but the acoustic system detects trembling), etc.]",

[0068] "user_need": "[Infer the user's potential needs, such as 'need for comfort,' 'need to confide in someone,' or 'need to solve a problem']",

[0069] "current_stage": "[Current dialogue stage: 'Initial Contact' / 'Emotion Recognition' / 'Emotion Exploration' / 'Cognitive Exploration' / 'Cognitive Restructuring' / 'Process Advancement' / 'Exception Handling']"

[0070] }

[0071] }

[0072] In this implementation, to ensure the smooth progress of subsequent interactions, large-scale language models such as GPT-4 and Claude will be provided in the cloud, as well as large model interface modules (LLM API) for calling and management. It should be noted that within the structured framework, the contextual understanding capabilities of large-scale language models can be used to dynamically generate personalized dialogue content.

[0073] Step S13: If the recognition is successful, the cognitive reconstruction thinking chain is triggered. The cognitive reconstruction thinking chain determines the user's initial cognition based on the user's emotions and target events. The cognition is reconstructed based on the user's first target dialogue to obtain new cognition and then informed to the user. The first target dialogue is a verification and analysis dialogue for the initial cognition.

[0074] In this embodiment, the user state includes the user's physiological feelings and the user's behavioral intentions.

[0075] In this embodiment, after the emotion is identified, the process automatically enters the cognitive restructuring thought chain.

[0076] In one specific embodiment, after entering the cognitive restructuring thought chain, the robot outputs the following through voice and display: 1. Automated thought extraction: "When you feel wronged, what thoughts flash through your mind? Is it 'This is unfair to me'?"; 2. Evidence verification and cognitive challenge: "What evidence supports the idea 'This is unfair to me'?", "What evidence or different perspectives oppose this idea?", "What are the worst, best, and most realistic outcomes?"; 3. Cognitive restructuring: Based on the dialogue, summarize and generate a new, more adaptive cognitive perspective ("So can we adjust our thought from 'This is absolutely unfair' to 'His way makes me feel wronged, but maybe he didn't mean any harm, I can try to communicate and express my feelings'?"), and submit it for user confirmation.

[0077] It should be noted that if the user does not accept the new understanding obtained from the reconstruction, the cognitive reconstruction thought chain will be retried, and the cognitive reconstruction process will be carried out again based on the previous dialogue and multimodal analysis results until the user accepts the new understanding obtained from the reconstruction.

[0078] It should be noted that if, for the same multimodal analysis result, the number of new cognitions that users do not accept exceeds the cognitive threshold, then consider jumping to the step of collecting multimodal information of the target user or analyzing multimodal information to obtain multimodal analysis results, and start cognitive reconstruction again.

[0079] In this embodiment, after reconstructing cognition based on the first target dialogue with the user to obtain new cognition and informing the user, the process further includes: after the user accepts the new cognition, performing embodied behaviors to reinforce the new cognition. The embodied behaviors include actions, facial expressions, and language. In one specific embodiment, embodied reinforcement occurs when the user completes cognitive reconstruction; the robot performs preset empathic actions (such as slow nodding or making an affirmative gesture) accompanied by an encouraging tone of voice to reinforce the new cognition.

[0080] It should be noted that after embodying the action, you can also give users behavioral suggestions. For example, based on the new understanding, you can provide a small and specific action suggestion ("Now, let's take three deep breaths together to calm down, okay?"), and guide the user to complete it together to calm the user's emotions.

[0081] It should be noted that this application utilizes the physical form of a humanoid robot to enhance empathy and trust through non-verbal information such as eye contact, gestures, and body language. It transcends the limitations of virtual interaction, providing a more natural and profound user experience, and achieving the true role of an "emotional coach."

[0082] Step S14: If identification fails, the progressive emotion mining mind chain is triggered. The progressive emotion mining mind chain identifies the user's emotions and target events based on the second target dialogue with the user, and then jumps to the step of determining whether the current identification is successful. The second target dialogue includes a dialogue to determine the user's emotions by asking about the user's state and a dialogue to connect emotions and events through contextual association hypothesis.

[0083] In this embodiment, a progressive emotion mining thought chain is triggered when the user is unable to express their emotions.

[0084] In one specific embodiment, after initiating the progressive emotion mining thought chain, the robot outputs guiding questions through voice and display: 1. Anchoring physiological feelings: "You look like you need to rest. Can you feel any discomfort in any part of your body? For example, a tightness in your chest?" 2. Guiding behavioral impulses: "If there is an action that can express this feeling, what would it be? Do you want to hide or shout?" 3. Providing dynamic emotion options: Based on responses a and b, the large model dynamically generates a set of the most relevant emotion words (such as "anger, resentment, helplessness") for the user to choose from. 4. Contextual association hypothesis: Proposing possible contextual hypotheses ("Is this feeling of [resentment] related to the previous conversation?") to help the user connect emotions with events. 5. Emotion confirmation and quantification: Scoring the intensity of identified emotions, specifically based on a fusion score of acoustic, visual, and transcribed text modalities. Based on the score, the large model can adjust the language of empathic expression and whether to prompt the user to request external help.

[0085] It should be noted that the above intensity score can influence embodied behavior. Specifically, after identifying user emotions and target events through the progressive emotion mining mind chain based on the second target dialogue with the user, it also includes: scoring the user's emotions for intensity and adjusting embodied behavior based on the score results.

[0086] It should be noted that the emotion mining thought chain engine includes a prompt word engineering library, which stores prompt templates carefully designed for the above process, for interaction with large models.

[0087] In summary, this application, based on psychological theories such as cognitive behavioral therapy, establishes a standardized thought process (including a progressive emotion-mining thought process and a cognitive restructuring thought process) to ensure a complete therapeutic loop from emotion recognition to cognitive restructuring. It overcomes the shortcomings of general chatbots, such as divergent dialogue and lack of therapeutic effect, ensuring the scientific validity and effectiveness of the intervention. Furthermore, this application provides in-depth "one-on-one" coaching tailored to the user's real-time emotional state and specific cognitive experiences, achieving a unity of standardized process and personalized content.

[0088] It should be noted that after processing multimodal information using a visual language model, this application can design a carefully constructed prompt word template to integrate S_audio and the image frame It into a unified question-answering task. An example of the prompt word is shown below:

[0089] "You are a professional psychological counseling robot, observing users through a camera to provide timely emotional support. You must integrate visual and audio information, conduct in-depth analysis, and determine the best intervention strategy."

[0090] [Audio Transcription and Analysis]

[0091] {S_audio};

[0092] Please analyze the image.

[0093] <image: I_t> ;

[0094] # Workflow

[0095] 1. Multimodal Perception Analysis: Analyze the provided visual snapshots and user statements to determine the user's emotional state, behavioral intentions, and potential needs.

[0096] 2. Process Status Judgment: Based on the current progress of the dialogue, determine which step in the [Ordered Strategy Process] should be executed.

[0097] 3. Decision-making: Select the steps to be executed from the [Ordered Strategy Process] (Progressive Emotion Mindset Chain or Cognitive Restructuring Chain).

[0098] 4. Generation: Based on the selected strategy, generate a natural and empathetic guided dialogue text to respond to the user or perform a specific action.

[0099] Complete the cognitive reconstruction process based on the above prompts.

[0100] As can be seen, this application automatically collects multimodal information of the target user through sensors pre-installed in the humanoid robot, and analyzes the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modality information and visual modality information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information; based on the multimodal analysis results, it determines whether to trigger the interaction process; if so, it identifies the user's emotion and the target event corresponding to the user's emotion based on the multimodal analysis results, and determines whether the identification is successful; if the identification is successful, it triggers the cognitive reconstruction thinking chain, which determines the user's initial cognition based on the user's emotion and the target event, reconstructs the cognition based on the first target dialogue with the user to obtain new cognition, and informs the user; the first target dialogue is a verification analysis dialogue of the initial cognition; if the identification fails, it triggers the progressive emotion mining thinking chain, which identifies the user's emotion and the target event based on the second target dialogue with the user, and jumps to the step of determining whether the identification is successful; the second target dialogue includes a dialogue to determine the user's emotion by asking the user's state and a dialogue to connect the emotion and the event through contextual association hypothesis. Therefore, it is evident that the humanistic robot of this application can autonomously intervene in emotions without user initiation and without missing the best intervention opportunity; this application integrates auditory modality information and visual modality information to identify emotional features and improve the accuracy of emotion recognition; this application establishes an emotion recognition process, a cognitive reconstruction thinking chain, and a progressive emotion mining thinking chain, which can realize a complete closed loop from emotion recognition to cognitive reconstruction, overcome the shortcomings of machine-generated dialogue divergence and lack of therapeutic effect, and improve scientificity and effectiveness.

[0101] This application discloses a specific cognitive reconstruction method based on a humanoid robot. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. See also... Figure 2 As shown, it specifically includes:

[0102] Step S21: If the three emotions determined based on the speech-transcribed text information, acoustic emotion information, and visual emotion information are the same target emotion, then the target emotion is identified as the user's emotion.

[0103] In this embodiment, if a multimodal signal synergy occurs, meaning multiple modal signals appear simultaneously and create a synergistic effect, and if the three emotions identified by the speech-transcribed text information, acoustic emotion information, and visual emotion information are the same target emotion, then that target emotion is identified as the user's emotion. If one or two of them represent the same target emotion, and the others represent other emotions, then the correct user emotion has not been identified.

[0104] Step S22: If the three emotions determined based on the speech-transcribed text information, acoustic emotion information, and visual emotion information are not the same target emotion, and / or, the confidence level of any one of the speech-transcribed text information, acoustic emotion information, and visual emotion information is greater than the confidence level threshold corresponding to any one of the information, then no user emotion is identified.

[0105] In this embodiment, if one or two of the speech-transcribed text information, acoustic emotion information, and visual emotion information represent the same target emotion, and the others represent other emotions, then the correct user emotion has not been identified. Specifically, if there is a conflict between information of different modalities, such as semantic and acoustic conflict signals, that is, the ASR-transcribed text content is neutral or positive, but the acoustic features show a strong negative emotion, then it is considered that the correct user emotion has not been identified.

[0106] In this embodiment, if the confidence level of any of the speech-transcribed text information, acoustic emotion information, and visual emotion information is greater than the confidence level threshold corresponding to any of the information, then the user's emotion is not identified. Specifically, if there is a high-intensity signal in a single modality, that is, if the emotion confidence level of any modality exceeds its independent trigger threshold, it is considered that the correct user emotion has not been identified. For example, in visual emotion analysis, the confidence level of "sadness" exceeds 0.85 for 3 seconds, or in acoustic emotion analysis, the confidence level of the "crying" feature exceeds 0.9.

[0107] Therefore, if the three emotions determined by this application based on speech-transcribed text information, acoustic emotion information, and visual emotion information are the same target emotion, then this target emotion is identified as the user's emotion. If the three emotions determined by these three information are not the same target emotion, and / or, the confidence level of any one of the three information is greater than the confidence level threshold corresponding to that information, then the user's emotion is not identified. Thus, this application determines whether a user's emotion is identified by considering the relationship between the three emotions determined by speech-transcribed text information, acoustic emotion information, and visual emotion information, as well as whether any special cases exist for any emotion, rather than directly identifying an emotion determined by a single modality as the user's emotion. This improves the accuracy of user emotion recognition.

[0108] Accordingly, this application also discloses a cognitive reconstruction device, see [link to relevant documentation]. Figure 3 As shown, the device, applied to a humanoid robot, includes:

[0109] The information acquisition module 11 is used to automatically acquire multimodal information of the target user through sensors pre-installed in the humanoid robot, and analyze the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modal information and visual modal information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information;

[0110] The recognition module 12 is used to determine whether to trigger the interaction process based on the multimodal analysis results. If so, it identifies the user's emotions and the target events corresponding to the user's emotions based on the multimodal analysis results, and determines whether the recognition is successful.

[0111] The cognitive reconstruction module 13 is used to trigger the cognitive reconstruction thinking chain if the recognition is successful. The cognitive reconstruction thinking chain determines the user's initial cognition based on the user's emotions and target events. The cognition is reconstructed based on the user's first target dialogue to obtain new cognition and inform the user. The first target dialogue is a verification and analysis dialogue of the initial cognition.

[0112] The emotion mining module 14 is used to trigger a progressive emotion mining thought chain if the identification fails. The progressive emotion mining thought chain identifies the user's emotions and target events based on the second target dialogue with the user, and jumps to the step of determining whether the current identification is successful. The second target dialogue includes a dialogue to determine the user's emotions by asking the user for state information and a dialogue to connect emotions and events through contextual association hypothesis.

[0113] The more specific working process of each of the above modules can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0114] As can be seen, this application automatically collects multimodal information of the target user through sensors pre-installed in the humanoid robot, and analyzes the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modality information and visual modality information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information; based on the multimodal analysis results, it determines whether to trigger the interaction process; if so, it identifies the user's emotion and the target event corresponding to the user's emotion based on the multimodal analysis results, and determines whether the identification is successful; if the identification is successful, it triggers the cognitive reconstruction thinking chain, which determines the user's initial cognition based on the user's emotion and the target event, reconstructs the cognition based on the first target dialogue with the user to obtain new cognition, and informs the user; the first target dialogue is a verification analysis dialogue of the initial cognition; if the identification fails, it triggers the progressive emotion mining thinking chain, which identifies the user's emotion and the target event based on the second target dialogue with the user, and jumps to the step of determining whether the identification is successful; the second target dialogue includes a dialogue to determine the user's emotion by asking the user's state and a dialogue to connect the emotion and the event through contextual association hypothesis. Therefore, it is evident that the humanistic robot of this application can autonomously intervene in emotions without user initiation and without missing the best intervention opportunity; this application integrates auditory modality information and visual modality information to identify emotional features and improve the accuracy of emotion recognition; this application establishes an emotion recognition process, a cognitive reconstruction thinking chain, and a progressive emotion mining thinking chain, which can realize a complete closed loop from emotion recognition to cognitive reconstruction, overcome the shortcomings of machine-generated dialogue divergence and lack of therapeutic effect, and improve scientificity and effectiveness.

[0115] Furthermore, embodiments of this application also provide an electronic device. Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0116] Figure 4 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the cognitive reconstruction method based on a humanoid robot disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0117] In this embodiment, the power supply 26 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 24 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0118] Furthermore, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored thereon can include computer programs 221, and the storage method can be temporary storage or permanent storage. The computer programs 221 may include, in addition to computer programs capable of performing the humanoid robot-based cognitive reconstruction method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, computer programs capable of performing other specific tasks.

[0119] Furthermore, embodiments of this application also disclose a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned cognitive reconstruction method based on a humanoid robot.

[0120] The specific steps of this method can be found in the relevant content disclosed in the foregoing embodiments, and will not be repeated here.

[0121] The various embodiments in this application are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts between the various embodiments, refer to each other. As for the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to in the method section.

[0122] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0123] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0124] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0125] The foregoing has provided a detailed description of a cognitive reconstruction method, system, device, and storage medium based on a humanoid robot. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A cognitive reconstruction method based on a humanoid robot, characterized in that, The cognitive reconstruction system pre-built into the humanoid robot includes: The humanoid robot automatically collects multimodal information of the target user through sensors pre-installed in it, and analyzes the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modality information and visual modality information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information; Based on the results of the multimodal analysis, determine whether to trigger the interaction process. If so, identify the user's emotions and the target events corresponding to those emotions based on the results of the multimodal analysis, and determine whether the identification was successful. If the recognition is successful, the cognitive reconstruction mind chain is triggered. The cognitive reconstruction mind chain determines the user's initial cognition based on the user's emotions and target events. The cognition is reconstructed based on the user's first target dialogue to obtain the new cognition and inform the user. The first target dialogue is a verification and analysis dialogue for the initial cognition. If identification fails, a progressive emotion mining mindset is triggered. This mindset identifies the user's emotions and target events based on the second target dialogue with the user, and then proceeds to the step of determining whether the current identification was successful. The second target dialogue includes a dialogue to determine the user's emotions by asking about the user's state and a dialogue to connect emotions and events through contextual association hypotheses.

2. The cognitive reconstruction method based on a humanoid robot according to claim 1, characterized in that, The analysis of multimodal information yields multimodal analysis results, including: The speech-transcribed text content is obtained by processing auditory modal information through an automatic speech recognition model; Acoustic emotion information is obtained by analyzing auditory modal information through a speech emotion recognition model; the acoustic emotion information includes acoustic feature description, acoustic feature confidence, acoustic emotion description, and acoustic emotion confidence. The speech-transcribed text content, acoustic emotion information, and visual modal information are processed simultaneously using a visual language model to obtain speech-transcribed text information, visual emotion information, and the acoustic emotion information. The speech-transcribed text information includes the semantic emotion state and semantic emotion confidence of the transcribed text. The visual emotion information includes the visual emotion state and the visual emotion confidence.

3. The cognitive reconstruction method based on a humanoid robot according to claim 2, characterized in that, The process of identifying user emotions and corresponding target events based on multimodal analysis results includes: If the three emotions identified based on the speech-transcribed text information, acoustic emotion information, and visual emotion information are the same target emotion, then the target emotion is identified as the user's emotion. If the three emotions determined based on the speech-transcribed text information, acoustic emotion information, and visual emotion information are not the same target emotion, and / or, the confidence level of any of the speech-transcribed text information, acoustic emotion information, and visual emotion information is greater than the confidence level threshold corresponding to any of the information, then no user emotion was identified.

4. The cognitive reconstruction method based on a humanoid robot according to claim 1, characterized in that, The user state includes the user's physiological feelings and the user's behavioral intentions.

5. The cognitive reconstruction method based on a humanoid robot according to claim 1, characterized in that, After reconstructing the cognition based on the initial target dialogue with the user to obtain new cognition and informing the user, the process further includes: After users accept new knowledge, they perform embodied behaviors that reinforce the new knowledge.

6. The cognitive reconstruction method based on a humanoid robot according to claim 5, characterized in that, After identifying user emotions and target events through a progressive emotion mining mindset based on the second-target dialogue with the user, the process also includes: Users' emotions are scored based on their emotional intensity, and their embodied behaviors are adjusted based on the scores.

7. The cognitive reconstruction method based on a humanoid robot according to any one of claims 1 to 6, characterized in that, The step of determining whether to trigger the interaction process based on the multimodal analysis results includes: Identify dangerous behavior signals based on multimodal analysis results; If the dangerous behavior signal is detected, the crisis intervention protocol is triggered to intervene in the user's behavior, and the interaction process is triggered according to the dangerous behavior signal.

8. A cognitive reconstruction system, characterized in that, Applications in humanoid robots, including: The information acquisition module is used to automatically acquire multimodal information of the target user through sensors pre-installed in the humanoid robot, and analyze the multimodal information to obtain multimodal analysis results; the multimodal information includes auditory modal information and visual modal information; the multimodal analysis results include speech-transcribed text information, acoustic emotion information, and visual emotion information; The recognition module is used to determine whether to trigger the interaction process based on the multimodal analysis results. If so, it identifies the user's emotions and the target events corresponding to the user's emotions based on the multimodal analysis results, and determines whether the recognition is successful. The cognitive reconstruction module is used to trigger a cognitive reconstruction thought chain if recognition is successful. The cognitive reconstruction thought chain determines the user's initial cognition based on the user's emotions and target events. The cognition is reconstructed based on the user's first target dialogue to obtain a new cognition and then informed to the user. The first target dialogue is a verification and analysis dialogue for the initial cognition. The emotion mining module is used to trigger a progressive emotion mining thought chain if identification fails. The progressive emotion mining thought chain identifies the user's emotions and target events based on the second target dialogue with the user, and then jumps to the step of determining whether the current identification is successful. The second target dialogue includes a dialogue to determine the user's emotions by asking the user for state information and a dialogue to connect emotions and events through contextual association assumptions.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the cognitive reconstruction method based on a humanoid robot as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the cognitive reconstruction method based on a humanoid robot as described in any one of claims 1 to 7.