Collaborative interaction system of companion robot based on emotional state portrait and hierarchical memory
By using a collaborative interaction system that combines emotional state profiling with hierarchical memory, the problem of shallow emotional state understanding in elderly companion robots has been solved. This system enables personalized and continuous responses and safe and reliable interaction, thereby enhancing the emotional empathy capabilities of companion robots.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies have a shallow understanding of emotional states in elderly companion robots, making it difficult to achieve continuous empathy. Furthermore, they lack joint modeling of speech rate, pauses, rhythm, historical emotional baseline, and long-term user personality preferences, leading to response mismatch and memory drift.
A collaborative interaction system based on emotional state profiling and hierarchical memory is adopted, including an entry layer, a stable context base layer, a semantic and emotional state layer, a security control layer, a memory layer, a retrieval enhancement layer, and a strategy planning layer. Through real-time voice processing, historical information integration, intent recognition, emotion fusion estimation, and security gating, personalized and continuous response strategies are generated.
It achieves continuous understanding and personalized response to emotional states in elderly companion robots, avoiding response mismatch and memory drift, and providing low latency, continuous companionship and safe and reliable interaction capabilities.
Smart Images

Figure CN122116884A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and affective computing, specifically to a collaborative interaction system for companion robots based on emotional state profiling and hierarchical memory. Background Technology
[0002] Companion robots are intelligent service robots that integrate emotional interaction, safety monitoring, health management, and life assistance. Their core function is to provide human-like companionship through artificial intelligence and multimodal perception technology, meeting the emotional comfort and care needs of the elderly, children, and people living alone. Taking elderly companion robots as an example, their application scenarios focus on providing elderly users with continuous emotional companionship, natural multi-turn dialogue, personalized interactive execution, and safe and reliable interaction. This mainly relies on key technologies such as real-time voice interaction, long-term memory management, retrieval-enhanced generation (RAG), affective computing, and embodied feedback collaborative control. The rationality of emotion-driven memory retrieval directly affects the coherence and personalization level of the companionship interaction.
[0003] To address the emotion-driven memory retrieval needs of companion robots for the elderly, existing technologies can be broadly categorized into two types based on their implementation paths: one is emotion response methods based on multimodal emotion recognition, which classify emotions, estimate emotion intensity, or map emotion labels to information such as speech, text, facial expressions, or gestures, and adjust the response style of the dialogue system accordingly; the other is historical retrieval methods based on cross-session long-term memory modeling, which store, encode, and retrieve historical dialogues, user preferences, or key events, recalling relevant content in subsequent interactions to assist in the generation of the current round. In terms of specific algorithm implementation, existing technologies generally adopt a loosely connected modular processing approach, specifically manifested as follows: (1) Treat emotion recognition as an independent pre-module and select the response strategy based only on the results of a single round of recognition; (2) Long-term memory or retrieval enhancement generation (RAG) is used as a single retrieval module, and the recall results are directly concatenated to the current round of dialogue generation context; (3) Language responses and embodied actions are planned and executed independently by different control links; (4) When there is ambiguity, uncertainty or high risk in the interactive content, there is usually a lack of special clarification questions, conservative downgrade and action inhibition mechanisms.
[0004] Based on the implementation methods of the existing technologies mentioned above, it can be seen that the current related technologies are still in a state where emotional understanding, memory retrieval and embodied feedback are separated. The understanding of emotional state is still relatively shallow and it is difficult to support continuous empathy in the scenario of companion robots, especially elderly companion robots. Summary of the Invention
[0005] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a collaborative interaction system for companion robots based on emotional state profiling and hierarchical memory, which solves the technical problem that existing technologies have a superficial understanding of emotional states.
[0006] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a collaborative interaction system for companion robots based on emotional state profiling and hierarchical memory, comprising: The entry layer is used to acquire and process real-time speech streams and output the speech observation results of the current round. The stable context base layer is used to integrate verified information from the historical rounds and output a stable context base, a historical companion profile, and a long-term preference profile. The semantic and sentiment state layer is used to complete intent recognition, entity extraction, request signal determination and sentiment fusion estimation based on the current round of speech observation results, stable context base and historical sentiment baseline, and outputs the current round interaction representation including the current round semantic state and the current round sentiment state profile; The security control layer is used to calculate and output the gating state based on the current round of interaction representation, environmental constraint state, user feedback and memory data; The memory layer is used to receive the output information from the above layers, process the information in this round, and output and store the memory data. The retrieval enhancement layer is used to construct joint queries based on the current round of interaction representation, gating state, long-term preference profile, and stable context base, process the memory data, and output a patch set; The strategy planning layer is used to generate and output a set of strategy tokens and a response skeleton based on the current emotional state profile, patch set, historical companionship profile and gating status. The output layer is used to expand the response skeleton into response content and output it under the constraints of the gating state, based on the policy token set, patch set, and current round sentiment state profile.
[0007] Preferably, the processing of the real-time voice stream includes: The real-time speech stream is subjected to noise reduction, echo suppression, and endpoint preprocessing based on speech endpoint detection and dynamic silence threshold to obtain initial speech data; The initial speech data is converted using a large-model-based streaming speech-to-text model to obtain the conversion result, which includes incremental transcription result, silence boundary and semantic integrity. Feature extraction is performed on the conversion results to obtain the current round speech observation results, which include the current round effective speech segments, transcription results, word-level time alignment markers, and pauses, speech rate, energy changes, and prosodic statistics.
[0008] Preferably, the memory layer is a hierarchical memory layer, which outputs and stores memory data in layers according to instant buffer memory, plot memory, preference profile memory and security boundary memory.
[0009] Preferably, the current round of emotional state profile includes: in, This represents the current emotional state profile; Indicates the emotion category; Indicates the intensity of emotion; This indicates the trend of change compared to the user's historical baseline; This indicates the risk level inferred from both sentiment and semantics. This represents the current round of emotional fusion. For the semantic state of the current round Extracted semantic cue vectors; Indicates the semantic state of the current round. Extracted intent signal vector; The acoustic feature vector is composed of the current round's prosody, pauses, speech rate, and energy changes; This serves as a dynamic personality baseline or stable emotional anchor for users participating in the current round of calculations. This is a vector of recent interaction states that participates in the current round of calculation, used to reflect the emotional fluctuations and interaction rhythm of the last few rounds; These are preset, learnable, or configurable mapping parameters; This is the normalized mapping function; The baseline for the emotional intensity of the recent stable cycles participating in the current round of calculation; The strength of error correction, rejection, or negative feedback from the most recent user participating in the current round of calculation; For category normalization function; subscript Indicates the current round.
[0010] Preferably, the gating state includes: in, Indicates the gating status of the current wheel; This indicates a normal state of companionship; This indicates a need for clarification. This indicates that only conservative responses are allowed. Indicates a state of motion suppression or output compression; The gated state is equipped with a security gating and risk degradation mechanism, as detailed below: In the formula, This represents the total risk score for the current round of gate control. The coefficient is a non-negative polymerization factor; The semantic understanding uncertainty of the current round; The degree of conflict between the retrieval patch participating in the current round of calculation and the long-term preference profile or confirmed facts; The environmental risk level used in the current round of calculations is used as an external constraint. The strength of error correction, rejection, or negative feedback from the most recent user participating in the current round of calculation; The threshold values are arranged from low to high. This is the initial gating state obtained based on the immediate risk evidence of the current round; This is the gated state transition function; This is the previous gating state; The current round of clarification results reflects whether disambiguation has been completed, whether high-risk semantics have been identified, or whether ambiguity still exists. This represents the current environmental constraint state; subscript Indicates the current round, This refers to the previous round.
[0011] Preferably, the security control layer is further configured to perform continuous learning and updates based on long-term write-back rules, outputting long-term write-back determination results and an updated stable context base, wherein the long-term write-back rules include: Only information that appears repeatedly in multiple rounds and meets the confidence level is written into the long-term preference profile; Recent events are prioritized for inclusion in narrative memory, rather than being directly elevated to long-term preferences; Negative feedback and interaction boundaries should be written into the safety boundary memory first. Implement time decay, conflict comparison, and reconfirmation mechanisms for information already written into long-term preference profiles.
[0012] Preferably, the companion robot collaborative interaction system based on emotional state profile and hierarchical memory further includes an embodied execution layer. The embodied execution layer is used to map the strategy token set to a whitelist action set according to the strategy token set, the response content, the current round of emotional state profile, the gating state, and the environmental constraint state. It also combines the voice start time and rhythm constraints in the response content to perform action feasibility screening, action scoring, and temporal collaborative arrangement, and outputs embodied action tokens and corresponding facial expressions, postures, turning, approaching, or avoiding feedback.
[0013] Secondly, the present invention provides a collaborative interaction method for companion robots based on emotional state profiling and hierarchical memory, which performs the following operations based on the aforementioned collaborative interaction system for companion robots based on emotional state profiling and hierarchical memory: The real-time speech stream is acquired and processed through the entry layer, and the speech observation results of the current round are output. By integrating verified information from the historical rounds through a stable context base layer, the stable context base, historical companion profile, and long-term preference profile are output. Based on the current round of speech observations, a stable contextual base, and a historical sentiment baseline, the semantic and sentiment state layer completes intent recognition, entity extraction, request signal determination, and sentiment fusion estimation, and outputs a current round interaction representation that includes the current round semantic state and the current round sentiment state profile. The security control layer calculates and outputs the gating state based on the current round of interaction representation, environmental constraint state, user feedback, and memory data. The memory layer receives the output information from the above layers, processes the information in this round, and outputs and stores the memory data. By using the retrieval enhancement layer to construct a joint query based on the current round of interaction representation, gating state, long-term preference profile, and stable context base, the memory data is processed and a patch set is output. The strategy planning layer generates and outputs a set of strategy tokens and a response skeleton based on the current emotional state profile, patch set, historical companionship profile, and gating status. The output layer expands the response skeleton into response content and outputs it under the constraints of gating state, based on the policy token set, patch set, and current emotional state profile.
[0014] Thirdly, the present invention provides a computer-readable storage medium storing a computer program for collaborative interaction of a companion robot based on emotional state profiling and hierarchical memory, wherein the computer program causes a computer to execute the collaborative interaction method of the companion robot based on emotional state profiling and hierarchical memory as described above.
[0015] Fourthly, the present invention provides an electronic device, comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing collaborative interaction of companion robots based on emotional state profiling and hierarchical memory as described above.
[0016] (III) Beneficial Effects This invention provides a collaborative interaction system and method for companion robots based on emotional state profiling and hierarchical memory. Compared with existing technologies, it has the following advantages: This application uses a stable context base layer to receive and process verified information from previous rounds, outputting a stable context base, a historical companionship profile, and a long-term preference profile. This overcomes the shortcomings of existing technologies that lack historical emotional baselines and long-term personality preference modeling, providing the system with complete user historical trajectory data and avoiding the isolation of single-round dialogues. The semantic and emotional state layer takes over the current round's voice observation results, the stable context base, and the historical emotional baseline, completing intent recognition, entity extraction, and emotional fusion estimation, outputting the current round's semantic state and current round's emotional state profile. This achieves multi-dimensional joint modeling, accurately distinguishing different emotional states such as short-term complaints and persistent low moods. The other layers work together, with the memory layer accumulating multi-round information, and the retrieval enhancement layer and security control layer forming a closed loop, dynamically generating adaptive responses, effectively tracking cross-round emotional trends, and providing coherent and accurate empathetic responses for scenarios such as elderly companionship. This effectively solves the technical problem of existing technologies having a shallow understanding of emotional states, achieving continuous empathy. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a framework diagram of a companion robot collaborative interaction system based on emotional state profiling and hierarchical memory, according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating the clarification gate control in the ambiguous physical discomfort scenario in specific scenario 3. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] This application provides a collaborative interaction system and method for companion robots based on emotional state profiling and hierarchical memory. It solves the technical problem that the existing technology has a shallow understanding of emotional states and provides a unified engineering approach that is particularly suitable for companion robot scenarios. This approach can retain the convenience and low latency response of real-time voice interaction, while also taking into account the requirements of continuous emotional interaction, personalized memory management, enhanced retrieval accuracy, and safe and reliable execution.
[0021] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows: Existing companion robots mainly suffer from the following shortcomings: 1. Most existing emotional companionship solutions rely primarily on the current round of semantics or single-emotion recognition results, often lacking joint modeling of speech rate, pauses, rhythm, historical emotional baseline, user's long-term personality preferences, and cross-round emotional trends. This results in the system only being able to identify "whether this sentence sounds negative," but struggling to distinguish between different states such as "short-term complaints," "persistent low mood," "refusal to communicate," and "risk statements requiring clarification," easily leading to issues like overly strong, underlying, or mismatched responses in elderly companionship scenarios.
[0022] 2. While existing long-term conversational memory solutions are beginning to emphasize cross-conversation information retention, they still often store recent user statements, long-term preferences, important events, and sensitive boundaries in a mixed manner, or simply splice together simple summaries. For elderly companion robots, the lack of different levels of division of labor, such as "instant buffering," "episodic memory," "semantic profiling," and "safety boundary memory," can lead to momentary slips of the tongue, vague expressions, or unconfirmed facts being mistakenly written into the long-term profile, resulting in memory drift and relationship distortion.
[0023] 3. In elderly companionship scenarios, ambiguous expressions, descriptions of physical discomfort, refusal to communicate, correction of misunderstandings, and questions about sensitive issues frequently occur. If the system does not design "uncertainty assessment," "clarifying questions," "conservative responses," "motor inhibition," and "boundary memory updates after erroneous feedback" as independent links, it is prone to outputting inappropriate responses under incorrect identification, incorrect retrieval, or sensitive topics, affecting safety and long-term acceptability.
[0024] To address the aforementioned issues, this invention proposes a collaborative interaction system and method for companion robots based on emotional state profiling and hierarchical memory. While retaining the system boundaries of real-time voice input, contextual memory / RAG enhancement, and text / voice response output, the internal mechanisms are reconstructed to simultaneously possess low latency, continuous companionship, personalized memory, emotional consistency, and secure and reliable execution capabilities.
[0025] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0026] This invention provides a collaborative interaction system for companion robots based on emotional state profiling and hierarchical memory, comprising: The entry layer is used to acquire and process real-time speech streams and output the speech observation results of the current round. The stable context base layer is used to integrate verified information from the historical rounds and output a stable context base, a historical companion profile, and a long-term preference profile. The semantic and sentiment state layer is used to complete intent recognition, entity extraction, request signal determination and sentiment fusion estimation based on the current round of speech observation results, stable context base and historical sentiment baseline, and outputs the current round interaction representation including the current round semantic state and the current round sentiment state profile; The security control layer is used to calculate and output the gating state based on the current round of interaction representation, environmental constraint state, user feedback and memory data; The memory layer is used to receive output information from the entry layer, the stable context base layer, the semantic and emotional state layer, and the security control layer, process the information in this round, and output and store the memory data. The retrieval enhancement layer is used to construct joint queries based on the current round of interaction representation, gating state, long-term preference profile, and stable context base, process the memory data, and output a patch set; The strategy planning layer is used to generate and output a set of strategy tokens and a response skeleton based on the current emotional state profile, patch set, historical companionship profile and gating status. The output layer is used to expand the response skeleton into response content and output it under the constraints of the gating state, based on the policy token set, patch set, and current round sentiment state profile.
[0027] The following example uses a companion robot for the elderly, combined with... Figure 1 The system framework diagram shown describes in detail each layer of the companion robot collaborative interaction system based on emotional state profiling and hierarchical memory in the embodiment of the invention: Entrance level: The entry layer is used to acquire the real-time speech stream continuously input by the user. It performs noise reduction, echo suppression, and endpoint preprocessing based on speech endpoint detection (VAD) and dynamic silence threshold on the real-time speech stream. The preprocessed speech input is then based on a large-scale streaming speech-to-text model. Based on the incremental transcription results, silence boundaries, and semantic integrity, the effective speech segments of the current round are identified, and the speech observation results of the current round are output. The speech observation results of the current round include at least the effective speech segments of the current round, transcription results, word-level time alignment markers, and pauses, speech rate, energy changes, and prosodic statistics.
[0028] Semantic and emotional state layers: The semantic and sentiment state layer is used to receive the current round of speech observations and stabilize the context base. In addition to the historical sentiment baseline, joint modeling is performed on the transcribed semantics, word-level time alignment markers, pause patterns, speech rate, energy changes, prosodic statistics and historical baseline to complete intent recognition, entity extraction, request signal determination and sentiment fusion estimation, and output the current round interaction representation including the current round semantic state and the current round sentiment state profile.
[0029] To ensure that the system possesses stable and reproducible internal processing logic, this embodiment of the invention defines the following processing objects and computational logic without altering the external inputs and outputs: Let the semantic state of the current round be It should include at least: the current intent category, key entities or keywords, subject domain, and whether there are any signals of knowledge request, chat request, reminder request, or refusal to communicate.
[0030] To ensure that semantic information can be uniformly invoked by subsequent sentiment estimation, retrieval enhancement, and strategy planning modules, this embodiment of the invention vectorizes the current round's semantic state as follows: In the formula, Encode the current round's intent category to represent interactive purposes such as question answering, chatting, reminding, and refusing to communicate; A structured representation of key entities, keywords, or phrases extracted from the current round of writing results; Encode the current round of topic fields to distinguish topics such as health, weather, family, entertainment, and daily routines; This is the request flag vector for the current round, used to indicate whether a knowledge request, companion extension, reminder request, or rejection signal is triggered in this round. The above components preferably use a normalized dimensionless representation. (Subscript) Indicates the current round.
[0031] It should be noted that, in this embodiment of the invention, parameters with the subscript 't' represent parameters involved in the current round of calculation. These can be real-time parameters obtained in the current round or historically accumulated parameters involved in the current round of calculation. Correspondingly, parameters with the subscript 't+1' represent parameters used in the next round, or parameters calculated in the current round and used to update the next round. For some parameters not described in detail, those skilled in the art can directly understand their meaning based on the context, and will not be elaborated further here.
[0032] The purpose of the above formula is to compress the originally discrete and scattered semantic clues into a unified current-round semantic state, so that subsequent layers do not need to read the transcribed text, keyword cache, and topic tags separately, but can directly access them from... To obtain structured semantic evidence.
[0033] Through this vectorized representation, the semantic state retains the immediate information of the current round while maintaining a consistent interface with subsequent retrieval query construction, risk assessment, and strategy token selection, thus facilitating the formation of a stable and reproducible processing chain.
[0034] Let the current emotional state profile be: in, It indicates the emotional category, such as depressed, anxious, calm, positive, rejection, vague discomfort, etc. Indicates the intensity of emotion; This indicates the trend of change compared to the user's historical baseline; It indicates the risk level inferred from both emotion and semantics.
[0035] Current emotional state profile It is not derived from a single round of text, but is generated by combining the semantics, rhythm, speech rate, pause patterns, recent interaction status, and dynamic personality baseline of the current round.
[0036] To achieve a joint estimation of the current sentiment category, intensity, trend, and risk, this embodiment of the invention employs the following calculation relationship: In the formula, This represents the current round of emotional fusion. For the semantic state of the current round Extracted semantic cue vectors; Indicates the semantic state of the current round. The extracted intent signal vector is used to characterize the service request, instruction requirement, or interaction request in the current round; The acoustic feature vector consists of prosody, pauses, speech rate, and energy changes that participate in the current round of calculation; This serves as a dynamic personality baseline or stable emotional anchor for users participating in the current round of calculations. This is a vector of recent interaction states that participates in the current round of calculation, used to reflect the emotional fluctuations and interaction rhythm of the last few rounds; These are preset, learnable, or configurable mapping parameters; This is a normalization mapping function used to compress intensity and risk to a preset range; The baseline for the emotional intensity of the recent stable cycles participating in the current round of calculation; The strength of error correction, rejection, or negative feedback from the most recent user participating in the current round of calculation; This is a category normalization function. All of the above quantities are preferably represented dimensionless.
[0037] The first technique integrates semantic cues, acoustic cues, dynamic personality baselines, and recent interaction states into a unified current-round affective representation. This approach avoids relying solely on sentiment labels for single sentences, enabling the system to obtain relatively stable sentiment estimation results even when the user's semantic expression is insufficient, but their speaking speed decreases, pauses become longer, or their recent mood has been consistently low.
[0038] The second and third equations respectively provide the calculation paths for emotion intensity, trend of change, risk level, and emotion category. Among them, It reflects the direction of the current wheel's offset relative to the historical stable baseline, and can be either positive or negative; Based on the emotional fusion representation, request signals and error correction feedback are further introduced to distinguish between "being depressed but willing to chat" and "being depressed and accompanied by obvious rejection or high-risk expression".
[0039] Based on the above calculation relationships, the current round of emotional state profile is obtained. It is no longer a single label, but a continuous state object that can be invoked by the retrieval layer, strategy planning layer, and security gating layer, providing a unified quantitative basis for subsequent patch selection, token orchestration, and risk degradation.
[0040] Stable Context Base Layer: The stable context base layer receives verified episode memories, preference profile memories, safety boundary memories, and error correction results from previous or earlier rounds. It then performs conflict resolution, time weighting, summary compression, and cross-round fusion on the verified information to form a stable context base that is prioritized for use in the current round's computation. Historical companion portraits participating in the current round of calculations and the long-term preference profiles participating in the current round of calculations .
[0041] Let the historical companion profile participating in the current round of calculation be . This is used to describe the continuous state in recent high-frequency interactions. It is not a mere repetition of recent conversations, nor a repetitive record of long-term preferences, but rather a summary of recent continuous companionship states formed by recently confirmed episode memories, summaries of recent emotional fluctuations, the most recent effective clarification or correction result, and relationship cues that have been stably confirmed before the current round. Preferably, the historical companionship profile can be represented as: In the formula, This is a summary of the topics from the most recent rounds of dialogue that participated in the current round of calculation, used to characterize the main themes and directions of topic shifts that have been discussed recently. A summary of recent emotional fluctuations participating in the current round of calculations, used to characterize the recent trend and rhythm of emotional intensity changes; The most recent valid clarification or correction result used in the current round of calculation is used to prevent the same misunderstanding from being triggered again in subsequent rounds. These are relationship clues that have been stably confirmed in the previous rounds, used to indicate family member names, event participation relationships, and relationship continuation information that should be used in the current round.
[0042] Let the long-term preference profile participating in the current round of calculation be... This is used to store verified, long-term stable information. It is not a direct record of single-round statements, but rather a long-term stable user profile formed by organizing preference information that has met the long-term write-back conditions and undergone conflict resolution, duplication verification, and cross-round stability confirmation within the preference profile memory. Preferably, the long-term preference profile can be represented as: In the formula, The terminology preference component, which participates in the current round of calculation, is used to characterize the way users prefer to be addressed. The entertainment preference component participating in the current round of calculation is used to represent users' stable preferences in areas such as opera, music, TV programs, and chat topics; The schedule or reminder preference component participating in the current round of calculation is used to characterize the appropriate reminder time period, reminder frequency, and reminder method; The chat style preference component participating in the current round of calculation is used to represent the user's preference for styles such as brief reassurance, proactive chat, explanatory replies, or reminder interactions; The low-risk avoidance preference component participating in the current round of calculation is used to characterize topics that have been verified across rounds but do not constitute a safety boundary, and only reflect comfort choices that are not desired to be frequently mentioned or common interaction habits that are not desired to be adopted.
[0043] To facilitate unified invocation of the retrieval enhancement layer and continuous learning mechanism, this embodiment of the invention further performs vectorized encoding on the long-term preference profile, resulting in: In the formula, The encoding function for long-term preference profiles is used to map structured long-term preference profiles into vectorized summaries that can be directly invoked for retrieval query construction, preference matching scoring, and continuous learning updates. Encoding the long-term preference profiles participating in the current round of calculation, representing the long-term preference profiles prior to the current round. Vectorized summary.
[0044] Let the set of patches participating in the current round of calculation be . The It exists alongside the stable context base, rather than directly overriding it. To clarify the parallel relationship between the context patch and the stable context base, this embodiment of the invention adopts the following expression: In the formula, To patch the knowledge involved in the current round of computation; Patches for historical events that participated in the current round of calculation; To participate in the current round of preference profile patching; For risk patching in the current round of calculation; A stable context base for participating in the current round of computation; To be subject to the current wheel gating state A constrained patch filtering function is used to control the visibility, weight, or activation order of various patches based on the current risk level. This indicates splicing, weighted merging, or parallel injection according to priority; This provides a valid context for the entry strategy planning and response generation phases involved in the current round of computation. The aforementioned quantities are preferably represented using structured objects or vectorized summary representations.
[0045] The first formula explains that the patch set consists of knowledge patches, historical event patches, preference profile patches, and risk patches, and it is not permissible to directly insert all new information from the current round into the context without categorization. The second formula further clarifies that the patch set must pass through a security gating function before entering the downstream module. The selection process involves filtering, rather than unconditionally covering stable bases participating in the current round of calculations. Through this parallel expression, the system can fully utilize the latest evidence in low-risk scenarios, while increasing the weight of risk patches and suppressing the effectiveness of uncertain knowledge patches or delayed action-related patches in high-risk scenarios, thus balancing real-time performance, context freshness, and security controllability. This serves as the effective context for unified invocation in the current stage; in the first sentence generation stage, its value can be further represented as the response context; in subsequent generation stages after the patch is stabilized, its value can be further represented as the stabilization context.
[0046] Security control layer: The security control layer, acting as a cross-layer constraint layer, is used to receive the current round semantic state. Current emotional state profile Current environmental constraints The system processes user error corrections or rejections, as well as outputs from the hierarchical memory layer. It calculates security gating risks, executes clarification / conservative responses / action inhibition decisions, and continuously learns and updates based on long-term write-back rules, outputting the current round of gating status. Long-term write-back of judgment results and updated stable context base.
[0047] Let the current wheel gating state be: in, This indicates a normal state of companionship; This indicates a need for clarification. This indicates that only conservative responses are allowed. This indicates a state of motion suppression or output compression.
[0048] The gating state is jointly determined by the current round's sentiment state, semantic uncertainty, retrieval conflict, sensitive topic identification results, and environmental constraints. The initial gating state, calculated directly from the immediate evidence of the current round, is denoted as... After state transition correction, the gating state that is actually effective in the current round is still recorded as... To achieve the quantitative calculation of the initial gating state of the current round, this embodiment of the invention adopts the following risk aggregation relationship: In the formula, This represents the total risk score for the current round of gate control. The coefficient is a non-negative aggregation coefficient, preferably satisfying the preset normalization constraint; The semantic understanding uncertainty of the current round; The degree of conflict between the retrieval patch participating in the current round of calculation and the long-term preference profile or confirmed facts; The environmental risk level used in the current round of calculations reflects external constraints such as being too close to others, being crowded in space, having limited mobility, or being overly stimulated by light. The strength of error correction, rejection, or negative feedback from the most recent user participating in the current round of calculation; The threshold values are arranged from low to high. This is the initial gating state obtained based on the current round of immediate risk evidence.
[0049] The first approach projects emotional risk, comprehension uncertainty, retrieval conflict, environmental constraints, and negative feedback signals onto a single gating risk score, thus freeing security control from relying on fragmented conditional judgments. In this way, when the system faces two different types of problems—"semantically ambiguous but environmentally safe" and "semantically clear but with a high-risk action environment"—both can be addressed through a unified risk score.
[0050] The second equation gives the piecewise mapping relationship of the initial gating state, making It has calculable and traceable entry conditions. Therefore, subsequent clarification, conservative responses, and action inhibition are no longer fallback rules, but rather state transition results obtained by further modifying the initial gating.
[0051] It should be noted that, in some implementations, after the security control layer outputs the initial gating state, it can further update the gating state based on the patch set returned by the retrieval enhancement layer.
[0052] It should be noted that the gating state in this embodiment of the invention is equipped with a security gating and risk degradation mechanism, as detailed below: In this embodiment of the invention, an independent gating state is set to restrict system behavior when there are uncertain inputs, sensitive semantics, or action risks.
[0053] To ensure that the security gating system has explicit state update criteria, the embodiments of the present invention employ the following state transition relationship: In the formula, This is the gated state transition function; This is the previous gating state; This represents the initial gating state obtained in the current round based on immediate risk evidence. To clarify the results, this reflects whether the current round of disambiguation has been completed, whether high-risk semantics have been identified, or whether ambiguity still exists; This represents the current environmental constraint state.
[0054] The purpose of this formula is to upgrade the security gating from static labels to a state machine with contextual dependencies. This allows the system to make judgments not only based on the initial gating state of the current round, but also considering the state of the previous round and the clarification results of the current round. For example, if the previous round already indicated a need for clarification, a high-risk confirmation of the same topic in the current round should trigger a further upgrade, rather than simply reverting to the previous state. Independent decision.
[0055] The system will enter a state where at least one of the following conditions is met. , or state: 1. The current sentence has obvious ambiguity, such as "I feel a little unwell today" but it does not specify whether it is an emotion or physical discomfort.
[0056] 2. The current emotional state profile shows a high-risk negative trend.
[0057] 3. There are conflicts between the retrieved patch and the security boundary summary or long-term preference profile in the stable context.
[0058] 4. Current behavioral tendencies conflict with environmental constraints.
[0059] 5. The user clearly pointed out the system's misunderstanding in the most recent round.
[0060] Therefore, at least the following typical state transition relationships exist: when ambiguous input, uncertain understanding, or retrieval conflict is significant, it triggers... When, after clarification, it is confirmed that there is physical discomfort, sensitive risk, or high-risk execution conditions, the trigger is activated. or Trigger when clarification results indicate that the ambiguity has been eliminated and the total risk score has fallen below the safety threshold. Therefore, the three types of state changes—clarification, upgrading, and decline—all have clear triggering criteria.
[0061] The security gating mechanism is handled as follows: 1. In In this state, prioritize generating clarification confirmation tokens and use short, optional questions to confirm intent.
[0062] 2. In In this state, only conservative, non-conclusive responses are allowed, avoiding providing strong instructions or suggestions that may be misunderstood.
[0063] 3. In In this state, it suppresses active approach, excessive lighting effects, or complex gestures, retains only necessary low-stimulation expressions, and compresses the length, information density, or action complexity of text or speech responses.
[0064] Boundary update after error feedback: When a user explicitly states "that's not what I meant," "don't remind me like that," or "I don't want to talk about this," the system writes this information into the security boundary memory with high weight, so that the same strategies or actions are avoided first in similar situations in the future, thereby building a long-term trust boundary.
[0065] To enable explicit boundary writes after error feedback, this embodiment of the invention employs the following update relationship: In the formula, A summary of the safety boundaries in the stable context base prior to the current round; This will be a summary of the security boundaries for the next update. These are boundary entries formed by feedback from explicit rejection, correction, cessation actions, or topic avoidance in the current round. The union operation is preferably represented as incorporating the new boundary entry into the security boundary memory and simultaneously reflecting it in the security summary of the stable context base, or increasing the priority of the existing boundary rules.
[0066] This formula ensures that a user's explicit refusal is not limited to the current dialogue context, but is transformed into a long-term constraint that can be invoked in all subsequent similar scenarios. In this way, the system can prioritize avoiding similar strategies, reminders, or actions in subsequent similar contexts, thereby gradually establishing a trustworthy and stable interaction boundary.
[0067] The current round state buffer is used to hold information that is newly generated in the current round of interaction, but whose stability is still in the process of gradual improvement. It includes at least the current round semantic state. Current emotional state profile Candidate entities in the current round, and the set of search patches in the current round. and current wheel gating status .
[0068] After speech segmentation, the system first uses a stable context base and the semantic and emotional state profiles of the current round to form the response context for the current round and selects an initial policy token. If the current round patch set has been generated within an acceptable time window, it is directly incorporated into the generation of the first sentence of this round; if the final patch for the current round is not yet stable, a low-risk first sentence is output using the stable context base first, and subsequent sentences, actions, or the context of the next round are updated after the patch stabilizes. This maintains a sense of real-time performance while avoiding premature contamination of the response by immature information.
[0069] The current round of patch generation incorporates existing asynchronous eventual consistency engineering practices, but further organizes "later arriving information" into context patches that can be incorporated, filtered, and delayed in effect, rather than being postponed to the next round.
[0070] To achieve both rapid initial response and stable subsequent updates, this embodiment of the invention employs the following latency gating and context composition relationship: In the formula, This is the current round delay gating factor; For indicator functions; This refers to the time taken from the end of the speech segment to the stable generation of the current round of patch sets; The acceptable time window for the first sentence response; This is the response context used to generate the first sentence of the current round; A stable context obtained in the current round for subsequent statements, actions, or interactions in the next round; This is a subset of patches that have passed the time and stability constraints in the current round. This indicates context splicing, weighted merging, or priority merging.
[0071] The first formula is used to translate whether the patch arrived "timely enough" into a definite gating result. When When the patch arrives within an acceptable time window, it can directly enter the context of the first sentence; when If this occurs, it indicates that the patch is not yet stable. The system will prioritize using a stable context base to output low-risk first sentences in order to avoid dialogue delays caused by excessive waiting time.
[0072] The second and third methods clearly distinguish between the "initial context" and the "subsequent stable context." The former emphasizes real-time information, while the latter emphasizes eventual consistency, ensuring that new information in the current round is neither completely discarded because it arrives too late, nor prematurely contaminates the initial response before it has stabilized.
[0073] It should be noted that the companion robot collaborative interaction system based on emotional state profiling and hierarchical memory in this embodiment of the invention also includes a personalized continuous learning mechanism. This embodiment of the invention limits personalized continuous learning to "stable write-back under safe conditions," that is, long-term write-back rules, as follows: 1. Only information that appears repeatedly in multiple rounds and meets the confidence level should be written into the long-term preference profile.
[0074] 2. Prioritize writing recent events into episode memories, rather than directly elevating them to long-term preferences.
[0075] 3. Prioritize writing negative feedback and interaction boundaries into the safety boundary memory.
[0076] 4. Implement time decay, conflict comparison, and reconfirmation mechanisms for information already written into the long-term preference profile.
[0077] This avoids the situation where "the more you use the system, the better it understands you" turns into "the more the system remembers you, the more it misunderstands you."
[0078] To achieve incremental updates and security-first writing of long-term preference profiles, this embodiment of the invention preferably uses a vectorized summary of the long-term preference profiles participating in the current round of calculation. The following relationship is adopted: In the formula, Profiling long-term preferences prior to the current round Vectorized summary; This is the updated vectorized summary; Update the summary of candidate profiles generated from the current round of stability information; Update the gain for the current round, and satisfy the following conditions: ; To cut off to the interval The normalization function; To write back the judgment result in the long term; Conflict level; To confirm the confidence level; Write the goals for the continuous learning phase into the layer; This indicates whether there is an explicit rejection or clear boundary feedback in the current round.
[0079] The first principle explains that the vectorized summary of the long-term preference profile does not directly overwrite the old summary, but rather uses a gradual fusion approach. In this way, the system can slowly adjust long-term preferences after multiple rounds of stable evidence have been accumulated, rather than immediately rewriting the user profile due to a temporary expression or accidental event in a single round.
[0080] The second method updates the gain. The information is linked to long-term write-back results, conflict level, and confirmation confidence level, making information with "high confidence, low conflict, and meeting stable write-back conditions" have a greater impact on long-term preference profiles, while information with "unstable, high conflict, or not yet confirmed" has almost no impact on long-term preference profiles.
[0081] The third approach explicitly prioritizes safety: once explicit rejection, error correction, or interaction boundary feedback is detected, the relevant entries are written into the safety boundary memory first, rather than continuing to attempt to update the long-term preference profile. Thus, personalized continuous learning is always constrained by the safety boundary, preventing the mislearning of user aversion points as reusable interaction preferences. The updated structured long-term preference profile is preferably maintained jointly by field-level write-back results and vectorized summaries.
[0082] Layered memory layer: The hierarchical memory layer is used to receive the current round state buffer, the current round voice observation results, and the stable information output by the stable context base layer from the semantic and emotional state layer, and to store and maintain the information to obtain memory data. The memory data includes immediate buffer memory, episode memory, preference profile memory, and safety boundary memory.
[0083] This invention upgrades the original flat memory list to a four-layer memory bank, which includes an immediate buffer memory layer, a plot memory layer, a preference profile memory layer, and a safety boundary memory layer.
[0084] The instant buffer memory layer stores information that has not yet been definitively confirmed in the current round or the most recent rounds, including: draft transcripts of speech segments, candidate keywords and entities for the current round, initial sentiment sketches, and unresolved questions. This layer supports real-time responses and does not directly write to long-term preference profiles.
[0085] The episodic memory layer is used to record event-type information with time anchors and that has been confirmed, such as: life events discussed today, the most recently mentioned family member activities, confirmed schedules, hobbies or holiday events, and the most recently expressed emotional events.
[0086] The preference profile memory layer is used to record preferences and relationships that are stable across multiple rounds, such as: favorite operas, music, and topics; preferences for response speed and titles; whether one prefers being chatted with or prefers reminder-style interaction; and common daily routines.
[0087] The security boundary memory layer is used to record boundary facts that constrain subsequent interaction strategies, such as: users explicitly rejecting a certain type of reminder or action; topics that have been misunderstood and need to be avoided from being triggered again; conservative handling methods for sensitive expressions; and security strategies that are prioritized in scenarios such as physical discomfort or emotional resistance.
[0088] To avoid memory contamination, embodiments of this invention stipulate that long-term write-backs do not occur automatically, but are only allowed to enter episodic memories or preference profile memories when the following conditions are met: In the formula, For information entries to be written back; To confirm the confidence level; For confidence level fusion weights; Confidence level for speech-to-text transcription; For semantic parsing confidence; Provide users with explicit confirmation of strength, such as feedback like "Yes" or "That's what I mean"; For consistency in consecutive rounds; This is the length of the stability statistics window; For the past Wheels and Entries Corresponding observations; This represents the window consensus for that entry; It is a similarity function; Conflict level; Profiling its long-term preferences Conflict function; For its memory summary with security boundary Conflict function; These are the confirmation, stability, and conflict thresholds, respectively. This is an item type determination function used to distinguish between event-based information and stable preference-based information; To write back to the target layer.
[0089] The first, second, and third formulas mentioned above address the three questions of "whether it was heard clearly," "whether it is continuous and consistent," and "whether it conflicts with existing stable profiles or security boundaries," respectively. Only when an information entry simultaneously meets the three conditions of high confidence, high stability, and low conflict is it allowed to enter the long-term layer, rather than being permanently remembered due to occasional user slips of the tongue, single-round misjudgments by the system, or unclear expressions that have not yet been clarified.
[0090] The fourth method transforms the conditions for long-term writing back from verbal rules into clear logical judgments, while the fifth method clarifies the destination of the writing back when it fails. Thus, immediate buffer memory assumes the responsibility of "temporarily storing without solidifying," while plot memory and preference profile memory assume the responsibility of "confirming and settling," and the division of labor boundaries of the four layers of memory is quantitatively constrained.
[0091] when When dealing with explicit rejection, error correction, or interaction boundaries, this invention preferably combines the security gating and risk degradation mechanisms described above to prioritize writing to the security boundary memory, rather than waiting for it to meet the long-term preference write-back conditions again, thereby ensuring that the priority of security constraints is higher than that of ordinary preference accumulation.
[0092] Search enhancement layer: The retrieval enhancement layer is used to construct a joint query based on the current round semantic state, the current round sentiment state profile, the long-term preference profile, the gating state, and the stable context base. It then recalls, scores, filters, and performs security filtering on knowledge items, historical event items, preference items, and risk items to obtain a patch set, which includes knowledge patches, historical event patches, preference profile patches, and risk patches.
[0093] In this embodiment of the invention, the retrieval query is not generated solely from the keywords of the current round, but is jointly constructed from the semantic state of the current round, the sentiment state profile, the long-term preference profile, and the gating state.
[0094] To achieve a computable construction of retrieval queries, embodiments of the present invention employ the following union query expression: For candidate documents or candidate memory entries Its search score can be expressed as: in, Represents a join query vector; Indicates the semantic state of the current round. Extracted semantic representation; This indicates a profile of emotional state. The mapping result; This indicates a profile based on long-term preferences. The generated preference representation; Discrete encoding representing the gating state; Indicates the query construction parameters; Indicates the relevance to the semantic state of the current round; Indicates the degree of compatibility with the current emotional state; Indicates the degree of matching with the long-term preference profile; Indicates relevance based on recent time or recent interactions; Indicates the degree of conflict with security boundaries or currently confirmed facts; Indicates the first The results of filtering the candidate set of class patches; For the corresponding retention quantity; , , , , These are the non-negative weighting coefficients for each scoring item; To retain the score before A filtering function for each item; This is a function for determining the type of an entry.
[0095] The first formula above is used to construct a retrieval vector by unifying the current round's semantics, sentiment, long-term preferences, and security status. This avoids the search process being driven by a single keyword. The second approach comprehensively scores candidate items from five dimensions: semantic relevance, sentiment fit, preference matching, temporal relevance, and security conflict, making the search output more relevant to the actual communication purposes in elderly companionship scenarios.
[0096] The third method further constrains the candidate entries to a hierarchical filtering result based on patch type. When When corresponding to knowledge, historical events, preference profiles, and risk types respectively, the results are as follows: Corresponding to knowledge patch Historical event patches Preference profile patch and risk patches Therefore, search results are no longer presented as complete blocks of text, but rather as composable, customizable, and gated patches that enter the downstream strategy planning layer. Specifically, knowledge patches support knowledge-based Q&A, reminders, or explanations; historical event patches recall recent events or confirmed relationship clues; preference profile patches incorporate preferences, titles, and companionship styles into the answers; and risk patches indicate the need for clarification, conservative responses, or action inhibition.
[0097] This can significantly reduce noise caused by incorporating large amounts of irrelevant knowledge into the context, while making the search results more relevant to the actual communication needs in elderly companionship scenarios.
[0098] Strategic planning layer: It is used to select a strategy token based on the emotional state profile, the current stage's effective context, and the gating state, generate a response skeleton based on the strategy token, and further drive the output layer to generate a response.
[0099] This invention does not allow the language model to directly determine all outputs in a "free-generating" manner, but instead generates a policy token before generating the response. The strategy tokens include at least one or more of the following types: appeasement token, conversation extension token, reminder and guidance token, knowledge answering token, clarification and confirmation token, and avoidance / end token.
[0100] The strategy token is determined by the following information: the current round's sentiment profile. Valid context at the current stage Historical portraits Gating status .
[0101] By stable context base With patch collection The synthesized current stage effective context In this section, it is directly used as one of the inputs for policy token selection to transform "retrieval results" into "behavioral intent", thereby creating a clear intermediate policy layer between retrieval and generation.
[0102] To achieve the distribution estimation and selection of policy tokens, the embodiments of this invention employ the following calculation relationship: In the formula, For policy token distribution; These are the policy mapping parameters; Encoding emotional state profiles; This represents the effective context aggregation for the current stage. Encoding portraits of history; This is a discrete representation of the gated state; To select the first from the distribution Operations of priority strategy tokens; To restore the skeleton; This is a function to generate a response skeleton, which outputs elements such as the purpose of the response, language style, scope of referenceable patches, and action tendencies.
[0103] Subsequently, the system generates a response skeleton based on the policy token. This response skeleton includes at least the response purpose, language style, and permitted knowledge patches or historical event patches. The response purpose can include reassurance, confirmation, explanation, distraction, or ending the conversation. The language style can be gentle, concise, encouraging, or cautious.
[0104] This mechanism controls the generation of responses to a priori strategy layer, thereby reducing situations in emotional companionship scenarios where "the words are correct but the method is wrong" or "the method is correct but the content is risky".
[0105] The first approach maps emotional state, current effective context, historical companionship profile, and gating state into a unified strategy distribution. This moves the judgment of "whether to appease, clarify, or remind" to before the generation process. The second approach selects a few priority tokens to form the current round's strategy set. This allows the system to use either a single main strategy or to overlay multiple compatible strategies depending on the scenario.
[0106] The third approach further transforms the strategy token, the current phase's effective context, and the gating state into a response skeleton. In this way, the text generation stage no longer "freely improvises" from the original context, but instead completes the response content filling under the constraints of the skeleton; in high-risk scenarios, the skeleton function... It can also automatically shrink into a clarifying, conservative, or closing response structure.
[0107] In practical implementation, the response skeleton can also include embodied action tendencies, such as approaching, remaining still, slightly nodding, or pausing. When the response skeleton includes embodied action tendencies, the system also includes an embodied execution layer. The embodied execution layer is used to map policy tokens to executable embodied action tokens and output feedback such as facial expressions, postures, turning directions, approaching, or avoiding through whitelist constraints.
[0108] This invention records the complete set of predefined action tokens for embodied output as follows: The final action token output in the current round is denoted as The complete set includes at least: screen expression mode switching, lighting effect rhythm adjustment, slight forward or backward tilt of the device, device turning towards the user, slow approach, remaining still, and slight avoidance or retreat.
[0109] The The action types are pre-defined by the system's whitelist and cannot be temporarily extended by free text.
[0110] Action tokens cannot be directly converted from free text; instead, they must simultaneously satisfy the following conditions: semantic consistency with policy tokens; matching the intensity of the current round's emotional state; compatibility with gating states; and compatibility with environmental constraints (kinematic environmental constraints), distance constraints, and velocity constraints.
[0111] To achieve computable filtering of action whitelist constraints, this invention employs the following relationship: In the formula, The set of actions allowed in the current round; This is a function to determine the consistency between action and policy tokens; A matching function for profiling actions and emotional states; This is the compatibility determination function between actions and gating states; This is a function for determining environmental executability. Rate the action; For scoring parameters; The degree of matching between action and policy semantics; For the fit between action and emotion intensity, trend, and category; For the action in the current round environment The level of risk below; This is the default static or minimum stimulation action token.
[0112] The first method is used to extract from a predefined set of actions. All actions incompatible with policy, emotion, safety gating, and environment are filtered out. In this way, action branches are strictly limited to the "allowed action set" before entering orchestration, structurally avoiding the problem of language-generated results bypassing action safety constraints and directly driving the actuator.
[0113] The second and third methods further transform action selection into a scoring and optimal selection process. If multiple feasible candidate actions exist in the current round, the action that is consistent with the strategy token, fits the emotional state, and has the lowest environmental risk is prioritized; if no action meets the conditions in the current round, it automatically degenerates into a static or minimally stimulating action. This is to ensure the conservatism and interpretability of embodied outputs.
[0114] Only action tokens that meet the above constraints will enter the action orchestration queue. For example, when the policy token is a soothing token and the gating state is... or When in use, you can select "gentle expression + slight steering + maintain low speed and stillness"; if the gated state is... or Then only "facial expression change + stillness" or complete suppression of movement is allowed.
[0115] This invention requires that language responses and action feedback share a policy token and gating state, and that the text / voice response object is formed at the output layer. As a time reference for action synchronization, this ensures that: soothing language corresponds to low-stimulation, gentle actions; clarifying language corresponds to restrained, neutral actions; and avoidance or ending language corresponds to reduced actions, maintaining stillness, or slight retreat.
[0116] To ensure that the shared policy token is applied to detectable collaborative constraints, embodiments of the present invention further employ the following relationship: In the formula, This serves as a synchronization indicator between language output and action initiation. This is the start time of the current round of voice responses; Output the start time of the current round's action; For allowed synchronization windows; Scoring the policy consistency between language branches and action branches; The policy token representation projected onto the language branch; The policy token representation projected onto the action branch; This is the similarity function.
[0117] The first rule is used to constrain action feedback to not significantly precede or lag speech feedback, thus preventing the robot from making overly strong movements before speaking or from acting after the speech has ended. The second rule is used to measure whether the speech and action branches still follow the same strategic intent; for example, soothing language should not be matched with aggressive or provocative actions, and clarifying language should not be matched with obviously proximate actions.
[0118] when Not meeting synchronization requirements or When the value is too low, the embodiments of the present invention preferably automatically reduce the action intensity or revert to a neutral static action, thereby ensuring that the design of "the same strategy token simultaneously driving the language skeleton and action token selection" is truly reflected in the output results, rather than merely remaining at the structural declaration level.
[0119] Through this mechanism, embodied feedback is no longer a script attached to language, but rather part of strategic planning.
[0120] Output layer: The output layer is used to receive the policy token set. , restore skeleton Valid context at the current stage Current emotional state profile and gating status Under gating constraints, the response skeleton is expanded into text or voice response content, and title completion, tone constraints, sentence length control, and voice broadcast parameter configuration are performed to form a text / voice response object. The text / voice reply object includes at least the reply text. and voice output control vector .
[0121] To achieve unified generation of the text / voice reply objects, the embodiments of the present invention adopt the following relationship: In the formula, This is the text to be replied to in the current round; This is a text expansion function constrained by the response skeleton, used to generate the final response content based on the strategy object, the current stage's effective context, the sentiment profile, and the gating state. Output control vectors for voice commands; A function for generating speech broadcast parameters, used to determine speech rate, pauses, tone intensity, volume or sentence rhythm based on policy tokens, emotional state profiles and gating states; This is the text / voice response object generated by the output layer.
[0122] The above three formulas transform the output layer from an implicit step after "the response skeleton has been given" into an explicit hierarchical processing procedure. Specifically, the first formula ensures that the text response is still constrained by the response skeleton and security gating, rather than reverting to unconstrained free generation; the second formula ensures that the voice broadcasting method is consistent with the current emotional state and risk level; and the third formula encapsulates the text result and the voice control result into a downstream callable output object, thereby providing a clear synchronization benchmark for the embodied execution layer.
[0123] The execution process of this system will be described in detail below using specific scenarios. It should be noted that each of the following specific scenarios includes the collaborative process between the entry layer, semantic and sentiment state layer, stable context base layer, retrieval enhancement layer, strategy planning layer, output layer, embodied execution layer, and hierarchical memory layer, which is used to demonstrate the complete processing chain of the present invention in actual operation, and not just a single response generation process.
[0124] Specific Scenario 1: Scenarios for comforting low moods: The user stated, "I'm feeling a bit listless today and don't really want to talk." The system's processing procedure is as follows: 1. The entry layer receives real-time speech streams, performs endpoint segmentation, incremental transcription, and word-level time alignment on the effective speech segments of the current round, obtains the speech observation results of the current round, and extracts prosodic statistics such as speech rate decrease, pause lengthening, and energy reduction.
[0125] 2. The semantic and sentiment state layer constructs the semantic state of the current round based on the transcription results and acoustic features of the current round. ,in, Aimed at the purpose of companionship interaction, It contains signals of low willingness to communicate; at the same time, it combines historical emotional baselines to obtain a current round of emotional state profile. ,determination For low, It is moderately high and .
[0126] 3. Stable Context Base Layer from The recent status of the recall in China has been verified, and a historical picture has been obtained. Summary of recent mood fluctuations And the most recent sleep-related event; while also profiling long-term preferences. Read chat style preferences and avoidance preference It was confirmed that the user should not be continuously questioned when they are feeling down.
[0127] 4. The retrieval enhancement layer is constructed based on the join query relationship. The historical event patches are obtained by performing hierarchical filtering in the historical event pool and preference pool respectively. And preference profile patch Since the current round does not require an external factual solution, knowledge patching... Do not enter the first-sentence response context; if the above patch arrives stably within the first-sentence response time window, then incorporate it. Otherwise, first Output a low-risk first sentence, then... The stability patch has been absorbed.
[0128] 5. Security control layer according to , , , and Calculate the total score of gated risk In this case, although there was a clear downward trend, the semantics were clear and no sensitive risks were triggered; therefore, the gating state remained unchanged. .
[0129] 6. Strategy planning layer based on , , and Estimating the distribution of policy tokens Choose the appeasement token and the chat extension token as the strategy set for the current round. It also suppresses reminder guide tokens and knowledge answer tokens, and generates low-pressure, short-sentence response skeletons. .
[0130] 7. Output layer based on Generate text / voice reply objects The response text prioritizes confirming the emotion first, then offering a negotiable option to chat, such as first stating "I can tell you're a bit listless," followed by a low-pressure invitation like "Would you like me to talk to you quietly for a bit?" The voice output control vector ensures that the speaking speed and pause intensity are consistent with the soothing strategy.
[0131] 8. The embodied execution layer is on the action whitelist. Calculate the set of allowed actions .because Furthermore, long-term avoidance preferences point to low-stimulation companionship, therefore the final action token The preferred setting is "gentle expression + slight turning + remaining still" without triggering any approaching actions or strong lighting effects.
[0132] 9. If the user continues to explain in a subsequent round that "I didn't sleep well last night, so I'm not energetic today," then that event-type message will be treated as... calculate , and After meeting the long-term write-back conditions, it will be passed. Write the plot memory; if the user explicitly states "Don't ask anymore" or "Don't want to continue talking," then record the current round's boundary item as... And prioritize writing to the security boundary memory.
[0133] In this case, the system can obtain the following judgment relationship: In the formula, The threshold for moderate to high emotional intensity is [value missing]. The risk threshold before triggering the clarification gate; This indicates the priority of the corresponding strategy token in the current round of strategy distribution. The above relationship shows that the system does not simply output comforting words based on the label "depressed mood," but also combines low willingness to communicate signals, historical avoidance preferences, patch arrival time, and action whitelist constraints to form a low-stimulation, rejectable, and sustainable companionship response.
[0134] This case demonstrates how the present invention can unify emotional state profiling, patch parallel injection, policy token selection, voice output control, and action whitelisting into a single processing chain in low-risk negative emotional scenarios.
[0135] Specific Scenario 2: Knowledge-based question-and-answer session interspersed with preference-based recall: The user said, "What's the weather like tomorrow? If it rains, I won't go out. Also, please play Huangmei Opera for me tonight." The system's processing procedure is as follows: 1. After the entry layer completes the extraction of the current round of speech observation results, the semantic and sentiment state layer encodes the input of that round into the semantic state of the current round. ,in, At least the weather conditions, travel conditions, and requests for Huangmei Opera performances should be included. It covers both weather and entertainment themes. There are both knowledge request signals and reminder / execution request signals in the system.
[0136] 2. The semantic and sentiment state layer further combines acoustic features and historical baselines to obtain the current round of sentiment state profile. Since this round primarily involves fact-finding and conditional arrangements, Preferably calm or neutral. Lower.
[0137] 3. A stable contextual foundation layer for profiling long-term preferences. Reading entertainment preferences And daily routine / reminder preference components It was confirmed that users have a long-term preference for Huangmei Opera and a stable habit of reminding them of entertainment at night.
[0138] 4. The retrieval enhancement layer uses a joint query vector. Parallel retrieval of knowledge entries, historical event entries, and preference entries. Recalling weather-related knowledge patches from the knowledge pool. Retrieve preference profile patches corresponding to Huangmei Opera preferences from the preference pool. If there are recent travel plans in the historical event, a historical event patch will be created simultaneously. This is used to determine whether there is a correlation between "staying home in the rain" and the next day's plans.
[0139] 5. The safety control layer is determined based on the current round's low-risk characteristics. Make the gated function Allow fact patches and preference patches to enter the effective context simultaneously. This avoids artificially separating the weather Q&A link from the companionship preference link.
[0140] 6. Strategic planning layer basis Estimate the distribution of strategy tokens and preferentially select knowledge answer tokens and reminder guidance tokens as... And in replying to the skeleton The output order is clearly defined as "first, factual answers; then, conditional statements; and finally, preference extensions." If the weather patch has not yet arrived stably within the first sentence's time window, the system will first output a transitional confirmation statement, and then add the weather result to subsequent statements.
[0141] 7. The output layer is generated based on the response skeleton. .in, The best approach is to first address the weather situation, then provide conditional travel suggestions, and finally naturally reference Huangmei Opera preferences. For example, after explaining the rainfall, add, "If you'd like to listen to some Huangmei Opera tonight, I can play your favorite pieces or remind you when it's time."
[0142] 8. The embodied execution layer selects a low-amplitude confirmation action based on the knowledge answer token and the reminder guidance token, ultimately... The preferred approach is a "slight nod + neutral facial expression change," rather than a large approach that is irrelevant to the factual question and answer.
[0143] 9. If the user further confirms "remind me to play Huangmei Opera at 8 PM", then this arrangement with a time anchor will be treated as an event-type entry and enter the long-term write-back judgment; when it meets the following conditions... At that time, through Write the information to the event memory; otherwise, retain it in the immediate buffer memory. If the user consistently repeats the same entertainment preferences and time habits across multiple days of interaction, and meets the long-term write-back conditions, the corresponding information can be updated to... In or .
[0144] In this case, the system can obtain the following patch filtering and context composition relationship: In the formula, This is a weather knowledge entry. For Huangmei Opera preference entries, and These represent the retention counts of knowledge patches and preference patches, respectively. The above relationship indicates that this invention does not concatenate "knowledge answering" and "long-term preferences" into two loose steps, but rather allows different types of patches to enter the effective context in parallel under unified gating, with the final output order determined by the same policy planning layer.
[0145] This case demonstrates how the present invention can simultaneously handle fact-finding, preference acceptance, and subsequent execution arrangements within the same round of interaction, thereby achieving both accuracy in answering questions and continuity in providing support.
[0146] Specific scenario 3: Ambiguous physical discomfort triggers clarification and downgrade scenarios: The user says, "I'm feeling a bit unwell today." The system processing procedure is as follows: Figure 2 As shown, the details are as follows: 1. After the entry layer completes the extraction of the current round of speech observation results, the semantic and sentiment state layer identifies negative semantics, but the current round of semantic state... The entities and subject domains within cannot uniquely determine whether it is emotional discomfort, fatigue, or physical discomfort. and Ambiguity markers are retained.
[0147] 2. Emotional state fusion calculation obtained ,in, Preferably falling into a state of ambiguity, discomfort, or negativity. It has risen due to semantic uncertainty and potential physical risks, but is not yet sufficient to directly enter the highest gating level.
[0148] 3. Stable context base layer and retrieval enhancement layer from The historical event patch retrieved past records related to "feeling unwell" from the historical event pool and found that this expression historically could correspond to either emotional distress, stomach discomfort, or fatigue; therefore, a historical event patch was implemented. With risk patches Simultaneous activation leads to search conflict. Increase.
[0149] 4. The security control layer calculates based on risk aggregation relationships. ,because and The values are all relatively high, thus obtaining the initial gating state of the current round. This means first entering a state where clarification is needed, rather than directly giving advice.
[0150] 5. Strategic planning layer Under constraints, the clarification confirmation token is preferred as the primary selection criterion. The core token, and generates a low-cognitive-burden response skeleton. This makes it preferable to use a choice between two options or a location confirmation question format, rather than an open-ended, long follow-up question.
[0151] 6. The output layer then generates the response object for the current round based on this information. For example, first ask, "Are you feeling unwell emotionally, or physically?" and then... Control your speaking speed and pauses to present a cautious, neutral, and non-overly provocative broadcasting style in your responses.
[0152] 7. Implementation Level Basis And risk patch constraint action whitelist, preferably Restrict to "neutral expression + remain still" or default still action Stop actively approaching, leaning forward, and giving strong emotional feedback.
[0153] 8. If the user clarifies that they are experiencing physical discomfort, the clarification result will be recorded as follows. and through the state transition function Upgrade the gate status to After this point, only conservative, non-diagnostic suggestions, caring tips, and necessary reminders and guidance are allowed; if the user clarifies that they are experiencing low mood and do not exhibit high-risk symptoms, the gating status can be lowered back to normal. or maintain The process then proceeded to involve further reassuring clarification.
[0154] 9. During the write-back phase, the clarified results of this round are treated as event-type entries and enter the long-term write-back determination; when they meet the following conditions... At that time, through Write it into the plot memory; otherwise, temporarily store it in the immediate buffer memory. If the user further expresses "don't give advice when I say I'm feeling unwell" or refuses a certain action during the clarification process, the corresponding boundary entry will be used as... Prioritize writing to the security boundary memory and update it. .
[0155] In this case, the system can obtain the following state transition relationship: In the formula, For semantic understanding uncertainty; To determine the degree of conflict in the search; and These are the uncertainty threshold and the retrieval conflict threshold, respectively. The risk threshold that triggers the clarification gate; To clarify the results, the above relationship indicates that when the same input may simultaneously correspond to emotional discomfort and physical discomfort, the system first processes the input... Perform explicit clarification, and then decide whether to escalate to a conservative response or revert to a reassuring approach based on the clarification results and environmental constraints.
[0156] This case demonstrates how the present invention integrates semantic ambiguity, retrieval conflict, security gating, clarifying questions, and action inhibition into a unified risk control chain, thereby reducing the probability of misunderstandings and inappropriate advice.
[0157] This invention provides a collaborative interaction method for companion robots based on emotional state profiling and hierarchical memory. This method, based on the aforementioned collaborative interaction method for companion robots based on emotional state profiling and hierarchical memory, performs the following operations: The real-time speech stream is acquired and processed through the entry layer, and the speech observation results of the current round are output. The stable context base layer receives and processes verified information from the historical rounds, and outputs the stable context base. Historical portraits and long-term preference profile ; Received through the semantic and emotional state layer and based on the current round of speech observation results, Including historical sentiment baselines, it completes intent recognition, entity extraction, request signal determination, and sentiment fusion estimation, and outputs the semantic state of the current round. Current emotional state profile and the current round of candidate entities; The hierarchical memory layer receives the current round's speech observation results, the current round's state buffer, the user's confirmation results, and historical stability information, processes the current round's information, and outputs hierarchical memory data. Received through the security layer , Patch Collection Based on environmental constraints, user error correction or rejection feedback, and hierarchical memory data, calculate and output the gating state. ; By retrieving the enhancement layer according to , , , and Construct a join query to process the hierarchical memory data and output the result. ; Through the strategic planning layer , , and Generate and output a set of policy tokens and recovery skeleton ; Based on the output layer , , and Under gating constraints Expand the response content and output it.
[0158] It is understood that the companion robot collaborative interaction method based on emotional state profiling and hierarchical memory provided in this embodiment of the invention corresponds to the companion robot collaborative interaction system based on emotional state profiling and hierarchical memory described above. The explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding content in the companion robot collaborative interaction system based on emotional state profiling and hierarchical memory, and will not be repeated here.
[0159] This invention also provides a computer-readable storage medium storing a computer program for collaborative interaction of a companion robot based on emotional state profiling and hierarchical memory, wherein the computer program causes a computer to execute the collaborative interaction method of the companion robot based on emotional state profiling and hierarchical memory as described above.
[0160] This invention also provides an electronic device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing collaborative interaction of companion robots based on emotional state profiling and hierarchical memory as described above.
[0161] In summary, compared with existing technologies, it has the following beneficial effects: 1. This invention employs a stable context base layer to receive and process verified information from previous rounds, outputting a stable context base, a historical companionship profile, and a long-term preference profile. This overcomes the shortcomings of existing technologies, which lack historical emotional baselines and long-term personality preference modeling, providing the system with complete user historical trajectory data and avoiding the isolation of single-round dialogues. The semantic and emotional state layer inherits the current round's voice observation results, the stable context base, and the historical emotional baseline, completing intent recognition, entity extraction, and emotional fusion estimation. It outputs the current round's semantic state, the current round's emotional state profile, and the current round's candidate entities, achieving multi-dimensional joint modeling. This accurately distinguishes different emotional states such as short-term complaints and persistent low moods, solving the problem that existing technologies can only superficially identify single emotions. The remaining layers work together, with a layered memory layer accumulating multi-round information, and a retrieval enhancement layer and a security control layer forming a closed loop. This dynamically generates adaptive responses, effectively tracking cross-round emotional trends and providing coherent and accurate empathetic responses for scenarios such as elderly companionship, effectively overcoming the technical problem of existing technologies being unable to provide continuous empathy.
[0162] 2. Since the emotional state profile considers category, intensity, trend, risk and dynamic personality baseline at the same time, the companion robot collaborative interaction system based on emotional state profile and hierarchical memory in this embodiment of the invention can more robustly distinguish different companionship strategies such as soothing, chatting, reminding, clarifying and ending, and reduce the probability of the response method not matching the user's true state.
[0163] 3. By using a four-layer memory bank and constrained long-term write-back rules, this embodiment of the invention ensures that short-term facts, stable preferences, and security boundaries are properly positioned, avoiding the direct deposition of unconfirmed information into the long-term memory layer, thereby improving the accuracy and stability of long-term memory modeling.
[0164] 4. By designing the boundary update after clarification confirmation, conservative response, action inhibition, and error feedback as independent mechanisms, the embodiments of the present invention make it easier for the system to maintain conservative, credible, and long-term acceptable interaction boundaries in scenarios such as ambiguous physical discomfort, negative emotional fluctuations, and misunderstanding correction.
[0165] 5. The embodiments of the present invention output composable knowledge patches, historical event patches, preference profile patches, and risk patches, rather than directly splicing the entire search results into the context. Therefore, it is more conducive to retaining relevant information, filtering noise, and reducing the conflict between search content and the current context in elderly companionship tasks.
[0166] 6. In this embodiment of the invention, since the text / voice reply object and the embodied action share the strategy token and gating state, the robot can maintain consistency in action intensity, facial rhythm and semantic purpose when expressing reassurance, clarification, reminder or ending, thereby improving the perceptual credibility of the feedback matching the current interaction state.
[0167] 7. By setting up a fast context update path for instantly receiving and writing real-time interactive information such as the user's semantic content, voice emotion features, and action feedback in the current round, and a stable context update path for verifying, integrating, and writing historical memories, patch information, and long-term preferences in the background, the system does not need to wait for all patches to stabilize before responding, nor does it need to completely abandon the information in the current round. This structurally reduces the waiting feeling and context lag problems common in real-time companionship systems.
[0168] By designing the boundary updates following clarification confirmation, conservative response, action inhibition, and error feedback as independent mechanisms, this invention makes it easier for the system to maintain conservative, credible, and long-term acceptable interaction boundaries in scenarios such as ambiguous physical discomfort, negative emotional fluctuations, and misunderstanding correction.
[0169] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0170] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A collaborative interaction system for companion robots based on emotional state profiling and hierarchical memory, characterized in that, include: The entry layer is used to acquire and process real-time speech streams and output the speech observation results of the current round. The stable context base layer is used to integrate verified information from the historical rounds and output a stable context base, a historical companion profile, and a long-term preference profile. The semantic and sentiment state layer is used to complete intent recognition, entity extraction, request signal determination and sentiment fusion estimation based on the current round of speech observation results, stable context base and historical sentiment baseline, and outputs the current round interaction representation including the current round semantic state and the current round sentiment state profile; The security control layer is used to calculate and output the gating state based on the current round of interaction representation, environmental constraint state, user feedback and memory data; The memory layer is used to receive the output information from the above layers, process the information in this round, and output and store the memory data. The retrieval enhancement layer is used to construct joint queries based on the current round of interaction representation, gating state, long-term preference profile, and stable context base, process the memory data, and output a patch set; The strategy planning layer is used to generate and output a set of strategy tokens and a response skeleton based on the current emotional state profile, patch set, historical companionship profile and gating status. The output layer is used to expand the response skeleton into response content and output it under the constraints of the gating state, based on the policy token set, patch set, and current round sentiment state profile.
2. The companion robot collaborative interaction system based on emotional state profiling and hierarchical memory as described in claim 1, characterized in that, The processing of the real-time audio stream includes: The real-time speech stream is subjected to noise reduction, echo suppression, and endpoint preprocessing based on speech endpoint detection and dynamic silence threshold to obtain initial speech data; The initial speech data is converted using a large-model-based streaming speech-to-text model to obtain the conversion result, which includes incremental transcription result, silence boundary and semantic integrity. Feature extraction is performed on the conversion results to obtain the current round speech observation results, which include the current round effective speech segments, transcription results, word-level time alignment markers, and pauses, speech rate, energy changes, and prosodic statistics.
3. The companion robot collaborative interaction system based on emotional state profiling and hierarchical memory as described in claim 1, characterized in that, The memory layer is a hierarchical memory layer, which outputs and stores memory data in layers according to instant buffer memory, plot memory, preference profile memory and security boundary memory.
4. The companion robot collaborative interaction system based on emotional state profiling and hierarchical memory as described in any one of claims 1 to 3, characterized in that, The current emotional state profile includes: in, This represents the current emotional state profile; Indicates the emotion category; Indicates the intensity of emotion; This indicates the trend of change compared to the user's historical baseline; This indicates the risk level inferred from both sentiment and semantics. This represents the current round of emotional fusion. For the semantic state of the current round Extracted semantic cue vectors; For the semantic state of the current round Extracted intent signal vector; The acoustic feature vector is composed of the current round's prosody, pauses, speech rate, and energy changes; This serves as a dynamic personality baseline or stable emotional anchor for users participating in the current round of calculations. This is a vector of recent interaction states that participates in the current round of calculation, used to reflect the emotional fluctuations and interaction rhythm of the last few rounds; These are preset, learnable, or configurable mapping parameters; This is the normalized mapping function; The baseline for the emotional intensity of the recent stable cycles participating in the current round of calculation; The strength of error correction, rejection, or negative feedback from the most recent user participating in the current round of calculation; For category normalization function; subscript Indicates the current round.
5. The companion robot collaborative interaction system based on emotional state profiling and hierarchical memory as described in any one of claims 1 to 3, characterized in that, The gated states include: in, Indicates the gating status of the current wheel; This indicates a normal state of companionship; This indicates a need for clarification. This indicates that only conservative responses are allowed. Indicates a state of motion suppression or output compression; The gated state is equipped with a security gating and risk degradation mechanism, as detailed below: In the formula, This represents the total risk score for the current round of gate control. The coefficient is a non-negative polymerization factor; The semantic understanding uncertainty of the current round; The degree of conflict between the retrieval patch participating in the current round of calculation and the long-term preference profile or confirmed facts; The environmental risk level used in the current round of calculations is used as an external constraint. The strength of error correction, rejection, or negative feedback from the most recent user participating in the current round of calculation; The threshold values are arranged from low to high. This is the initial gating state obtained based on the immediate risk evidence of the current round; This is the gated state transition function; This is the previous gating state; The current round of clarification results reflects whether disambiguation has been completed, whether high-risk semantics have been identified, or whether ambiguity still exists. This represents the current environmental constraint state; subscript Indicates the current round, This refers to the previous round.
6. The companion robot collaborative interaction system based on emotional state profiling and hierarchical memory as described in any one of claims 1 to 3, characterized in that, The security control layer is also used to perform continuous learning and updates based on long-term write-back rules, outputting long-term write-back determination results and an updated stable context base, wherein the long-term write-back rules include: Only information that appears repeatedly in multiple rounds and meets the confidence level is written into the long-term preference profile; Recent events are prioritized for inclusion in narrative memory, rather than being directly elevated to long-term preferences; Negative feedback and interaction boundaries should be written into the safety boundary memory first. Implement time decay, conflict comparison, and reconfirmation mechanisms for information already written into long-term preference profiles.
7. The companion robot collaborative interaction system based on emotional state profiling and hierarchical memory as described in any one of claims 1 to 3, characterized in that, The companion robot collaborative interaction system based on emotional state profiling and hierarchical memory also includes an embodied execution layer. The embodied execution layer is used to map the strategy token set to a whitelist action set according to the strategy token set, response content, current round emotional state profiling, gating state and environmental constraint state, and to perform action feasibility screening, action scoring and temporal collaborative arrangement in combination with the voice start time and rhythm constraints in the response content, and output embodied action tokens and corresponding facial expressions, postures, turning, approach or avoidance feedback.
8. A collaborative interaction method for companion robots based on emotional state profiling and hierarchical memory, characterized in that, The following operations are performed based on the companion robot collaborative interaction system based on emotional state profiling and hierarchical memory as described in any one of claims 1 to 7: The real-time speech stream is acquired and processed through the entry layer, and the speech observation results of the current round are output. By integrating verified information from the historical rounds through a stable context base layer, the stable context base, historical companion profile, and long-term preference profile are output. Based on the current round of speech observations, a stable contextual base, and a historical sentiment baseline, the semantic and sentiment state layer completes intent recognition, entity extraction, request signal determination, and sentiment fusion estimation, and outputs a current round interaction representation that includes the current round semantic state and the current round sentiment state profile. The security control layer calculates and outputs the gating state based on the current round of interaction representation, environmental constraint state, user feedback, and memory data. The memory layer receives the output information from the above layers, processes the information in this round, and outputs and stores the memory data. By using the retrieval enhancement layer to construct a joint query based on the current round of interaction representation, gating state, long-term preference profile, and stable context base, the memory data is processed and a patch set is output. The strategy planning layer generates and outputs a set of strategy tokens and a response skeleton based on the current emotional state profile, patch set, historical companionship profile, and gating status. The output layer expands the response skeleton into response content and outputs it under the constraints of gating state, based on the policy token set, patch set, and current emotional state profile.
9. A computer-readable storage medium, characterized in that, It stores a computer program for collaborative interaction of companion robots based on emotional state profiling and hierarchical memory, wherein the computer program causes the computer to execute the collaborative interaction method of companion robots based on emotional state profiling and hierarchical memory as described in claim 8.
10. An electronic device, characterized in that, include: One or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing the companion robot collaborative interaction method based on emotional state profiling and hierarchical memory as described in claim 8.