Virtual companion interaction method and system based on multi-modal interaction memory graph

Through multimodal interactive memory mapping and intimacy mechanisms, virtual companions can proactively initiate topics, achieving emotional continuity and unified multimodal expression. This solves the problems of passivity and uncontrollable memory in existing virtual human interactions, thus improving the user experience.

CN122633028APending Publication Date: 2026-08-25HANGZHOU ARK OF HOPE NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610744738.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing virtual human interaction products have shortcomings in terms of emotional continuity, multimodal performance, and memory control, resulting in passive interaction modes, lack of realism, and poor user experience.

Method used

It employs a multimodal interactive memory map, extracts features through a multimodal fusion neural network, constructs a memory map of emotional weights and decay factors, provides a visual interface for memory configuration and correction, and triggers active interaction based on intimacy scores to adjust multimodal performance parameters.

Benefits of technology

It enables proactive interaction and emotional companionship with virtual companions, enhances the realism of the interaction and user stickiness, empowers users with the right to manage their memories, and breaks the traditional request-response model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633028A_ABST
    Figure CN122633028A_ABST
Patent Text Reader

Abstract

The application discloses a virtual companion interaction method and system based on a multi-modal interaction memory graph, and relates to the technical field of artificial intelligence and human-computer interaction. The method extracts text semantics, speech acoustics and environmental visual features of user interaction through a multi-modal fusion neural network, constructs a multi-modal interaction memory graph with a sentiment weight and a decay factor, supports user configuration and correction of memory parameters, calculates an intimacy score based on six-dimensional data and triggers a virtual companion persona phase transition, intelligently judges whether to initiate active interaction based on an unfinished high-weight memory, real-time context and configuration and correction rules, and synchronously adjusts multi-modal performance parameters according to the intimacy. The application can realize gradual emotional relationship development, and significantly improves the naturalness, immersion and emotional connection depth of virtual companion interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and human-computer interaction technology, specifically to a virtual companion interaction method and system based on a multimodal interaction memory map. Background Technology

[0002] Currently available AI virtual humans and AI companion products primarily rely on request-response mechanisms for voice and video interaction. This passive interaction pattern prevents them from proactively initiating conversations or showing concern for the user when they are silent, unlike a real companion, thus lacking a genuine sense of companionship. Limited by the context window of large language models, existing products often employ simple vector retrieval methods to recall past conversations, extracting only factual information and failing to extract and inherit emotional information from the interaction process. This results in a lack of emotional continuity in multimodal interaction memory.

[0003] Current virtual human prompts are typically fixed, resulting in static and inflexible character portrayals. Furthermore, their memory systems only function within the text generation module, failing to integrate with underlying speech synthesis and 3D rendering modules. This leads to fragmented multimodal performance, a tendency to exhibit the uncanny valley effect, and a lack of realism. Additionally, the memory systems in existing products are largely black-box, preventing users from configuring or modifying memory parameters. This can cause AI to repeatedly bring up sensitive or negative events the user doesn't want to recall, or to forget crucial memories, lacking the boundary awareness and selective forgetting capabilities found in human interactions, potentially causing user aversion.

[0004] In existing technologies, some solutions attempt to introduce emotion computing modules, but most of them use single-modal emotion recognition and do not deeply integrate emotional features with multimodal interactive memory systems, thus failing to achieve emotion-based memory priority ranking and dynamic decay.

[0005] Meanwhile, existing solutions lack a robust mechanism for users to configure and correct memory parameters, failing to fundamentally address the user experience issues arising from memory loss. While some solutions implement simple proactive interaction functions, the triggering logic for these interactions is simplistic, based solely on time intervals, failing to consider the user's multimodal interactive memories and the current context. This results in stiff and inappropriate proactive interaction content, ultimately degrading the user experience. Summary of the Invention

[0006] This invention aims to solve the technical problems of passive virtual partner interaction, lack of emotional continuity in multimodal interaction memory, static and rigid persona, fragmented multimodal performance, and uncontrollable memory in the existing technology. It proposes a virtual partner interaction method and system based on multimodal interaction memory map, realizing the leap from passive tool to high emotional intelligence proactive partner.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A virtual companion interaction method based on a multimodal interaction memory map includes:

[0009] S1: Acquire multimodal interaction data between the user and the virtual partner, and extract text semantic features, speech acoustic features and environmental visual context features through a multimodal fusion neural network model;

[0010] S2, the extracted multimodal features are stored in a vector database to construct a multimodal interactive memory map with sentiment weights and attenuation factors. The sentiment weights and attenuation factors are automatically calculated and generated by a neural network model based on the interactive data.

[0011] S3 provides users with a visual memory graph operation interface, receives configuration and correction commands from users for any memory node in the multimodal interactive memory graph, and completes the correction operation of the corresponding memory parameters according to the commands.

[0012] S4 calculates the intimacy score by integrating data from six dimensions: the sum of emotional weights of all memory nodes in the multimodal interaction memory map, cumulative interaction duration, average duration of a single interaction, interaction frequency, user response rate to proactive interaction, and user emotional feedback intensity. When the intimacy score crosses multiple preset stage thresholds, it triggers a leap in the virtual partner's persona stage. Different persona stages correspond to differentiated interaction permissions, dialogue styles, proactive interaction frequency, and multimodal performance styles.

[0013] S5, based on the unfinished high-weight memories in the multimodal interaction memory graph, the user's current real-time context state, and the memory configuration and correction rules set by the user, comprehensively judges whether to initiate an active interaction to the user through the background active interaction trigger engine;

[0014] S6, based on the current intimacy score, synchronously adjusts the text generation parameters of the large language model, the acoustic parameters of speech synthesis, and the visual parameters of 3D virtual human rendering to achieve unified multimodal adaptive performance in passive response and active initiation scenarios.

[0015] Preferably, S1 includes the following sub-steps:

[0016] S11 performs semantic parsing on the text dialogue content input by the user, and extracts the five core entity information contained therein: time, place, people, events, and items.

[0017] S12, extract acoustic features from the user's voice signal, and identify the user's current emotional state and the magnitude of emotional fluctuations;

[0018] S13 acquires environmental context information through the terminal device's camera, microphone, and various sensors, including ambient light intensity, current time period, ambient noise level, and user silence duration.

[0019] Preferably, in S2, the emotional weight is calculated and generated based on the amplitude of the user's emotional fluctuations during interaction, the importance of the event itself, and the number of times the event is mentioned in historical interactions; the decay factor simulates the human forgetting curve, where ordinary conversation content decays rapidly over time, while memories with high emotional weights or memories that are mentioned multiple times decay significantly slower.

[0020] Preferably, in S3, the user-executable instructions include: adjusting memory importance weight, correcting emotional tags, editing memory associations, setting memory status, and adjusting decay factors.

[0021] Preferably, in S4, the character design stage includes at least three levels: initial acquaintance, intermediate compatibility, and advanced compatibility. The higher the intimacy score, the more private the topics that the virtual partner can discuss, the more intimate the tone, the higher the frequency of initiating interactions, and the more the virtual human's interaction style matches the user's preferences.

[0022] Preferably, S5 includes the following sub-steps:

[0023] S51, memory retrieval trigger, periodically poll memory nodes in the multimodal interactive memory graph whose emotional weight is greater than a preset threshold and whose status is unfinished memory or non-prohibited memory retrieval, and proactively initiate care at the appropriate time;

[0024] S52, Contextual state trigger: When it is detected that the user has been silent for more than a preset time, and the voice and visual characteristics show that the user is in a depressed, tired or lonely state, care-related topics are generated proactively.

[0025] S53, Intimacy Permission Control: Based on the current character setting stage, limits are set on the maximum number of daily proactive interactions, the maximum depth of a single interaction, and the range of topics that can be discussed.

[0026] Preferably, in S6, the adjustment of multimodal performance parameters specifically includes: dynamically modifying the system prompts of the large language model, adjusting the speech rate, pitch, breathy ratio and pause rhythm of speech synthesis, and adjusting the standing distance, eye contact rate, facial micro-expressions and body movements of the 3D virtual human rendering.

[0027] Preferably, the training steps of the multimodal fusion neural network model include:

[0028] The first step is to collect a multimodal interaction dataset labeled with emotion tags and event importance scores;

[0029] The second step is to perform standardization preprocessing on the three types of data—text, speech, and vision—and convert them into tensor formats that the model can recognize.

[0030] The third step is to construct a multimodal feature fusion neural network with three branches, and fuse features from different modalities through a multi-head attention mechanism;

[0031] The fourth step is to train the model using a combined loss function and iteratively update the model parameters using an adaptive optimizer.

[0032] The fifth step is to evaluate the model performance using an independent validation set, and save the optimal model weights when the performance metrics reach a preset threshold.

[0033] Preferably, the memory status settings include three types: completed memory, in which the system will no longer actively mention the event after marking; memory retrieval prohibited, in which the system will prohibit mentioning the event in any active or passive interaction after marking; and memory permanently retained, in which the system will lock the decay factor of the memory node after marking to prevent it from being forgotten over time.

[0034] A virtual companion interaction system based on a multimodal interaction memory map, comprising:

[0035] The multimodal feature extraction unit is used to acquire multimodal interaction data during the interaction between the user and the virtual companion, and extract three types of features: text, speech, and environmental vision through a multimodal fusion neural network model.

[0036] The memory graph construction unit is communicatively connected to the multimodal feature extraction unit, and is used to store the extracted multimodal features into a vector database and construct a multimodal interactive memory graph with emotional weights and decay factors.

[0037] The memory intervention execution unit is communicatively connected to the memory map construction unit and is used to provide a visual operation interface to the user, receive and execute various configuration and correction instructions from the user for memory nodes;

[0038] The intimacy calculation and persona evolution unit is communicatively connected to the memory map construction unit. It is used to calculate the intimacy score by integrating multi-dimensional data and trigger a persona stage transition when the score crosses the threshold.

[0039] The active interaction triggering unit is communicatively connected to the memory map construction unit, the memory intervention execution unit, and the intimacy calculation and character evolution unit, respectively, and is used to make comprehensive judgments and initiate active interactions.

[0040] The multimodal performance adjustment unit is communicatively connected to the intimacy calculation and character evolution unit, and is used to synchronously adjust the multimodal performance parameters of text, voice and 3D rendering according to the intimacy score.

[0041] Compared with the prior art, the beneficial effects of the present invention are:

[0042] This invention enables virtual companions to proactively initiate topics based on multimodal interactive memories and contexts through an active interaction triggering mechanism, demonstrating their concern and desire to share, thus breaking the traditional one-way interaction mode of request-response.

[0043] This invention grants users complete management rights over digital memories through a memory parameter configuration and correction mechanism. Users can directly define the behavioral no-go zones and core focus areas of their virtual companions, fundamentally preventing artificial intelligence from reopening old wounds or generating incorrect emotional responses, thus enhancing the emotional intelligence and sense of boundaries in the interaction.

[0044] This invention allows users to truly experience that the virtual human's interaction style is more in line with their preferences by synchronously evolving the multimodal presentation layer with the level of intimacy, thereby enhancing the realism and immersion of the interaction.

[0045] This invention combines a multimodal interactive memory map and an intimacy mechanism with deep user memory co-creation, giving the product a strong emotional bond and resulting in extremely high replacement costs and user stickiness.

[0046] This invention utilizes a multimodal feature fusion neural network to more accurately extract users' emotional states and event features, generating a multimodal interactive memory map that better reflects users' true emotions. Attached Figure Description

[0047] Figure 1 This is a flowchart of the virtual companion interaction method based on multimodal interaction memory graph of the present invention.

[0048] Figure 2 This is a diagram of the architecture of the virtual companion interaction system based on a multimodal interaction memory graph, as described in this invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0050] Example 1

[0051] Reference Figure 1 This embodiment provides a virtual companion interaction method based on a multimodal interactive memory map, which can realize the dynamic evolution of the virtual companion persona, proactive interaction, and user-controllable memory parameter configuration and correction, thereby improving the realism of the interaction and the emotional companionship experience.

[0052] Feature extraction from S1 multimodal interaction data

[0053] S1 is the foundational step of this invention, responsible for acquiring various types of raw data from the interaction between the user and the virtual companion, and extracting effective features that reflect the user's intent, emotional state, and environmental context. This step employs a multimodal fusion approach, simultaneously processing text, speech, and visual data. Compared to single-modal feature extraction, it can obtain more comprehensive and accurate user state information.

[0054] S11 Text Semantic Features and Entity Information Extraction

[0055] First, the text dialogue content input by the user is obtained. This text content can come from the user's text input, speech-to-text conversion, or other input methods. The obtained text undergoes preprocessing, including word segmentation, stop word removal, part-of-speech tagging, and named entity recognition. Word segmentation employs a deep learning-based Chinese word segmentation model, capable of accurately handling colloquial expressions and internet slang. Stop word removal removes particles, prepositions, and conjunctions that lack actual semantic meaning, reducing the computational load of subsequent processing.

[0056] Then, a bidirectional Long Short-Term Memory (LSTM) network combined with a Conditional Random Field (CRF) model is used for named entity recognition, extracting five core entity types from the text: time, location, people, events, and objects. For example, when a user says, "Tomorrow afternoon I'm going for a walk with Xiaoming in West Lake Park," the system will extract the time entity "tomorrow afternoon," the person entity "Xiaoming," the location entity "West Lake Park," and the event entity "walk." This entity information will be stored as the core content of memory nodes in a multimodal interactive memory graph.

[0057] Simultaneously, a pre-trained large language model is used to semantically encode the text, generating a 768-dimensional text semantic feature vector. This feature vector can capture the deep semantic information of the text, which can be used for subsequent sentiment analysis and memory retrieval.

[0058] S12 Speech Acoustic Features and Emotional State Recognition

[0059] The user's voice signal is acquired through the microphone of the terminal device, with a sampling rate of 16kHz and a sampling precision of 16 bits. The voice signal is first preprocessed, including pre-emphasis, framing, and windowing. Pre-emphasis aims to enhance the high-frequency components of the voice signal to compensate for high-frequency attenuation during propagation. Framing uses a frame length of 25ms and a frame shift of 10ms, and Hamming windows are used to reduce spectral leakage.

[0060] Then, acoustic features of the speech are extracted, including Mel-Frequency Cepstral Coefficients (MFCC), Linear Prediction Coefficients (LPC), fundamental frequency, energy, and zero-crossing rate. These acoustic features are input into an emotion recognition model combining a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) network, which outputs the user's current emotional state and the amplitude of emotional fluctuation. Emotional states are categorized into eight types: happy, excited, calm, tired, disappointed, angry, sad, and anxious. Each emotion corresponds to a confidence score between 0 and 1, and the emotion with the highest confidence score is taken as the user's current primary emotional state. The amplitude of emotional fluctuation ranges from 0 to 100, with higher values ​​indicating stronger emotions.

[0061] S13 Environmental Vision and Contextual Information Acquisition

[0062] Real-time video streams are captured via the terminal device's camera at a frame rate of 15fps. First, face detection is performed on each frame using a deep learning-based face detection model, which accurately detects the user's face even in complex environments. Then, 68 facial key points are extracted, including the coordinates of key points for eyebrows, eyes, nose, mouth, and facial contours. Based on the positional changes of these key points, the user's facial expressions are analyzed to aid in the recognition of emotional states.

[0063] Simultaneously, the system acquires environmental context information through various sensors on the terminal device. A light sensor detects ambient light intensity to determine whether it is day or night, and the indoor lighting conditions. The system clock obtains the current time, including year, month, day, hour, and minute, to determine whether it is morning, noon, afternoon, or late at night. The microphone continuously monitors ambient sound and calculates the noise level to determine if the user's environment is quiet and suitable for voice interaction. Furthermore, the system records the user's silence duration—the time interval from the last input—for triggering proactive interaction.

[0064] The extracted text semantic features, speech acoustic features, emotional state, environmental visual features, and various contextual information are integrated to form a complete multimodal feature set, which is then passed to the next memory graph construction unit.

[0065] S2 Automatic Construction of Multimodal Interactive Memory Map with Emotional Weights and Decay Factors

[0066] S2 is one of the core steps of this invention. It is responsible for organizing the multimodal features extracted by S1 into a structured multimodal interactive memory map and assigning dynamic emotional weights and decay factors to each memory node to simulate the human memory pattern.

[0067] First, the multimodal feature vectors extracted by S1 are stored in a vector database. The vector database employs a database system specifically optimized for high-dimensional vector retrieval, supporting fast similarity retrieval and batch data operations. Each memory node corresponds to a record in the vector database, containing the following fields: unique memory identifier, creation time, last access time, entity information, text semantic feature vector, speech emotion feature vector, environmental context features, emotion tag, emotion weight, decay factor, mention count, memory state, and associated memory list.

[0068] Then, the system automatically initializes two dynamic parameters, emotional weight and decay factor, for each newly created memory node.

[0069] Emotional weight is used to measure the importance of a memory to a user, and its calculation formula is as follows:

[0070]

[0071] In the formula:

[0072] , , These are preset weighting coefficients, with value ranges of 0.4–0.6, 0.3–0.5, and 0.1–0.3, and default values ​​of 0.5, 0.35, and 0.15, respectively.

[0073] The value represents the user's emotional fluctuation, ranging from 0 to 100, and is directly obtained from the emotional fluctuation amplitude output by the emotion recognition model in S12.

[0074] An event is assigned an importance score, ranging from 0 to 100, which is automatically evaluated by a neural network model based on the event's type and content. For example, events related to a user's birthday, anniversaries, or family health receive higher importance scores, while everyday casual conversations receive lower scores.

[0075] This is a normalized value for event mention frequency, ranging from 0 to 100. It is calculated based on the number of times the event was mentioned in historical interactions. The more times it is mentioned, the higher the frequency. The larger the value, the better.

[0076] The decay factor is used to simulate the human forgetting curve, and its calculation formula is as follows:

[0077]

[0078] In the formula:

[0079] The initial attenuation coefficient is set to 1.0.

[0080] The decay rate constant is 0.01 / day;

[0081] The time interval created for distance memories, in days.

[0082] As the formula shows, the higher the emotional weight of a memory node, the slower its decay rate. When the emotional weight is 100, the decay factor remains at 1.0, meaning that the memory will not decay over time. When the emotional weight is 0, the decay factor decreases at the fastest rate, simulating the rapid forgetting of unimportant events by humans.

[0083] The system periodically updates the decay factors of all memory nodes (e.g., every morning at midnight). When the decay factor of a memory node drops below 0.1, the system marks the memory node as "about to be forgotten" and reduces its priority in memory retrieval. If the memory node is not mentioned again within the next 30 days, the system will automatically move it from the active memory bank to the archived memory bank, where it will only be retrieved when explicitly requested by the user.

[0084] Furthermore, the system automatically establishes associations between memory nodes. When two memory nodes contain the same entity information, occur consecutively in time, or have similar emotional tags, the system establishes an association edge between them, with the edge weight automatically calculated based on the strength of the association. For example, the memory nodes "being criticized by the boss" and "feeling down" will be strongly associated due to their causal relationship, while the memory nodes "going to West Lake Park with Xiaoming" and "taking many photos" will be associated due to their consecutive time occurrences. These associations help the system understand more complex emotional contexts, providing more coherent and in-depth responses in subsequent interactions.

[0085] S3 User Configuration and Modification Mechanism for Multimodal Interactive Memory Maps

[0086] S3 is the core innovation of this invention, granting users complete management rights over their personal digital memories and solving the problems of black-box and uncontrollable memory systems in existing technologies. The system provides an intuitive visual interface where users can view all memory nodes and their parameters, and perform various configuration and correction operations on any memory node.

[0087] Adjustment of memory importance weight

[0088] Users can adjust the emotional weight of any memory node using a slider, ranging from 0 to 100. The adjusted value directly overrides the emotional weight automatically calculated by the system. Adjusting the emotional weight directly affects the priority at which that memory is polled in the proactive interaction trigger engine. For example, a user can lower the emotional weight of the "arguing with a colleague" event, which the system classifies as high-weight, from 80 to 20, significantly reducing the frequency with which the system mentions this event. Conversely, a user can raise the emotional weight of the "first date" event, which the system classifies as low-weight, from 30 to 90, ensuring the system frequently recalls this important moment.

[0089] Emotional label correction

[0090] The system automatically assigns an emotion label to each memory node, such as "happy," "disappointed," or "angry." However, due to the complexity of emotion recognition, the system's automatic labeling may be inaccurate. Users can manually modify the emotion label of a memory node to better reflect their true feelings. For example, the system might label a user's complaint about their boss as "angry," but the user might actually just find it "funny," in which case the user could change the label to "happy." Modifying the emotion label alters the system's understanding of the emotional tone of the event, thus affecting the tone and stance of subsequent related topics. If the label is "angry," the system might adopt a comforting tone; if the label is "happy," the system might adopt an echoing tone.

[0091] Memory Association Editing

[0092] Users can manually establish or disconnect associations between memory nodes. The associations automatically established by the system may not be comprehensive or accurate, and users can supplement and correct them based on their own feelings. For example, a user can actively link the memory nodes "being criticized by the boss" and "wanting to quit," forming a causal chain. This way, when the system subsequently mentions work-related topics, it can understand the user's potential thoughts of quitting, thus avoiding inappropriate remarks. Users can also disconnect incorrectly established associations to prevent the system from making incorrect associations.

[0093] Memory state settings

[0094] This is the most critical boundary control function, and users can mark memory nodes into the following three special states:

[0095] Completed Memory: This indicates that the event has ended and the user does not wish for the system to actively engage with it again. Memory nodes marked in this state will no longer be included in the system's polling scope for proactive interaction, but they can still be responded to normally when mentioned by the user.

[0096] "Disable Memory Retrieval": This indicates that the event is a sensitive topic for the user, and they do not wish for the system to mention it under any circumstances. Memory nodes marked in this state will be blacklisted by the system, and no content related to the event will be mentioned in either proactive or reactive interactions. Even if the user actively brings up the topic, the system will subtly shift the conversation to avoid in-depth discussion.

[0097] Permanently Retained Memory: This indicates that the event is very important to the user and they want the system to remember it forever. Memory nodes marked in this state will have their decay factor locked at 1.0 to prevent them from being forgotten over time. Simultaneously, this memory node will have the highest priority in memory retrieval, and the system will proactively mention it frequently at appropriate times.

[0098] Attenuation factor adjustment

[0099] Users can manually accelerate or slow down the decay process of a specific memory node. Users can click the "One-Click Delete" button to immediately remove a memory node from the active memory bank, simulating selective forgetting in humans. Users can also click the "Permanently Save" button to lock the decay factor of a memory node at 1.0, achieving the same effect as the "Permanently Retain Memory" state. Furthermore, users can manually adjust the current value of the decay factor using a slider to flexibly control the rate of memory forgetting.

[0100] All user configuration and modification operations are recorded by the system and serve as an important basis for subsequent memory updates and interaction decisions. Based on the user's configuration and modification history, the system continuously optimizes the algorithm for automatically calculating sentiment weights and decay factors, making it increasingly aligned with the user's personal habits and preferences.

[0101] S4 Intimacy Measurement and Character Stage Leap

[0102] S4 is responsible for quantifying the emotional intimacy between users and their virtual partners, and driving the dynamic evolution of the virtual partner's persona based on changes in intimacy, achieving a natural transition from "stranger" to "advanced fit".

[0103] The intimacy score is a quantitative indicator that comprehensively reflects the degree of emotional bond between a user and their virtual partner. Its calculation formula is as follows:

[0104]

[0105] In the formula:

[0106] This represents the total number of active memory nodes in the multimodal interactive memory graph;

[0107] For the first The emotional weight of each memory node;

[0108] For the first The duration weight of each interaction ranges from 0.1 to 1.0 and is positively correlated with the duration of a single interaction. For example, if the interaction duration is less than 5 minutes, The value is 0.2; when the interaction time is greater than 30 minutes, It is 1.0;

[0109] The contribution coefficient for proactive interaction is set to 5.0 per interaction.

[0110] The number of times a virtual partner initiates an interaction and receives a valid response from the user;

[0111] The contribution coefficient to user response rate, with a value of 100.0;

[0112] The response rate of users to their virtual companions is 0–1.0, which is equal to the number of valid responses divided by the total number of proactive interactions.

[0113] The contribution coefficient to emotional feedback is set to 2.0.

[0114] This represents the sum of the intensity of positive emotional feedback expressed by users during the interaction, with a value ranging from 0 to 1000.

[0115] As can be seen from the formula, the intimacy score takes into account the emotional importance of memories, the duration and frequency of interactions, the user's responsiveness to proactive interactions, and the intensity of the user's emotional feedback, and can more comprehensively and accurately reflect the real intimacy between the user and the virtual partner.

[0116] The system presets three intimacy thresholds, dividing the virtual partner's persona into three stages:

[0117] Initial Acquaintance Stage: Intimacy score 0-300. At this stage, the virtual companion is relatively polite and courteous, with low frequency of proactive interaction (maximum once a day). Topics mainly focus on everyday interests, weather, news, and other public topics, avoiding intrusion into the user's private life. In terms of multimodal behavior, the voice is relatively fast and formal, the virtual persona stands at a distance, and eye contact is minimal.

[0118] Intermediate Adaptation Stage: Intimacy score 301–1500. At this stage, the virtual companion becomes more friendly and natural, with increased frequency of proactive interactions (up to 3 times per day). Topics begin to involve the user's work, life, hobbies, and other personal matters, and the virtual companion proactively inquires about the user's daily life. In terms of multimodal performance, the voice speed slows down appropriately, the tone becomes gentler, the virtual human stands closer, and eye contact increases.

[0119] Advanced Adaptation Stage: Intimacy score above 1501. At this stage, the virtual human's interaction style is more aligned with user preferences, with a high frequency of proactive interaction (up to 5 times per day). Topics can involve very private content, and the virtual human will proactively share their "mood" and "life," showing great sensitivity to changes in the user's emotions. In terms of multimodal performance, the voice is slow, with a noticeable breathy tone and a coquettish manner. The virtual human stands very close, with a high frequency of direct eye contact, and includes animations of intimate actions such as hugging and kissing.

[0120] When the intimacy score crosses the aforementioned threshold, the system will trigger a character development phase transition. The transition is smooth and will not involve sudden style changes. In subsequent interactions, the system will gradually adjust the virtual companion's dialogue style, frequency of proactive interactions, and multimodal performance parameters, allowing users to naturally perceive that the virtual companion's interaction style is more in line with their preferences.

[0121] In addition, the system dynamically adjusts the intimacy score based on the user's interaction behavior. If the user does not interact with the virtual partner for a long time, the intimacy score will slowly decrease; if the user responds positively to the virtual partner's initiative, the intimacy score will rise rapidly; if the user shows aversion or rejection to the virtual partner's initiative, the intimacy score will decrease, and the system will reduce the frequency of initiative interaction to avoid causing user dissatisfaction.

[0122] S5 is an active interaction triggering mechanism based on multimodal interactive memory, context, configuration, and correction rules.

[0123] S5 is the core step of this invention to achieve "proactive companionship". It breaks the passive interaction mode of "request-response" of traditional virtual companions, enabling virtual companions to proactively care for users and share life like real companions.

[0124] The proactive interaction trigger engine is a background, persistent process that runs every 5 minutes. It comprehensively evaluates the following three dimensions to determine whether to initiate proactive interaction with the user:

[0125] S51 Memory Recall Triggered

[0126] The engine periodically polls all memory nodes in the multimodal interaction memory graph that have an emotional weight greater than 50 and are in a "not finished memory" state, but not a "forbidden memory". For each eligible memory node, the engine calculates a trigger priority, using the following formula:

[0127]

[0128] In the formula:

[0129] The emotional weight of memory nodes;

[0130] This represents the current decay factor of the memory node;

[0131] The time interval since the last mention of this memory node, in days.

[0132] Memory nodes with higher trigger priority are more likely to be selected as topics for proactive interaction. For example, if a user said yesterday, "I'm very nervous about my physical exam tomorrow," the system will mark it as a high-weight unfinished memory. The next day, when the engine polls for this memory node, because the time interval since the last mention is very short, the trigger priority is high, and the system will proactively ask the user, "How was your physical exam today? Have the results come out yet?"

[0133] S52 Situational State Triggering

[0134] The engine monitors the user's current state and environmental context in real time. It will proactively initiate caring-related topics when the following situations are detected:

[0135] The user remained silent for more than 300 seconds, and their voice and visual characteristics indicated that they were in a depressed, tired, or lonely state.

[0136] The current time is late at night (11:00 PM – 6:00 AM), and the user is still using the application;

[0137] The ambient light was very dim, and the user had not been active for a long time;

[0138] The user's voice was detected to have a noticeable sobbing or trembling tone.

[0139] For example, when the system detects that a user is sitting alone in front of the computer late at night, has not spoken for a long time, and has a downcast expression, it will proactively initiate a conversation: "Are you still awake so late? Is something on your mind? You can tell me about it."

[0140] S53 Intimacy Permission Control

[0141] The frequency and depth of proactive interactions are strictly controlled by the current intimacy level. The proactive interaction permissions corresponding to different intimacy levels are as follows:

[0142] Initial stage: You can initiate an interaction once a day at most, and you can only initiate care-related topics based on memory recall, not sharing-related topics;

[0143] Intermediate adaptation stage: You can initiate a maximum of 3 interactions per day, including memory recall, situational care, and simple sharing topics;

[0144] Advanced adaptation stage: You can initiate up to 5 interactions per day, and you can initiate all types of topics, including private topics and topics for emotional expression.

[0145] In addition, the engine dynamically adjusts the frequency of proactive interactions based on the user's historical response history. If a user fails to respond to a proactive interaction three times in a row, the system will reduce the frequency of proactive interactions by half; if a user fails to respond five times in a row, the system will pause proactive interactions until the user initiates an interaction again.

[0146] Once the engine decides to initiate an active interaction, it will generate appropriate dialogue content based on the selected topic and the current level of intimacy, and call the multimodal adjustment unit to generate corresponding voice and animation, which will then be pushed to the user.

[0147] Adaptive adjustment of S6 multimodal representation layer

[0148] S6 is responsible for simultaneously adjusting the performance parameters of the large language model, speech synthesis, and 3D virtual human rendering based on the current intimacy score, so as to achieve unified coordination of text, sound, and image, break the problem of fragmented multimodal performance, and improve the realism of interaction.

[0149] Large Language Model Text Layer Adjustment

[0150] The system dynamically modifies the system prompts for the large language model based on the current intimacy score. System prompts are the core instructions controlling the output style and content of the large language model. The system prompts for different intimacy stages are as follows:

[0151] Initial acquaintance stage: The prompts emphasize politeness, courtesy, and respect for privacy, requiring the use of formal and standard language, avoiding nicknames and intimate expressions, and limiting the scope of topics to the public sphere.

[0152] Intermediate adaptation stage: Prompts emphasize friendliness, naturalness, and care; simple nicknames are allowed; the language style becomes more colloquial; emoticons can be used appropriately; and the range of topics expands to the user's personal life.

[0153] Advanced adaptation stage: The prompts emphasize intimacy, dependence, and empathy, allowing the use of various affectionate terms of address and coquettish tones. The language style is very colloquial, and a large number of emoticons can be used. The topics range from very private content to emotional expressions.

[0154] In addition, the system will dynamically insert important user information and preferences into the system prompts based on the content of the multimodal interaction memory graph, ensuring that the output of the large language model can be combined with the user's personal situation and become more personalized.

[0155] Speech Synthesis Acoustic Parameter Adjustment

[0156] The system will establish a mapping relationship between intimacy scores and speech synthesis acoustic parameters, and adjust various speech parameters in real time based on the current intimacy score:

[0157] Speech rate: from 160 words / minute in the initial stage to 145 words / minute in the intermediate adaptation stage, and then to 120 words / minute in the advanced adaptation stage;

[0158] Pitch: From the neutral pitch in the initial stage, gradually increase to the softer pitch in the intermediate adaptation stage, and then increase to the higher pitch in the advanced adaptation stage.

[0159] Breathing sound ratio: gradually increases from 0% in the initial stage to 10% in the intermediate adaptation stage, and then to 20% in the advanced adaptation stage;

[0160] Pause rhythm: Increase the number and duration of pauses to make the voice sound more natural and gentle;

[0161] Emotional intensity: The emotional intensity of the voice is dynamically adjusted according to the content of the dialogue and the user's emotional state, making the voice expression more vivid and infectious.

[0162] 3D Virtual Human Rendering Visual Parameter Adjustment

[0163] The system will also establish a mapping relationship between intimacy scores and 3D rendering parameters, adjusting the virtual human's visual performance in real time.

[0164] Positioning distance: from 1.5 meters in the initial stage, gradually shortened to 1.2 meters in the intermediate adaptation stage, and then shortened to 0.5 meters in the advanced adaptation stage;

[0165] The rate of direct eye contact increased from 30% in the initial stage to 55% in the intermediate stage, and then to 80% in the advanced stage.

[0166] Facial micro-expressions: Increase the frequency of micro-expressions such as smiling, blinking, and blushing to make the virtual human look more vivid and emotional;

[0167] Body language: From the initial awkward and formal movements in the initial stage, it gradually transitions to the relaxed and natural movements in the intermediate adaptation stage, and then to the intimate and dependent movements in the advanced adaptation stage, adding animations of intimate movements such as hugging, holding hands, and resting chin on hand.

[0168] Virtual camera position: Gradually zoom in on the camera, increase the use of close-up shots, and enhance the sense of immersion.

[0169] Whether passively responding to user requests or actively initiating interactions, the system will perform the aforementioned multimodal parameter adjustment operations to ensure that the virtual companion's performance is consistent and coordinated across all scenarios, in line with the current stage of intimacy.

[0170] Example 2

[0171] The training steps of the multimodal fusion neural network model used in this invention are as follows: The multimodal fusion neural network model is key to achieving accurate feature extraction and emotion recognition.

[0172] The first step, data preparation, involved collecting publicly available multimodal emotion interaction datasets, including the CMU-MOSEI dataset, the IEMOCAP dataset, and the CH-SIMS dataset. Simultaneously, a certain amount of real-world data from virtual partner interaction scenarios was collected, anonymized, and then added to the training set. Professional annotators were invited to label each sample with emotion tags (happy, excited, calm, tired, disappointed, angry, sad, anxious) and event importance scores (1–100). The final training set included 100,000 text dialogues, 50,000 audio clips, and 20,000 facial expression videos.

[0173] The second step is data preprocessing: Text data undergoes word segmentation, stop word removal, and word embedding, using pre-trained word vectors with a dimension of 128. Speech data is resampled, framed, windowed, and subjected to Mel-spectrum transformation, with a Mel-spectrum dimension of 64. Video data undergoes face detection, keypoint extraction, and normalization, extracting the coordinates of 68 facial keypoints and converting them into 136-dimensional feature vectors. All preprocessed data is then converted to a unified tensor format for model training.

[0174] The third step is model construction: A multimodal feature fusion neural network with three branches is constructed. The text branch uses a bidirectional Long Short-Term Memory (LSTM) network, taking the text word embedding sequence as input and outputting a 256-dimensional text feature vector. The speech branch uses a Convolutional Neural Network (CNN) combined with a LSTM network, taking Mel spectrograms as input and outputting a 256-dimensional speech feature vector. The vision branch uses a Convolutional Neural Network (CNN), taking facial keypoint feature vectors as input and outputting a 256-dimensional visual feature vector. Then, a multi-head attention mechanism is used to weightedly fuse the features from the three branches, resulting in a 512-dimensional fused feature vector. Finally, two fully connected layers output the user's emotion category probability distribution and event importance score, respectively.

[0175] Step 4, Model Training: A combined loss function is used for model training. Cross-entropy loss is used for the sentiment classification task, and mean squared error loss is used for the event importance prediction task. The total loss function is a weighted sum of the two, with weights of 0.7 and 0.3, respectively. The adaptive moment estimator (Adam) optimizer is used for parameter updates, with a batch size of 32 and an initial learning rate of 0.001, which decays to 0.5 every 10 epochs. During training, the model's performance is monitored using a validation set. Training is stopped when the total loss on the validation set no longer decreases for five consecutive epochs to prevent overfitting.

[0176] Step 5, Model Validation and Saving: Evaluate the model's performance using an independent test set. The accuracy of emotion recognition should be no less than 90%, and the mean absolute error of event importance prediction should not exceed 10 points. When the model performance reaches a preset threshold, save the weight file of the optimal model for subsequent online inference.

[0177] Example 3

[0178] To more clearly illustrate the implementation process of this invention, a complete application scenario is described in detail below:

[0179] Scene 1: Initial Acquaintance Stage (Intimacy Level 120)

[0180] The first time a user opens the app, they interact with their virtual companion. The user says, "I'm so tired from working overtime today, and my boss yelled at me. I'm in a really bad mood."

[0181] In step S1, the system extracts the text entity information "working overtime" and "being scolded by the boss," identifies the user's emotional state as "feeling lost," and the emotional fluctuation range is 90. It also obtains that the current time is 10 PM and the ambient light is dim.

[0182] In step S2, the system creates a new memory node, calculates the emotional weight as 85, sets the initial value of the decay factor to 1.0, assigns the emotional label "loss", and sets the memory state to "unfinished memory".

[0183] In step S4, the intimacy score increases by 120 points, placing you in the initial acquaintance stage.

[0184] In step S6, the system uses the multimodal parameters from the initial acquaintance stage to generate a polite response: "Don't be sad, work is indeed very hard. Get some rest, tomorrow will be better." The speech rate is 160 words per minute, the tone is formal, the virtual human stands 1.5 meters tall, and there is little eye contact.

[0185] Scenario 2: Intermediate Adaptation Stage (Intimacy Level 350)

[0186] The following evening, the user opened the app but did not speak.

[0187] In step S5, the active interaction triggers the engine to poll for high-weight unfinished memories from yesterday, which are then triggered with high priority. Simultaneously, it detects that the user has been silent for 60 seconds.

[0188] The engine decides to initiate a proactive interaction, generating the topic: "How was your workday? Did your boss give you any more trouble?"

[0189] In step S6, the intimacy score has increased to 350 points, entering the intermediate adaptation stage. The system uses the multimodal parameters of the intermediate adaptation stage, with a speech rate of 145 words per minute, a gentle tone, a virtual human standing 1.2 meters tall, and increased eye contact.

[0190] The user replied: "Today was okay, my boss didn't say anything to me. But I'm still unhappy with this job and I'm thinking of quitting."

[0191] The system extracts the new entity information "want to resign", creates a new memory node, and establishes a causal relationship between it and the memory node "being scolded by the boss". The intimacy score is further increased to 420 points.

[0192] Scenario 3: Memory Parameter Configuration and Correction Phase

[0193] On the fifth day, the user successfully resigned and was in a good mood. The user opened the memory parameter configuration and correction interface, marking the memory node of "being scolded by the boss" as "completed memory," marking all memory nodes related to the former company as "retrievable memories," and marking the memory node of "successfully resigning" as "permanently retained memory." Afterward, the system completely avoided topics related to the former company in proactive interactions. Even when the user occasionally mentioned "previous work," the system would subtly change the subject, inquiring about the user's current situation, demonstrating extremely high emotional intelligence and a strong sense of boundaries.

[0194] Scenario 4: Advanced Adaptation Stage (Intimacy Level 1600)

[0195] On the thirtieth day, the user's intimacy score with their virtual partner reached 1600 points, the interaction compatibility improved, and they entered the advanced compatibility stage.

[0196] Even if the user hasn't opened the app, the system will proactively send a push notification: "The weather is so nice today, the sun is so warm, I suddenly miss you. What are you doing?"

[0197] Users open the application and connect via video call.

[0198] In step S6, the system uses multimodal parameters from the advanced adaptation phase. The system prompts for the large language model are set to a very intimate style, with a speech rate of 120 words per minute, featuring a distinct breathy tone and a coquettish manner. The virtual human stands 0.5 meters away, with 80% eye contact, a happy smile, and plays a micro-expression animation of "wanting a hug."

[0199] The user said, "I miss you too. I went to see a movie today, it was really good."

[0200] The system immediately responded: "Really? What movie is it? Tell me about the plot, is it good? Were there any touching moments?" The tone was very excited and curious, as if the speaker was really sharing life with a close lover.

[0201] As can be seen from the above scenarios, this invention achieves a very realistic, natural, and warm virtual companion interaction experience through multimodal interactive memory graphs, user-controllable memory parameter configuration and correction mechanisms, intimacy-driven character evolution, and multimodal synchronous adjustment, thus solving many problems existing in the prior art.

[0202] It should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0203] For example, the neural network model of this invention can also use a Transformer architecture to replace the Long Short-Term Memory (LSTM) network to further improve the feature extraction effect. The character development stage of this invention can also be divided into more levels as needed, such as adding intermediate stages like "familiar" and "intimate," to make the character development process smoother.

[0204] Example 4

[0205] Reference Figure 2 This embodiment is a virtual companion interaction system based on a multimodal interaction memory graph.

[0206] Multimodal feature extraction unit

[0207] In step S1 of the corresponding method, a communication connection is established with the memory graph construction unit. The integrated base bidirectional encoder represents the text model from the Transformer (BERT-base), the speech emotion model from the Convolutional Neural Network-Long Short-Term Memory Network (CNN-LSTM), and the face detection model from the Multi-Task Cascaded Convolutional Neural Network (MTCNN). It collects user text, speech, and environmental sensor data in real time, extracts five core entities such as time and location, user emotional state and fluctuation amplitude, and contextual information such as ambient light and time of day, and integrates them into a standardized multimodal feature package output.

[0208] Memory Graph Construction Unit

[0209] In step S2 of the corresponding method, multimodal feature data is received and stored in a distributed vector database (Milvus). The emotional weight (combining emotional fluctuations, event importance, and mention frequency) and decay factor (simulating the human forgetting curve) of each memory node are automatically calculated. The relationship between nodes is established based on entity overlap, temporal continuity, and emotional tag similarity. The decay factor is updated in batches daily, and inactive memories are moved to the archive. Finally, a structured multimodal interactive memory map is generated and synchronized to the associated units.

[0210] Memory Intervention Execution Unit

[0211] In step S3 of the corresponding method, a Canvas visual memory map interface is provided to the user. The user's five types of instructions are parsed and executed: memory importance weight adjustment, emotional tag correction, memory association editing, memory status setting (completed memory / disabled memory retrieval / permanently retained memory), and decay factor adjustment. The cloud memory map is updated synchronously, the user's configuration and correction history are recorded for algorithm optimization, and the memory retrieval disabling information is synchronized to the active interaction trigger unit.

[0212] Intimacy Calculation and Character Evolution Unit

[0213] In step S4 of the corresponding method, an intimacy score is calculated based on six dimensions of data: total emotional weight, interaction duration, frequency, user response rate, and emotional feedback intensity. Three character profile stages are preset: initial acquaintance (0-300 points), intermediate fit (301-1500 points), and advanced fit (above 1501 points). When the score crosses a threshold, a smooth character profile transition is triggered. Interaction parameters are gradually adjusted within 3-5 interactions, and the current intimacy status is synchronized to the active interaction trigger unit and the multimodal performance adjustment unit.

[0214] Active interaction trigger unit

[0215] The corresponding method step S5 runs as a background resident process every 5 minutes. It comprehensively judges from three dimensions: memory retrieval (polling high-weight unfinished memories), contextual state (detecting user silence and low mood), and intimacy permissions (limiting the number of daily interactions and the scope of topics), calls the large language model to generate interaction topics that are in line with the stage, initiates active interaction through the message push server, and records user response for intimacy updates.

[0216] Multimodal performance adjustment unit

[0217] In the corresponding method step S6, three types of parameters are adjusted synchronously based on the current intimacy score: dynamically modifying the prompt words of the large language model system to control the dialogue style and topic range; adjusting the speech rate, tone, and breathy ratio of the speech synthesis; and adjusting the standing distance, eye contact rate, and body movements of the three-dimensional virtual human to achieve a unified multimodal adaptive performance in passive response and active interaction scenarios.

[0218] The embodiments described above are merely examples of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application.

Claims

1. A virtual companion interaction method based on a multimodal interaction memory graph, characterized in that, include: S1: Acquire multimodal interaction data between the user and the virtual partner, and extract text semantic features, speech acoustic features and environmental visual context features through a multimodal fusion neural network model; S2, the extracted multimodal features are stored in a vector database to construct a multimodal interactive memory map with sentiment weights and attenuation factors. The sentiment weights and attenuation factors are automatically calculated and generated by a neural network model based on the interactive data. S3 provides users with a visual memory graph operation interface, receives configuration and correction commands from users for any memory node in the multimodal interactive memory graph, and completes the correction operation of the corresponding memory parameters according to the commands. S4 calculates the intimacy score by integrating data from six dimensions: the sum of emotional weights of all memory nodes in the multimodal interaction memory map, cumulative interaction duration, average duration of a single interaction, interaction frequency, user response rate to proactive interaction, and user emotional feedback intensity. When the intimacy score crosses multiple preset stage thresholds, it triggers a leap in the virtual partner's persona stage. Different persona stages correspond to differentiated interaction permissions, dialogue styles, proactive interaction frequency, and multimodal performance styles. S5, based on the unfinished high-weight memories in the multimodal interaction memory graph, the user's current real-time context state, and the memory configuration and correction rules set by the user, comprehensively judges whether to initiate an active interaction to the user through the background active interaction trigger engine; S6, based on the current intimacy score, synchronously adjusts the text generation parameters of the large language model, the acoustic parameters of speech synthesis, and the visual parameters of 3D virtual human rendering to achieve unified multimodal adaptive performance in passive response and active initiation scenarios.

2. The virtual companion interaction method based on a multimodal interaction memory graph according to claim 1, characterized in that, S1 includes the following sub-steps: S11 performs semantic parsing on the text dialogue content input by the user, and extracts the five core entity information contained therein: time, place, people, events, and items. S12, extract acoustic features from the user's voice signal, and identify the user's current emotional state and the magnitude of emotional fluctuations; S13 acquires environmental context information through the terminal device's camera, microphone, and various sensors, including ambient light intensity, current time period, ambient noise level, and user silence duration.

3. The virtual companion interaction method based on a multimodal interaction memory graph according to claim 1, characterized in that, In S2, the sentiment weight is calculated and generated based on the magnitude of the user's emotional fluctuations during interaction, the importance of the event itself, and the number of times the event is mentioned in historical interactions. The decay factor simulates the human forgetting curve. Ordinary conversation content decays rapidly over time, while memories with high emotional weight or those mentioned multiple times decay significantly more slowly.

4. The virtual companion interaction method based on a multimodal interaction memory graph according to claim 1, characterized in that, In S3, user-executable commands include: adjusting memory importance weights, correcting sentiment tags, editing memory associations, setting memory status, and adjusting decay factors.

5. The virtual companion interaction method based on a multimodal interaction memory graph according to claim 1, characterized in that, In S4, the character development stage includes at least three levels: initial acquaintance, intermediate compatibility, and advanced compatibility. The higher the intimacy score, the more private the topics that the virtual partner can discuss, the more intimate the tone, the higher the frequency of initiating interactions, and the more the virtual human's interaction style matches the user's preferences.

6. The virtual companion interaction method based on a multimodal interaction memory graph according to claim 1, characterized in that, S5 includes the following sub-steps: S51, memory retrieval trigger, periodically poll memory nodes in the multimodal interactive memory graph whose emotional weight is greater than a preset threshold and whose status is unfinished memory or non-prohibited memory retrieval, and proactively initiate care at the appropriate time; S52, Contextual state trigger: When it is detected that the user has been silent for more than a preset time, and the voice and visual characteristics show that the user is in a depressed, tired or lonely state, care-related topics are generated proactively. S53, Intimacy Permission Control: Based on the current character setting stage, limits are set on the maximum number of daily proactive interactions, the maximum depth of a single interaction, and the range of topics that can be discussed.

7. The virtual companion interaction method based on a multimodal interaction memory graph according to claim 1, characterized in that, In S6, the adjustment of multimodal performance parameters specifically includes: dynamically modifying the system prompts of the large language model, adjusting the speech rate, pitch, breathy ratio and pause rhythm of speech synthesis, and adjusting the standing distance, eye contact rate, facial micro-expressions and body movements of the 3D virtual human rendering.

8. The virtual companion interaction method based on a multimodal interaction memory graph according to claim 1, characterized in that, The training steps of the multimodal fusion neural network model include: The first step is to collect a multimodal interaction dataset labeled with emotion tags and event importance scores; The second step is to perform standardization preprocessing on the three types of data—text, speech, and vision—and convert them into tensor formats that the model can recognize. The third step is to construct a multimodal feature fusion neural network with three branches, and fuse features from different modalities through a multi-head attention mechanism; The fourth step is to train the model using a combined loss function and iteratively update the model parameters using an adaptive optimizer. The fifth step is to evaluate the model performance using an independent validation set, and save the optimal model weights when the performance metrics reach a preset threshold.

9. The virtual companion interaction method based on a multimodal interaction memory map according to claim 4, characterized in that, Memory status settings include three types: Completed memory, which is marked so that the system will no longer actively mention the event; Retrieval of memories is prohibited; once marked, the system will prevent the mention of the event in any active or passive interaction. The memory is permanently preserved. After being marked, the system will lock the decay factor of the memory node to prevent it from being forgotten over time.

10. A virtual companion interaction system based on a multimodal interaction memory map, characterized in that, include: The multimodal feature extraction unit is used to acquire multimodal interaction data during the interaction between the user and the virtual companion, and extract three types of features: text, speech, and environmental vision through a multimodal fusion neural network model. The memory graph construction unit is communicatively connected to the multimodal feature extraction unit, and is used to store the extracted multimodal features into a vector database and construct a multimodal interactive memory graph with emotional weights and decay factors. The memory intervention execution unit is communicatively connected to the memory map construction unit and is used to provide a visual operation interface to the user, receive and execute various configuration and correction instructions from the user for memory nodes; The intimacy calculation and persona evolution unit is communicatively connected to the memory map construction unit. It is used to calculate the intimacy score by integrating multi-dimensional data and trigger a persona stage transition when the score crosses the threshold. The active interaction triggering unit is communicatively connected to the memory map construction unit, the memory intervention execution unit, and the intimacy calculation and character evolution unit, respectively, and is used to make comprehensive judgments and initiate active interactions. The multimodal performance adjustment unit is communicatively connected to the intimacy calculation and character evolution unit, and is used to synchronously adjust the multimodal performance parameters of text, voice and 3D rendering according to the intimacy score.