Intelligent accompanying-oriented emotional anthropomorphic multi-mode voice interaction large model system

By using a multimodal voice interaction big model system for smart companion toys, the problems of templated emotional interaction, low accuracy of emotion recognition, difficulty in synchronizing multimodal data, insufficient cross-modal understanding, and high real-time interaction latency have been solved, realizing emotional and anthropomorphic interaction and personalized services, thus improving the user experience.

CN121997254APending Publication Date: 2026-05-08SUZHOU PINGPING MAINLAND TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU PINGPING MAINLAND TECHNOLOGY CO LTD
Filing Date
2026-01-12
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing smart companion toys face technical bottlenecks in areas such as templated emotional interaction, low accuracy of emotion recognition, difficulty in synchronizing multimodal data, insufficient cross-modal understanding, lack of user profiles, poor multi-round interaction effects, and high real-time interaction latency.

Method used

The system adopts an emotional and anthropomorphic multimodal voice interaction model for intelligent companionship, including a user interaction module, a multimodal data fusion module, a database module, a data analysis module, and a functional processing module. Through technologies such as cross-modal attention mechanism, hierarchical memory architecture, and edge-cloud collaborative reasoning architecture, it realizes feature extraction, alignment, and fusion of multimodal data, generates a unified multimodal feature representation, constructs dynamic user profiles, and provides anthropomorphic feedback and personalized services.

Benefits of technology

It enhances the naturalness and realism of emotional interaction, strengthens users' long-term immersion in interaction, reduces real-time interaction latency, realizes personalized interaction solutions, is applicable to multiple user groups, and improves user experience and market applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997254A_ABST
    Figure CN121997254A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent accompanying-oriented emotional personification multi-modal voice interaction large model system, which relates to the field of language processing and is characterized by comprising the following modules: a user interaction module, a multi-modal data fusion module, a database module, a data analysis module and a function processing module, the user interaction module comprises touch interaction, voice interaction, text interaction and local interaction. The intelligent accompanying toy has the advantages that by integrating core technologies such as multi-modal data fusion, hierarchical memory and RAG, dynamic emotion response and edge-cloud collaboration, the problems of emotion interaction templating, lack of role consistency, high response delay, insufficient utilization of complex modals, weak long-term memory ability and the like of an existing intelligent accompanying toy are systematically solved; and the reality sense, the individuation degree and the real-time performance of interaction are remarkably improved, so that high-quality accompanying experience with higher emotional value and immersion is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of language processing, and in particular to a large-scale model system for emotional, anthropomorphic, multimodal voice interaction for intelligent companionship. Background Technology

[0002] In recent years, the rapid development of artificial intelligence technology has driven multimodal voice interaction to become a research hotspot in the field of intelligent companion toys. With the emergence of open-source large models such as DeepSeek and Llama3, the development threshold for AI toys has been significantly lowered, enabling them to evolve from simple voice interaction into intelligent companions with long-term memory, emotional computing, and scene self-learning capabilities. Product forms have also expanded from traditional robots to various forms such as plush toys, figurines, and story machines, and application scenarios have extended from children's education to adult companionship and elderly care, covering the entire age market.

[0003] With the rapid development of artificial intelligence and multimodal interaction technologies, smart companion toys are gradually becoming an important product type to meet users' emotional companionship needs. Currently, some smart companion toys have adopted multimodal processing solutions, integrating multiple modal inputs such as auditory, visual, facial recognition, voice perception, and touch perception, and are attempting to achieve collaborative processing of multimodal information through relevant algorithms to enhance the user's interactive experience.

[0004] Currently, while smart companion toys, as an important application of artificial intelligence in human-computer interaction, have achieved basic interactive functions based on single-modal technologies such as speech recognition and natural language processing, they still face key technical bottlenecks: insufficient emotional and anthropomorphic qualities lead to stiff and rigid interactions lacking emotional resonance; single-modal emotion recognition has limited accuracy and lacks a dynamic weight allocation mechanism; multimodal data synchronous processing faces challenges in temporal alignment and feature matching; insufficient cross-modal understanding capabilities result in indirect and complex information processing; a lack of accurate user profiling technology for personalized services; weak multi-turn interaction context memory leads to insufficient user stickiness; and model inference and network latency result in poor real-time interactive experiences. These technical problems restrict the further development and industrial application of smart companion toys, urgently requiring new solutions to overcome existing limitations. Summary of the Invention

[0005] The purpose of this invention is to provide a large-scale model system for emotional and anthropomorphic multimodal voice interaction for intelligent companionship, which solves the technical problems of existing intelligent companion toys in terms of templated emotional interaction, low accuracy of emotion recognition, difficulty in multimodal data synchronization, insufficient cross-modal understanding, lack of user profiles, poor multi-round interaction effect, and high real-time interaction latency.

[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0007] A large-scale model system for emotional, anthropomorphic, and multimodal voice interaction aimed at intelligent companionship is characterized by comprising the following modules: a user interaction module, a multimodal data fusion module, a database module, a data analysis module, and a function processing module. The user interaction module includes touch interaction, voice interaction, text interaction, and local interaction. The multimodal data fusion module is connected to the user interaction module and employs a cross-modal attention mechanism to extract, align, and fuse features from the multimodal interaction data, generating a unified multimodal feature representation. The multimodal data fusion module includes feature extraction, data alignment, and fusion algorithms. The database module is responsible for storing multi-dimensional data. Storage and management support the long-term memory and personalized services of the intelligent agent, including storing user and cloud information, user preference records, and multimodal historical interaction records. The data analysis module is connected to the database module. The data analysis relies on multi-source data from the database to achieve in-depth mining of user behavior and needs, including user profile construction, behavior pattern analysis, and dynamic preference updates. The functional processing module is connected to the multimodal data fusion module and the data analysis module respectively. The functional processing module includes action feedback, facial expression display, emotional companionship voice dialogue, long-term memory, user profile and data mining, dynamic adjustment of response strategies, and mobile terminal interaction.

[0008] The user interaction module collects multimodal interaction data from users, including tactile, voice, image, and text inputs. The multimodal data fusion module uses a cross-modal attention mechanism to extract, align, and fuse features from the multimodal interaction data, generating a unified multimodal feature representation. The database module stores user attributes, interaction history, user preferences, and role attribute information hierarchically. The data analysis module constructs and dynamically updates user profiles based on the stored data and analyzes user behavior patterns. The functional processing module, connected to both the multimodal data fusion module and the data analysis module, drives the execution of anthropomorphic action feedback, facial expression display, and emotionally supportive voice dialogue based on the multimodal feature representation and user profile. It also utilizes retrieval enhancement generation technology to access long-term memory information in the database module to achieve role-consistent interaction.

[0009] As an improvement, the touch interaction is used to collect the user's touch force, position and frequency and convert them into interaction commands. The voice interaction is used to collect the text content and acoustic features of the user's voice through automatic speech recognition technology. The text interaction supports text input on the APP and the device, collects user questions, casual conversation content and semantic tags, as well as local interaction (lightweight interaction on the device) and cloud interaction (cloud processing of complex tasks).

[0010] As an improvement, the cross-modal attention mechanism in the multimodal data fusion module can map multimodal features such as "force-emotion" of touch, "tonality-semantics" of speech, and "visual-scene" of images to a shared semantic space, realizing intermodal information linkage. The fusion algorithm in the multimodal data fusion module adopts a multimodal Transformer and a dynamic fusion network to integrate the features of each modality to generate a unified representation, providing multi-dimensional decision-making basis for functional processing.

[0011] As an improvement, the functional processing module retrieves relevant historical information and role attributes from the long-term memory layer of the database module through a retrieval enhancement generation method, and inputs them as context into the large language model to generate responses that conform to the role settings and have long-term consistency.

[0012] As an improvement, the system adopts an edge-cloud collaborative inference architecture, in which edge devices are responsible for performing basic emotion recognition, simple command response and local data encryption, while cloud servers are responsible for performing complex multimodal fusion analysis, long text generation and model training. The edge devices and cloud servers collaborate through a task scheduling algorithm and transmit data through homomorphic encryption technology.

[0013] The functional processing module also includes a response strategy dynamic adjustment submodule, which is used to select or generate appropriate response content and style from a preset response strategy library based on the real-time fused emotional intensity and multimodal input signals.

[0014] As an improvement, the database module adopts a hierarchical memory architecture with a "short-term-medium-long-term" hierarchical memory storage mechanism and completes information retrieval through the Retrieval Enhancement Generation (RAG) method. Combined with the database module and data analysis module, it realizes dynamic updating of user profiles and memory-based personalized recommendation algorithms.

[0015] The system's multimodal voice interaction method is characterized by the following steps: Step 1: Collecting multimodal interaction data of the user's touch, voice, image, and text through the user interaction module;

[0016] Step 2: Using the multimodal data fusion module, a cross-modal attention mechanism is employed to perform feature fusion on the multimodal interaction data, generating a unified multimodal feature representation;

[0017] Step 3: Using the data analysis module, build and update user profiles based on historical data stored in the database module;

[0018] Step 4: Through the functional processing module, based on the multimodal feature representation and user profile, generate and execute anthropomorphic multimodal feedback. At the same time, during the generation process, retrieve and enhance the generation technology to call long-term memory to maintain role consistency.

[0019] As an improvement, in step two, the fusion of tactile signals specifically includes: mapping touch force, trajectory and duration information into an emotion intensity sequence through a deep learning model, and performing cross-modal semantic alignment with intonation features in speech and facial expression features in images.

[0020] The beneficial effects of this invention are as follows: by integrating multimodal data collection of "touch + voice + image + text" with cross-modal attention mechanism, and combining real-time multimodal emotion perception model and randomized expression module, it effectively solves the problems of templated emotional interaction and lack of dynamic adaptation in existing systems, generates non-fixed and differentiated emotional responses, greatly reduces the mechanical feeling, and improves the naturalness and realism of interaction.

[0021] Based on a hierarchical memory storage architecture of "short-term-medium-long-term" and retrieval-enhanced generation (RAG) technology, coupled with an anthropomorphic character attribute library and character memory module, it breaks through the limitations of existing systems' insufficient long-term memory capacity and lack of consistency of anthropomorphic characters, and achieves cross-session memory continuation and stable character characteristics, enhancing the immersive experience of users' long-term interactions.

[0022] We construct multi-dimensional dynamic user profiles and reinforcement learning-driven personalized adaptation strategies to provide customized interaction solutions for users of different ages, interests, and usage habits. At the same time, through edge-cloud collaborative inference and lightweight model compression, we significantly reduce system response latency, solve the problem of poor real-time interactivity, and meet the real-time needs of emotional companionship.

[0023] Innovative tactile signal acquisition and processing methods fully utilize tactile modal information that is difficult for existing systems to process. Combined with a full-link interactive closed-loop design of "user input-data fusion-decision-feedback", the accuracy of user intent recognition in complex scenarios is improved, providing users with emotional value that better meets their deep emotional needs.

[0024] By employing privacy-preserving technologies such as local-cloud tiered storage and federated learning, cross-device data synchronization and personalized services are achieved while ensuring user data security and privacy compliance. This approach is suitable for low-cost deployment scenarios in consumer-grade smart toys and can also cover various user groups such as children's education, adult companionship, and elderly care, significantly improving the user experience and market applicability of smart companion products. Attached Figure Description

[0025] Figure 1 This is a system architecture diagram of the large-scale emotional and anthropomorphic multimodal voice interaction system for intelligent companionship, based on the present invention. Detailed Implementation

[0026] To make the content of this invention easier to understand, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Identical components are represented by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "up," and "down" used in the following description refer to directions in the accompanying drawings, while the terms "inner" and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.

[0027] like Figure 1 As shown, a large-scale model system for emotional and anthropomorphic multimodal voice interaction for intelligent companionship is characterized by comprising the following modules: a user interaction module, a multimodal data fusion module, a database module, a data analysis module, and a function processing module. The user interaction module includes touch interaction, voice interaction, text interaction, and local interaction. The multimodal data fusion module is connected to the user interaction module and employs a cross-modal attention mechanism to extract, align, and fuse features from multimodal interaction data, generating a unified multimodal feature representation. The multimodal data fusion module includes feature extraction, data alignment, and fusion algorithms. The database module is responsible for the storage and management of multi-dimensional data, supporting the agent's long-term memory and personalized services, including storing user and cloud information, user preference records, and multimodal historical interaction records. The data analysis module is connected to the database module. Data analysis relies on multi-source data from the database to achieve in-depth mining of user behavior and needs, including user profile construction, behavior pattern analysis, and dynamic preference updates. The function processing module is connected to both the multimodal data fusion module and the data analysis module. The function processing module includes action feedback, facial expression display, emotional companionship-style voice dialogue, long-term memory, user profiling and data mining, dynamic adjustment of response strategies, and mobile terminal interaction.

[0028] The system comprises several modules: a user interaction module for collecting multimodal interaction data from users, including tactile, voice, image, and text data; a multimodal data fusion module for extracting, aligning, and fusing features from the multimodal interaction data using a cross-modal attention mechanism to generate a unified multimodal feature representation; a database module for hierarchically storing user attributes, interaction history, user preferences, and role attribute information; a data analysis module for constructing and dynamically updating user profiles based on the stored data and analyzing user behavior patterns; and a function processing module, connected to the multimodal data fusion module and the data analysis module, for driving the execution of anthropomorphic action feedback, facial expression display, and emotionally supportive voice dialogue based on the multimodal feature representation and user profile, and utilizing retrieval enhancement generation technology to call long-term memory information from the database module to achieve role-consistent interaction.

[0029] like Figure 1As shown, touch interaction is used to collect the user's touch force, position and frequency and convert them into interaction commands. Voice interaction is used to collect the text content and acoustic features of the user's voice through automatic speech recognition technology. Text interaction supports text input on the APP and the device, collects user questions, chat content and semantic tags, as well as local interaction (lightweight interaction on the device) and cloud interaction (cloud processing of complex tasks).

[0030] like Figure 1 As shown, the cross-modal attention mechanism in the multimodal data fusion module can map multimodal features such as "force-emotion" of touch, "tonality-semantics" of speech, and "visual-scene" of image to a shared semantic space, realizing information linkage between modalities. The fusion algorithm in the multimodal data fusion module adopts multimodal Transformer and dynamic fusion network to integrate the features of each modality to generate a unified representation, providing multi-dimensional decision-making basis for functional processing.

[0031] like Figure 1 As shown, the functional processing module retrieves relevant historical information and role attributes from the long-term memory layer of the database module through a retrieval enhancement generation method, and inputs them as context into the large language model to generate responses that conform to the role settings and have long-term consistency.

[0032] like Figure 1 As shown, the system adopts an edge-cloud collaborative inference architecture, in which edge devices are responsible for performing basic emotion recognition, simple command response and local data encryption, while cloud servers are responsible for performing complex multimodal fusion analysis, long text generation and model training. Edge devices and cloud servers collaborate through task scheduling algorithms and transmit data through homomorphic encryption technology.

[0033] like Figure 1 As shown, the functional processing module also includes a response strategy dynamic adjustment submodule, used to select or generate suitable response content and style from a preset response strategy library based on real-time fused emotional intensity and multimodal input signals. The database module adopts a hierarchical memory architecture, using a "short-term-medium-term" hierarchical memory storage mechanism and a retrieval-enhanced generation (RAG) method to complete information retrieval. Combined with the database module and data analysis module, it enables dynamic updating of user profiles and memory-based personalized recommendation algorithms.

[0034] like Figure 1As shown, the system's multimodal voice interaction method is characterized by the following steps: Step 1: Collecting multimodal interaction data of the user's touch, voice, image, and text through the user interaction module; Step 2: Using a multimodal data fusion module, performing feature fusion on the multimodal interaction data using a cross-modal attention mechanism to generate a unified multimodal feature representation; Step 3: Using a data analysis module, constructing and updating user profiles based on historical data stored in the database module; Step 4: Using a function processing module, generating and executing anthropomorphic multimodal feedback based on the multimodal feature representation and user profile, while simultaneously using retrieval enhancement generation technology to call upon long-term memory to maintain role consistency during the generation process. In Step 2, the fusion of tactile signals specifically includes: mapping touch force, trajectory, and duration information into an emotion intensity sequence using a deep learning model, and performing cross-modal semantic alignment with intonation features in speech and facial expression features in images.

[0035] Example 1: This example provides an emotional interaction system that analyzes the acoustic features of user speech, such as fundamental frequency and speech rate, and text semantics through a real-time multimodal emotion perception model. Combined with a randomized expression module, it generates 3-5 differentiated responses for the same emotional need. The system also constructs a context-prosodic mapping library, dynamically binding emotional scenarios with parameters such as intonation, speech rate, and volume. Furthermore, a paralinguistic enhancement module incorporates natural interjections, making emotional expression more realistic and natural.

[0036] Example 2: This example employs a hierarchical memory architecture and retrieval-enhanced generation technology, storing user historical interaction data and role attributes through a vector database. Before interaction, the system searches the role attribute database to ensure that the response matches the set personality traits and language style. Simultaneously, it dynamically adjusts the interaction strategy based on the user profile, achieving long-term memory across conversations and personalized companionship.

[0037] Example 3: Interaction Based on Tactile Fusion This example uses infrared detection and touch sensors to collect information on the force, location, and trajectory of the user's touch. A deep learning model is then used to convert the tactile signals into a sequence of emotional features. These tactile features are fused with multimodal information such as speech and vision through a cross-modal attention mechanism, ultimately generating a comprehensive interactive response that includes corresponding action feedback and voice responses.

[0038] Example 4: This example employs an edge-cloud collaborative architecture. Lightweight computations such as speech recognition and basic emotion recognition are performed at the edge to reduce latency, while complex multimodal fusion analysis and long text generation are executed in the cloud. High real-time emotional interaction is achieved through task scheduling algorithms and streaming output technology, while a layered storage strategy ensures the security of user privacy data.

[0039] Example 5: Interaction based on dynamic scene recognition. This example constructs a complex scene library containing more than 15 emotional scenarios, and identifies users' ambiguous intentions through keyword extraction and contextual reasoning. When the system detects that a user is in a complex emotion such as "loss" or "anxiety," it will proactively ask follow-up questions and adopt a multi-level empathy response strategy, gradually upgrading from ordinary comfort to deep care, to achieve precise emotional companionship.

[0040] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A large-scale model system for emotional, anthropomorphic, multimodal voice interaction for intelligent companionship, characterized in that: The system comprises the following modules: a user interaction module, a multimodal data fusion module, a database module, a data analysis module, and a function processing module. The user interaction module includes touch interaction, voice interaction, text interaction, and local interaction. The multimodal data fusion module, connected to the user interaction module, employs a cross-modal attention mechanism to extract, align, and fuse features from the multimodal interaction data, generating a unified multimodal feature representation. The multimodal data fusion module includes feature extraction, data alignment, and fusion algorithms. The database module is responsible for storing and managing multi-dimensional data, supporting the agent's long-term memory and personalized services, including storing user and cloud information, user preference records, and multimodal historical interaction records. The data analysis module, connected to the database module, relies on multi-source data from the database to achieve in-depth mining of user behavior and needs, including user profile construction, behavior pattern analysis, and dynamic preference updates. The function processing module, connected to both the multimodal data fusion module and the data analysis module, includes action feedback, facial expression display, emotionally supportive voice dialogue, long-term memory, user profiling and data mining, dynamic adjustment of response strategies, and mobile interaction.

2. The emotional, anthropomorphic, multimodal voice interaction system for intelligent companionship as described in claim 1, characterized in that, The touch interaction is used to collect the user's touch force, position and frequency and convert them into interaction commands. The voice interaction is used to collect the text content and acoustic features of the user's voice through automatic speech recognition technology. The text interaction supports text input on the APP and the device, collects user questions, casual conversation content and semantic tags, as well as local interaction (lightweight interaction on the device) and cloud interaction (cloud processing of complex tasks).

3. The emotional, anthropomorphic, multimodal voice interaction system for intelligent companionship as described in claim 2, characterized in that, The cross-modal attention mechanism in the multimodal data fusion module can map multimodal features such as "force-emotion" of touch, "tonality-semantics" of speech, and "visual-scene" of images to a shared semantic space, realizing intermodal information linkage. The fusion algorithm in the multimodal data fusion module adopts multimodal Transformer and dynamic fusion network to integrate the features of each modality to generate a unified representation, providing multi-dimensional decision-making basis for functional processing.

4. The emotional, anthropomorphic, multimodal voice interaction system for intelligent companionship as described in claim 1, characterized in that, The functional processing module retrieves relevant historical information and role attributes from the long-term memory layer of the database module using a retrieval enhancement generation method, and inputs them as context into the large language model to generate responses that conform to the role settings and have long-term consistency.

5. The emotional, anthropomorphic, multimodal voice interaction system for intelligent companionship as described in claim 1, characterized in that, The system adopts an edge-cloud collaborative inference architecture, in which edge devices are responsible for performing basic emotion recognition, simple command response and local data encryption, while cloud servers are responsible for performing complex multimodal fusion analysis, long text generation and model training. The edge devices and cloud servers collaborate through a task scheduling algorithm and transmit data through homomorphic encryption technology.

6. The emotional, anthropomorphic, multimodal voice interaction system for intelligent companionship as described in claim 1, characterized in that, The functional processing module also includes a response strategy dynamic adjustment submodule, which is used to select or generate appropriate response content and style from a preset response strategy library based on the real-time fused emotional intensity and multimodal input signals.

7. The emotional, anthropomorphic, multimodal voice interaction system for intelligent companionship as described in claim 1, characterized in that, The database module adopts a hierarchical memory architecture with a "short-term-medium-long-term" hierarchical memory storage mechanism and completes information retrieval through the Retrieval Enhancement Generation (RAG) method. Combined with the database module and data analysis module, it realizes dynamic updating of user profiles and memory-based personalized recommendation algorithms.

8. The system multimodal voice interaction method according to claim 1, characterized in that, Includes the following steps: Step 1: Collect multimodal interaction data of the user, including tactile, voice, image, and text, through the user interaction module; Step 2: Using the multimodal data fusion module, a cross-modal attention mechanism is employed to perform feature fusion on the multimodal interaction data, generating a unified multimodal feature representation; Step 3: Using the data analysis module, build and update user profiles based on historical data stored in the database module; Step 4: Through the functional processing module, based on the multimodal feature representation and user profile, generate and execute anthropomorphic multimodal feedback. At the same time, during the generation process, retrieve and enhance the generation technology to call long-term memory to maintain role consistency.

9. The system multimodal voice interaction method according to claim 8, characterized in that, In step two, the fusion of tactile signals specifically includes: mapping touch force, trajectory, and duration information into an emotional intensity sequence through a deep learning model, and performing cross-modal semantic alignment with intonation features in speech and facial expression features in images.