Real-time personalized interaction system
By designing a real-time personalized interaction system, using multimodal input fusion algorithm and other technical means, the problem of difficulty in realizing low latency, multimodal interaction and personalized services in existing systems is solved, and an efficient and personalized interaction system is realized.
Patent Information
- Application Number
- CN202510146410.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
Existing interactive systems are difficult to achieve low latency, multimodal interaction, emotional and memory systems, and cannot provide personalized and robust virtual and physical environment execution capabilities.
A real-time personalized interaction system is designed, using multi-modal input fusion algorithm, real-time interaction protocol, voice interruption processing mechanism, emotional computing and expression module and memory module. Through interaction with the environment, it gradually learns user preferences and habits.
It realizes ultra-low latency multimodal interaction, has emotional and memory systems, can provide personalized services, enhance the generalization ability of the model and task adaptability.
Smart Images

Figure CN120066270A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of interactive systems, specifically a real-time personalized interactive system. Background Art
[0002] An interactive system is a software or hardware combination that allows users to communicate with a computer or other devices in a two-way manner. The purpose of such a system is to make the communication between humans and machines more natural and smooth by providing an intuitive, easy-to-use, and efficient interface. In the context of artificial intelligence technology, the evolution of interactive systems is moving towards the direction of "agents". It will not only be combined with common devices such as mobile phones and computers, but also exist in more and more new hardware terminals.
[0003] From the perspective of the personal user field, how to achieve real-time feedback with low latency, visual understanding, and high emotional interaction, how to build a personalized memory system, and how to have robust execution capabilities in both virtual and physical environments have become important challenges for the evolution of "personal basic agents" and personalized interaction. Summary of the Invention
[0004] The present invention aims to solve the deficiencies in the background art and provides a real-time personalized interactive system. This system has ultra-low latency, supports multi-modal interaction, has an emotion and memory system, and can continuously optimize its own behavior strategy through interaction with the environment, creating a "personal basic agent" for users. It will gradually learn the user's preferences and habits and provide more personalized services.
[0005] To achieve the above object, the present invention provides the following technical solutions. A real-time personalized interactive system includes:
[0006] S1. A multi-modal input fusion algorithm;
[0007] S2. A real-time interaction protocol;
[0008] S3. A voice interruption processing mechanism;
[0009] S4. An emotion calculation and expression module;
[0010] S5. A designed emotion module;
[0011] S6. A memory module, which continuously optimizes its own behavior strategy through interaction with the environment. It will gradually learn the user's preferences and habits and provide more personalized services.
[0012] Further, step S1 includes:
[0013] S11. A data input module, which includes an audio collection module, an image acquisition module, a text processing module, and a centralized processing module;
[0014] S12. Among them, the centralized processing module preprocesses the input data, including data alignment and scale normalization;
[0015] S13. Among them, the audio data collected by the audio collection module is encoded by the Audio Encoder;
[0016] S14. Among them, the image data collected by the image acquisition module is encoded by the Image Encoder;
[0017] S15. A keyword detection module is provided in the centralized processing module for keyword detection. The core of the voice activation trigger lies in keyword detection, that is, the system can listen for and recognize specific wake-up words or command words.
[0018] Further, step S2 includes:
[0019] S21. Select an efficient transport layer protocol;
[0020] S22. Data compression and encoding optimization;
[0021] S23. Network path optimization. By using advanced routing algorithms to find the shortest path between the source node and the target node, the transmission distance of data packets can be effectively shortened, thereby reducing latency.
[0022] Further, step S3 includes:
[0023] S31. Introduce a voice activation trigger and a context-aware algorithm, so that the system can quickly respond after detecting the user's voice command, adjust or pause the ongoing task;
[0024] S32. Dialogue management. To ensure a smooth and natural multi-round dialogue experience, a powerful dialogue management system must be available. This system needs to understand the current state of the dialogue and decide how to handle new inputs accordingly. Specifically, when a user interruption signal is detected, the currently playing content should be immediately stopped and switched to the listening mode to wait for further instructions;
[0025] S33. Voice recognition engine configuration. The sensitivity of the voice recognition engine can be appropriately adjusted. Usually, the default values have been carefully tuned to balance accuracy and user experience. However, in some special application scenarios, these parameters may need to be fine-tuned. For example, in a noisy environment, the sensitivity can be appropriately reduced to avoid excessive interference, while in a quiet environment, the sensitivity can be increased to capture the user's intention faster;
[0026] S34. Voice termination timeout. Set a short time window. If no new sound is detected during this period, it is considered that the user has finished speaking. A reasonable timeout setting helps prevent premature truncation of the user's expression and also avoids unnecessary waiting.
[0027] Further, step S4 includes:
[0028] S41. The emotion module includes: an expression storage module, a language module, and a display module;
[0029] S42. Identify the output data, select appropriate expression data from the expression storage module and output it to the display module, select appropriate tone data from the language module and transmit it to the voice module, and read aloud the output data.
[0030] The present invention provides a real-time personalized interaction system, having the following beneficial effects:
[0031] The advantages of the present invention are that this system can simultaneously process data of text, audio, and images, and realize the conversion of cross-modal tasks. At the same time, it conducts end-to-end optimization design, emphasizing the full-process learning from input to output. Among them, synthetic data is the key in the optimization process, mainly used to generate large-scale training data, including various types of data augmentation such as generating text and speech from pictures or speech, and generating text from speech. This method effectively improves the generalization ability and task adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of the overall steps of the present invention.
[0033] Figure 2 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.
[0035] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or reference letters in different examples. Such repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed. In addition, the present application provides examples of various specific processes and materials, but those of ordinary skill in the art can be aware of the application of other processes and / or the use of other materials.
[0036] An embodiment of the present application provides a real-time personalized interaction system, which will be described in detail below. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0037] The present application will be described in detail below in conjunction with the accompanying drawings and specific embodiments. Please refer to Figure 1-2 In, a real-time personalized interaction system provided in this embodiment includes:
[0038] S1. Multimodal input fusion algorithm;
[0039] S2. Real-time interaction protocol;
[0040] S3. Voice interruption processing mechanism;
[0041] S4. Emotion computing and expression module, which constructs an emotion computing framework including an emotion recognition engine and a personalized response generator, and adjusts the answering style according to non-verbal cues such as the user's tone and expression, making the interaction more user-friendly;
[0042] S5. Designed emotion module;
[0043] S6. Memory module, which continuously optimizes its own behavior strategy through interaction with the environment. It will gradually learn the user's preferences and habits and provide more personalized services.
[0044] Among them, step S1 includes:
[0045] S11. Data input module, which includes an audio collection module, an image acquisition module, a text processing module, and a centralized processing module. The audio collection module collects the voice emitted by the user and converts it into coded input, while the image acquisition module captures the facial image data of the user and encodes it for input;
[0046] S12. These encoded messages are uniformly processed in the centralized processing module, and the input data is processed. The model generates an output by predicting the next token, so text or audio can be streamed in real time. Data processing includes data alignment. Data of different modalities usually have different timestamps or spatial positions, so they must be ensured to be aligned before fusion. For example, in video analysis, audio and visual information need to be synchronized. At the same time, due to the large differences in the data ranges of each modality, direct splicing may cause certain modalities to dominate the results. Therefore, each type of data should be standardized so that all inputs are at a similar order of magnitude;
[0047] S13. The audio data collected by the audio collection module is encoded by the Audio Encoder;
[0048] S14. The image data collected by the image acquisition module is encoded by the Image Encoder;
[0049] S15. A keyword detection module is set in the centralized processing module for keyword detection. The core of the voice activation trigger lies in keyword detection, that is, the system can listen for and recognize specific wake-up words or command words.
[0050] Among them, step S2 includes:
[0051] S21. Select an efficient transport layer protocol. For real-time audio and video communication, the Real-Time Transport Protocol (RTP) is a widely adopted standard. It provides end-to-end transport functions for multimedia data and monitors the quality of service (QoS) by combining with the RTP Control Protocol (RTCP);
[0052] S22. Data compression and coding optimization. Selecting an efficient codec can significantly reduce the coding time and bandwidth requirements. For example, the Opus audio codec can control the delay within 20 milliseconds while maintaining high quality; for video, formats such as H.264 / AVC or VP8 that support fast coding can be selected;
[0053] S23. Network path optimization. Using advanced routing algorithms to find the shortest path between the source node and the target node can effectively shorten the transmission distance of data packets and thus reduce latency.
[0054] Among them, step S3 includes:
[0055] S31. Introduce a voice activation trigger and a context-aware algorithm so that the system can quickly respond after detecting the user's voice command, adjusting or pausing the ongoing task;
[0056] S32. Dialogue management. To ensure a smooth and natural multi-turn dialogue experience, a powerful dialogue management system must be available. This system can understand the current state of the dialogue and decide how to handle new inputs accordingly. Specifically, when a user interruption signal is detected, the currently playing content should be immediately stopped and switched to the listening mode to wait for further instructions;
[0057] S33. Speech recognition engine configuration. The sensitivity of the speech recognition engine can be appropriately adjusted. Usually, the default values have been carefully calibrated to balance accuracy and user experience. However, in some special application scenarios, these parameters may need to be fine-tuned. For example, in a noisy environment, the sensitivity can be appropriately reduced to avoid excessive interference, while in a quiet environment, the sensitivity can be increased to capture the user's intention faster;
[0058] S34. Speech termination timeout. Set a short time window. If no new sound is detected during this period, it is considered that the user has finished speaking. A reasonable timeout setting helps prevent premature truncation of the user's expression and also avoids unnecessary waiting.
[0059] Among them, step S4 includes:
[0060] S41. The emotion module includes: an expression storage module, a language module, and a display module. The language module stores multiple languages internally, including Portuguese, Japanese, Arabic, Cantonese, etc.;
[0061] S42. The recognized output data. Select appropriate expression data from the expression storage module and output it to the display module, and select appropriate tone data from the language module and send it to the voice module, and read aloud the output data.
[0062] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0063] The above has introduced in detail a real-time personalized interaction system provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the technical solution and its core idea of the present application; those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A real-time personalized interaction system, characterized in that: include: S1, multimodal input fusion algorithm; S2, real-time interaction protocol; S3, voice interruption processing mechanism; S4, emotion calculation and expression module; S5. Design emotion module; S6, the memory module, continuously optimizes its own behavior strategy through interaction with the environment. It will gradually learn the user's preferences and habits and provide more personalized services.
2. The real-time personalized interaction system according to claim 1, characterized in that: Step S1 includes: S11, a data input module, which includes an audio collection module, an image acquisition module, a text processing module, and a centralized processing module; S12, wherein the centralized processing module pre-processes the input data, including data alignment and scale normalization; S13, wherein the audio data collected by the audio collection module is encoded by an Audio Encoder; S14, wherein the image data collected by the image acquisition module is encoded by an Image Encoder; S15. A keyword detection module is provided in the centralized processing module for keyword detection. The core of the voice activation trigger lies in keyword detection, that is, the system can monitor and recognize specific wake-up words or command words.
3. The real-time personalized interaction system according to claim 1, characterized in that: Step S2 includes: S21. Select an efficient transport layer protocol; S22, data compression and coding optimization; S23. Network path optimization uses advanced routing algorithms to find the shortest path from the source node to the target node, which can effectively shorten the transmission distance of data packets and thus reduce latency.
4. The real-time personalized interaction system according to claim 1, characterized in that: Step S3 includes: S31, introduce voice activation triggers and context-aware algorithms, so that the system can quickly respond after detecting the user's voice command, adjust or pause the ongoing task; S32. Dialogue management: In order to ensure a smooth and natural multi-round dialogue experience, a powerful dialogue management system is required. The system needs to understand the current state of the dialogue and decide how to handle new input accordingly. Specifically, when an interruption signal is detected from the user, the content being played should be stopped immediately and the user should be switched to listening mode to wait for further instructions. S33, Speech recognition engine configuration. You can choose to adjust the sensitivity of the speech recognition engine appropriately. Generally, the default values have been carefully adjusted to balance accuracy and user experience, but in some special application scenarios, you may need to fine-tune these parameters. For example, in a noisy environment, you can appropriately reduce the sensitivity to avoid excessive interference, while in a quiet environment, you can increase the sensitivity to capture the user's intention faster; S34, speech termination timeout, sets a short time window. If no new sound is detected during this period, it is considered that the user has finished speaking. Reasonable timeout settings help prevent premature truncation of the user's expression and avoid unnecessary waiting.
5. The real-time personalized interaction system according to claim 1, characterized in that: Step S4 includes: S41, the emotion module includes: an expression storage module, a language module and a display module; S42, identify the output data, select appropriate facial expression data from the expression storage module and output it to the display module, select appropriate tone data from the language module and transmit it to the sound module, and read the output data aloud.
Citation Information
Cited By
Adaptive interaction method based on multi-modal emotion calculation and robot
CN120631175A