Real-time voice interaction digital human agent based on localized multitask workflow

By using a real-time voice interaction digital human agent based on a localized multi-task workflow, combined with online large models and local processing, the problems of insufficient real-time performance and professionalism in existing technologies are solved, achieving efficient and natural emotional interaction and professional psychological services, and improving user experience and interaction continuity.

CN120998237APending Publication Date: 2025-11-21ZHONGKE QIMENG (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511200891.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing AI digital human interaction technologies in the field of mental health suffer from insufficient real-time performance and professionalism. They exhibit significant delays in speech synthesis, stiff emotional expression, and inadequate speech recognition accuracy. Furthermore, they suffer from poor thread coordination during multi-task collaborative processing, are unable to detect user crisis signals in a timely manner, and lack emotion analysis and psychological report generation functions, thus affecting the continuity of interaction and user experience.

Method used

It adopts a real-time voice interaction digital human agent based on localized multi-task workflow, combining online large model and local processing, including local working system, online large model, local speech recognition system and multi-task mechanism system, with efficient short-term context memory, accurate intent recognition, crisis recognition and emotion analysis capabilities, and improves recognition accuracy and interaction coherence through dual-channel speech recognition and seamless speech stream splicing technology.

Benefits of technology

It achieves a combination of high-efficiency real-time voice interaction and professional mental health services, improving the accuracy of voice recognition and the naturalness of emotions, timely detecting user crises and triggering interventions, generating professional psychological reports, and ensuring the continuity of interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
Patent Text Reader

Abstract

The invention discloses a real-time voice interaction digital human agent based on localized multi-task workflow. The real-time voice interaction digital human agent comprises a local working system, an online large model, a local voice recognition system, a multi-task mechanism system and a local voice player, the local working system comprises a short-term memory layer, an intention recognition route, an RAG local knowledge base, crisis recognition, an emotional map and psychological report generation; the multi-task mechanism system comprises a digital human front-end communication task thread, a large model question and answer task thread, a subtitle task, an instruction task and a TTS processing task. According to the invention, through a mixed architecture combining an online large model and localization processing, the real-time voice interaction performance and the psychological health service effect are effectively considered, and the limitation of a single architecture is broken through. The online speech synthesis technology guarantees accurate pronunciation, natural emotion and quick response, local dual-channel speech recognition is matched with an optimization mechanism, the recognition accuracy is remarkably improved, meanwhile, speech detection and hot word interruption are supported, and the interaction flexibility is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital interaction technology, and more specifically to a real-time voice-interactive digital human agent based on a localized multi-task workflow. Background Technology

[0002] In the field of AI-powered digital human interaction technology within mental health, existing real-time voice dialogue and emotional support platforms generally suffer from architectural design limitations. Some platforms rely solely on large online models, while others are limited to local processing, making it difficult to balance interactive performance and service effectiveness, and failing to meet users' dual demands for real-time performance and professionalism. At the voice interaction level, existing speech synthesis technologies often suffer from significant latency and unnatural emotional expression, and the speech recognition process lacks efficient optimization mechanisms, resulting in insufficient accuracy in recognition results. During dialogue, most platforms have weak capabilities in remembering and recalling user contextual information, resulting in limited accuracy in intent recognition and an inability to accurately route user consultation needs to the corresponding professional intervention modules.

[0003] Meanwhile, existing solutions lack core functionalities in mental health services: they are insufficient in detecting potential crisis signals and dynamically assessing risks, making it difficult to trigger timely crisis intervention; their analysis of users' emotional states is not deep enough, and they lack the function of automatically integrating conversation data and professional indicators to generate psychological reports. Furthermore, during multi-task collaborative processing, the coordination between threads is poor, and gaps easily appear in audio stream playback, severely affecting the continuity of interaction and user experience. Summary of the Invention

[0004] The purpose of this invention is to provide a real-time voice-interactive digital human intelligent agent based on a localized multi-task workflow.

[0005] To achieve the above objectives, the present invention provides the following technical solution: The real-time voice-interactive digital human agent based on a localized multi-task workflow includes a local working system, an online large model, a local speech recognition system, a multi-task mechanism system, and a local voice player. The local working system includes a short-term memory layer, intent recognition routing, a RAG local knowledge base, crisis recognition, emotion graph, and psychological report generation. The multi-task mechanism system includes a digital human front-end communication task thread, a large model question-and-answer task thread, a subtitle task, an instruction task, and a TTS processing task.

[0006] Preferably, the online large model includes a role-playing text large model and a speech synthesis large model.

[0007] Preferably, the local speech recognition system includes a dual-channel recognition engine and other functions.

[0008] Preferably, the other functions include VAD voice effective detection and hot word interruption functions. Beneficial effects

[0009] This invention employs a hybrid architecture combining online large-scale models with localized processing, effectively balancing real-time voice interaction performance with the effectiveness of mental health services, breaking through the limitations of a single architecture. Online speech synthesis technology ensures accurate pronunciation, natural emotion, and rapid response, while local dual-channel speech recognition, coupled with optimization mechanisms, significantly improves recognition accuracy. It also supports speech detection and hot word interruption, optimizing interaction flexibility.

[0010] Furthermore, the local workflow boasts efficient short-term contextual memory and precise intent recognition capabilities, accurately routing consultation needs to the corresponding intervention modules. Combined with a local knowledge base, this enhances service professionalism. The crisis identification function promptly captures risk signals and triggers interventions, while emotion mapping analysis and automatic psychological report generation further improve the professionalism and relevance of mental health services. A multi-task asynchronous processing mechanism ensures efficient collaboration across threads, and seamless audio stream splicing technology completely eliminates playback gaps, enhancing interactive coherence and user experience.

[0011] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the system of the present invention; Figure 2 This is a schematic diagram illustrating other functions of the local speech recognition system of the present invention. Detailed Implementation

[0013] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0015] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0016] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0017] like Figures 1-2 As shown, the real-time voice interaction digital human intelligent agent based on localized multi-task workflow includes a local working system, an online large model, a local speech recognition system, a multi-task mechanism system, and a local voice player. The local working system comprises a short-term memory layer, intent recognition and routing, a RAG local knowledge base, crisis identification, emotion mapping, and psychological report generation. The short-term memory layer uses a Memoripy-like sliding window cache to retain the most recent 5-7 rounds of dialogue and stores key dialogue entities, such as user preferences and unfinished tasks, through a vector database. Intent recognition and routing involves multi-label classification of user input, identifying consultation intent, and routing to the corresponding intervention protocol module. The RAG local knowledge base is a local database built to enhance retrieval of psychological and daily-related knowledge. Crisis identification uses a large online model to detect crisis signals in user text or speech in real time, such as suicidal ideation and self-harm statements, combining dynamic risk assessment to generate risk levels and automatically trigger crisis intervention. The emotion mapping analyzes the emotions in the text based on the dialogue between the client and the counselor, generating an emotional state transition map for the client. Psychological report generation automatically integrates conversation records, emotion mapping data, psychological scale scores, and other psychological professional indicators based on a structured output template to generate a phased consultation report.

[0018] The multi-task mechanism system includes a digital human front-end communication task thread, a large model question-and-answer task thread, a subtitle task, an instruction task, and a TTS processing task; The digital human front-end communication task thread uses the WebSocket protocol to receive the user question character stream from the front end; the large model question-and-answer task thread is used to process the received front-end user questions; the subtitle task is used to send the subtitle character stream to the front end, and combined with the voice playback event, the two are nearly synchronized; the instruction task is used to receive instructions sent by the front end, such as "stop playing voice", which can forcibly interrupt the voice playback task; the TTS processing task receives the character stream returned by the large model, performs semantic truncation, sends the character stream to the online TTS model, and then receives the voice chunks using a streaming transmission method.

[0019] The local voice player uses seamless splicing of voice stream chunks and can pre-allocate a memory pool to store audio chunks generated by TTS. The player can also read the buffer pool through a dual-pointer polling mechanism to achieve seamless splicing between segments.

[0020] Preferably, the online large model includes a role-playing text large model and a speech synthesis large model; the role-playing text large model is designed with a psychological counselor template, enabling long-term online memory retention, as well as safety and ethical filtering; the speech synthesis large model adopts real-time streaming TTS technology based on CosyVoice2.0, which features accurate pronunciation, low latency, and natural emotion, and also supports voice cloning of specific characters.

[0021] Preferably, the local speech recognition system includes a dual-channel recognition engine and other functions. The dual-channel recognition engine includes a first channel and a second channel. The first channel is in streaming mode, and the second channel is in offline mode. The first channel transmits the audio stream in real time via WebSocket to achieve millisecond-level real-time subtitle feedback in PCM format. The second channel triggers offline ASR after the audio stream is complete, and adopts an error correction mechanism to improve the final recognition accuracy. Preferably, other functions include VAD voice effective detection and hot word interruption function.

[0022] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0023] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A real-time voice-interactive digital human intelligent agent based on a localized multi-task workflow, characterized in that: It includes a local working system, an online large model, a local speech recognition system, a multi-task mechanism system, and a local speech player; the local working system includes a short-term memory layer, intent recognition routing, a RAG local knowledge base, crisis recognition, emotion graph, and psychological report generation; the multi-task mechanism system includes a digital human front-end communication task thread, a large model question-and-answer task thread, a subtitle task, an instruction task, and a TTS processing task.

2. The real-time voice interaction digital human intelligent agent based on localized multi-task workflow as described in claim 1, characterized in that, The online large-scale model includes a role-playing text large-scale model and a speech synthesis large-scale model.

3. The real-time voice interaction digital human intelligent agent based on localized multi-task workflow as described in claim 1, characterized in that, The local speech recognition system includes a dual-channel recognition engine and other functions.

4. The real-time voice interaction digital human agent based on localized multi-task workflow as described in claim 3, characterized in that, Other features include VAD voice valid detection and hot word interruption.