Large-model-driven method for real-time interaction with digital human

By using a large model-driven real-time interaction method for digital humans and leveraging pre-built video libraries and classifier technology, the latency problem in the real-time question-and-answer process of digital humans has been solved, resulting in a smoother interactive experience and a higher level of immersion.

WO2026098122A1PCT designated stage Publication Date: 2026-05-15WEIKE ZHIJIAN (FOSHAN) TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
WEIKE ZHIJIAN (FOSHAN) TECHNOLOGY CO LTD
Filing Date
2025-09-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, the real-time question-and-answer process of digital humans is delayed, which reduces the user's immersion and simulation level, making it difficult to achieve true real-time response.

Method used

A large-model-driven real-time interactive approach for digital humans is adopted. By using a pre-built video library and classifier technology, and utilizing pre-recorded digital human videos and summary videos, the generation latency is reduced and a smooth interactive experience is provided.

Benefits of technology

It achieves a smoother real-time question-and-answer interactive experience for digital humans, reduces the sense of distortion, and enhances the user's immersion and simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025125834_15052026_PF_FP_ABST
    Figure CN2025125834_15052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of digital humans. Disclosed is a large-model-driven method for real-time interaction with a digital human, comprising the following steps: S1, inputting a question text by means of a human-machine dialog frontend and sending the question text to a large model backend; S2, the large model backend performing large model inference on the basis of the question text and a knowledge base to generate an answer summary and a full answer and respectively sending the answer summary and the full answer to a digital human backend and the human-machine dialog frontend for display; S3, a classifier generating a plurality of classification labels on the basis of the answer summary, and the digital human backend performing retrieval and matching on a preset video library on the basis of the classification labels to obtain a target preset video and sending the target preset video to the human-machine dialog frontend; and S4, a digital human processing engine generating a digital human summary video on the basis of the answer summary and pushing the digital human summary video to a player, and the player sequentially playing the target preset video and the digital human summary video. Smoother experience of real-time question-and-answer interaction with a digital human is finally achieved, thereby reducing loss of realism in the experience.
Need to check novelty before this filing date? Find Prior Art

Description

A large-model-driven real-time interaction method for digital humans Technical Field

[0001] This invention relates to the field of digital human technology, and in particular to a real-time interactive method for digital humans driven by a large model. Background Technology

[0002] Digital humans, also known as virtual digital humans, are digital images created using computer graphics (CG) and artificial intelligence technologies that closely resemble human appearances. Digital virtual humans rely on display devices for their existence, not only visually mimicking human physical features but also being given specific character identities through technology to achieve more realistic emotional interaction. In nature, digital humans possess the following characteristics: (1) They have a human appearance, possessing specific physical features, gender, and personality traits; (2) They possess human behavior, having the ability to express themselves through language, facial expressions, and body language; (3) They possess human thought, having the ability to recognize the external environment and interact with others.

[0003] In recent years, the rapid development of large-scale models has injected new vitality into the development of digital humans. Large-scale model technology provides digital humans with almost unlimited training parameters and autonomous generation capabilities, enabling them to achieve richer and more complex motor expressions, including facial expressions and body movements. Combined with rich knowledge base models, digital humans can become experts in any field at any time, providing users with 24 / 7 uninterrupted, highly realistic professional consultation and services.

[0004] For example, patent document CN118377865A provides a real-time question-answering method and system for digital humans based on large models and deep learning. The method includes the following steps: generating silent audio; obtaining user questions; when the user questions are obtained, generating corresponding question-and-answer text using a large model, and then converting it into question-and-answer audio of several standard durations; when the user questions are not obtained, generating silent audio and using it cyclically; based on the question-and-answer audio, the silent audio, and the corresponding face image, using a deep model to calculate and render the corresponding face image frame; processing the question-and-answer audio, the silent audio, and the face image frame, inputting them into the corresponding channel to obtain a real-time rendered lip-syncing face video; and using real-time driving technology to push the lip-syncing face video to the user's end. Ultimately, this enables users to experience virtual reality products in real time and, by leveraging the characteristics of large models, generates more reasonable interactive templates, increasing the product's flexibility.

[0005] However, the above solution still has shortcomings. Specifically, the solution has multiple delays in several steps before generating the final digital human output, including the generation of question-and-answer text, the conversion of question-and-answer text to audio, the rendering of facial images based on the audio, and the creation of a digital human facial video with dynamic lip movements. The total delay caused by these multiple steps prevents the generated digital human from truly achieving "real-time responsiveness," making it difficult for users to maintain a high level of immersion and realism, and easily causing them to lose the smoothness and realism of interacting with the digital human. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides a large model-driven real-time digital human interaction method, which achieves a smoother real-time question-and-answer interaction experience for digital humans and reduces the distortion of the experience.

[0007] The technical solution of this invention is implemented as follows:

[0008] A large-model-driven real-time interaction method for digital humans includes a human-computer dialogue front-end, a large-model back-end, a digital human back-end, and a digital human front-end;

[0009] The large model backend has a knowledge base; the digital human backend has a digital human processing engine, a pre-set video library, and a classifier; the pre-set video library stores multiple pre-set videos, which are pre-recorded digital human videos; each pre-set video is associated with at least one classification tag; there are multiple types of classification tags; the digital human frontend has a player.

[0010] The method includes the following steps:

[0011] S1. The user inputs a question text through the human-computer dialogue front end and sends the question text to the large model back end;

[0012] S2. The large model backend performs large model reasoning based on the question text and the knowledge base to generate an answer summary and a complete answer; the complete answer is sent to the human-computer dialogue frontend for display; the answer summary is sent to the digital human backend;

[0013] S3. The classifier performs classification tag division processing based on the answer summary to generate at least one classification tag; the digital human backend performs retrieval and matching on the preset video library based on the generated classification tag to obtain the target preset video; the target preset video is sent to the human-computer dialogue frontend;

[0014] The pre-set videos mainly provide general, basic explanations and answers to related questions. These materials cover a variety of common question categories and scenarios; this category is not fixed and can be configured according to different vertical fields.

[0015] S4. The player plays the target preset video; the digital human processing engine generates a digital human summary video based on the answer summary and pushes the digital human summary video to the player; after the target preset video finishes playing, the digital human summary video continues to play.

[0016] The human-computer dialogue front-end and digital human front-end mentioned above do not specifically refer to two independent devices, but should be understood primarily as interfaces, different parts of the interactive interface on the same terminal, or other settings with similar logical interaction relationships. Of course, they can also be understood as two independent devices. Similarly, the large model back-end and digital human back-end can be understood as two independent modules of the same device / cloud platform, or two different devices / cloud platforms.

[0017] By playing pre-recorded videos, the latency in generating digital human videos and complete answers can be bridged. Users can learn about general explanations of relevant questions or some pre-stored answers related to specific vertical fields through pre-recorded videos, and then browse the digital human summary video to gain a comprehensive understanding of the final complete answer in advance. Ultimately, this achieves a smoother real-time question-and-answer interaction experience with digital humans and reduces distortion.

[0018] As a further optimization of the above solution, the preset video includes general-purpose videos and vertical-domain videos; in step S3, the target preset video is one of the general-purpose videos, or one of the vertical-domain videos, or any number of general-purpose videos and any number of vertical-domain videos spliced ​​together.

[0019] As a further optimization of the above scheme, in step S3, the classifier is a semantic classifier; the digital human processing engine predicts the required generation time of the digital human summary video based on the answer summary; the digital human backend generates the target preset video based on the classification tag matching in the preset video library, and the total duration of the target preset video is the same as the generation time.

[0020] The generation duration is the same as the total duration of the target preset video. The purpose is to make the playback of the final target preset video and the digital human summary video as seamless as possible, so as to achieve a smoother real-time question-and-answer interactive experience for digital humans and reduce the distortion of the experience.

[0021] As a further optimization of the above scheme, in step S2, the generation of the answer summary and the complete answer are processed in parallel; or the large model backend first generates the answer summary and sends it to the digital human backend, and then generates the complete answer and sends it to the human-computer dialogue frontend.

[0022] The ultimate goal is to prioritize displaying digital human summary videos, allowing users to focus on and understand the answer summary before browsing the complete answer.

[0023] As a further optimization of the above solution, the digital human backend performs text segmentation on the answer summary to obtain multiple script sequence texts; the digital human processing engine includes multiple digital human processing threads; the digital human processing threads correspond one-to-one with the script sequence texts;

[0024] The digital human processing thread generates behavioral expression video sequences and audio sequences based on the script sequence text, and performs lip-syncing on the obtained behavioral expression video sequences and audio sequences to generate digital human summary video clips.

[0025] Multiple digital human summary video clips are played sequentially; multiple digital human processing threads process them in parallel.

[0026] As a further optimization of the above solution, the digital human backend also stores an emotion dataset and a behavior dataset; the digital human processing engine analyzes the answer summary based on the emotion dataset and the behavior dataset to obtain emotion tags and behavior tags; the digital human processing engine generates the behavior expression video sequence based on the text corresponding to the answer summary, the emotion tags and the behavior tags, and generates the audio sequence based on the text corresponding to the answer summary and the emotion tags.

[0027] As a further optimization of the above solution, the player plays a carousel video when it does not receive the preset video or the digital human summary video.

[0028] The carousel can contain any video. Furthermore, the carousel displays videos of digital humans without any movement or sound.

[0029] As a further optimization of the above solution, the human-computer dialogue front end is also equipped with a voice input module and a pre-set question bank;

[0030] The user can input the question text by typing, or by inputting voice, or by selecting a question from the preset question bank as the question text.

[0031] The voice input module converts the input voice into text, which serves as the question text.

[0032] As a further optimization of the above solution, in step S4, the digital human backend directly pushes the digital human summary video to the player from memory.

[0033] Pushing from memory reduces read / write overhead and the total time required for users to go from inputting questions to seeing the digital human summary video.

[0034] Compared with the prior art, the present invention achieves the following beneficial effects:

[0035] This invention provides a large-model-driven real-time interactive method for digital humans. By playing pre-set videos, the delay in generating digital human videos and complete answers can be bridged. Users can learn about general explanations of relevant questions or some pre-stored answers related to specific vertical fields through the pre-set videos, and then browse the digital human summary video to gain a comprehensive understanding of the final complete answer in advance. Ultimately, this achieves a smoother real-time question-and-answer interactive experience with digital humans and reduces distortion. Attached Figure Description

[0036] Figure 1 is a flowchart illustrating a large-model-driven real-time interaction method for digital humans provided in an embodiment of the present invention;

[0037] Figure 2 is a schematic diagram of the process of matching target preset videos provided in an embodiment of the present invention;

[0038] Figure 3 is a schematic diagram of the process for generating digital human summary videos provided in an embodiment of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0040] As shown in Figures 1 to 3, this embodiment provides a real-time interactive method for digital humans driven by a large model, including a human-computer dialogue front-end, a large model back-end, a digital human back-end, and a digital human front-end; in this embodiment, the human-computer dialogue front-end and the digital human front-end are different display parts of the interactive interface on the same terminal.

[0041] The large model has a knowledge base on the backend;

[0042] The digital human backend is equipped with a digital human processing engine, a pre-built video library and a classifier, and also stores emotion datasets and behavior datasets; in this embodiment, the classifier is a semantic classifier.

[0043] The preset video library stores multiple preset videos, which are pre-recorded digital human video materials; each preset video is associated with at least one category tag; there are multiple types of category tags;

[0044] The pre-set videos mainly provide general, basic explanations and answers to related questions. These materials cover a variety of common question categories and scenarios; this category is not fixed and can be configured according to different vertical fields.

[0045] In this embodiment, the preset videos include general videos and vertical domain videos.

[0046] The digital human's front end is equipped with a player. In this embodiment, when no preset video or digital human summary video is received, the player plays a carousel video. The carousel video can be any video. In this embodiment, the carousel video displays a digital human video without any movement or sound.

[0047] The method includes the following steps:

[0048] S1. The user inputs a question text through the human-computer dialogue front end and sends the question text to the back end of the large model; in this embodiment, the human-computer dialogue front end is also equipped with a voice input module and a pre-set question bank; the user has multiple input methods.

[0049] Users can obtain the question text by typing it in using the keyboard;

[0050] Users can also input voice, which will be converted into text by the voice input module and used as the question text;

[0051] Users can also quickly select questions from a pre-set question bank as their question text.

[0052] S2. The large model backend performs large model reasoning based on the question text and knowledge base to generate an answer summary and a complete answer; the complete answer is sent to the human-computer dialogue frontend for display; the answer summary is sent to the digital human backend; in this embodiment, the generation process of the answer summary and the complete answer is processed in parallel.

[0053] S3. The classifier performs classification tag segmentation based on the answer summary, generating at least one classification tag. In this embodiment, the digital human processing engine predicts the required generation duration of the digital human summary video based on the answer summary. The digital human backend obtains the target preset video from the preset video library by matching the classification tags provided by the classifier, and the total duration of the target preset video is the same as the generation duration. In this embodiment, the target preset video is a general-purpose video or a vertical domain video. The target preset video is then sent to the human-computer dialogue frontend.

[0054] The generation duration is the same as the total duration of the target preset video. The purpose is to ensure that the playback of the final target preset video and the digital human summary video are played as seamlessly as possible, and to ensure the playback of the digital human summary video as quickly as possible.

[0055] S4. The player plays the target preset video; the digital human processing engine generates a digital human summary video based on the answer summary.

[0056] In this embodiment, the digital human backend segments the answer summary into multiple script sequence texts; the digital human processing engine includes multiple digital human processing threads; and there is a one-to-one correspondence between the digital human processing threads and the script sequence texts.

[0057] Each digital human processing thread analyzes the script sequence text based on the emotion dataset and the behavior dataset to obtain emotion tags and behavior tags. The digital human processing thread generates a video sequence of behavioral expressions based on the script sequence text, emotion tags, and behavior tags, generates an audio sequence based on the script sequence text and emotion tags, and performs lip-syncing on the obtained video sequence of behavioral expressions and audio sequence to generate a digital human summary video clip.

[0058] Multiple digital human processing threads process the data in parallel, resulting in multiple digital human summary video clips.

[0059] The digital human processing engine streams digital human summary video clips to the player; in this embodiment, the digital human backend directly streams the digital human summary video clips to the player from memory. Streaming from memory reduces read / write overhead and the total time required for the user to go from inputting a question to seeing the digital human summary video.

[0060] After the target preset video finishes playing, digital human summary video clips will be played in sequence, that is, digital human summary videos will be played.

[0061] By playing pre-recorded videos, the latency in generating digital human videos and complete answers can be bridged. Users can learn about general explanations of relevant questions or some pre-stored answers related to specific vertical fields through pre-recorded videos, and then browse the digital human summary video to gain a comprehensive understanding of the final complete answer in advance. Ultimately, this achieves a smoother real-time question-and-answer interaction experience with digital humans and reduces distortion.

[0062] Based on the disclosure and teachings of the foregoing specification, those skilled in the art can make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. Furthermore, although some specific terms are used in this specification, these terms are only for convenience of explanation and do not constitute any limitation on the present invention.

Claims

1. A large-model-driven real-time interaction method for digital humans, characterized in that, This includes the human-computer dialogue front-end, the large model back-end, the digital human back-end, and the digital human front-end; The large model backend has a knowledge base; the digital human backend has a digital human processing engine, a pre-set video library, and a classifier; the pre-set video library stores multiple pre-set videos, which are pre-recorded digital human videos; each pre-set video is associated with at least one classification tag; there are multiple types of classification tags; the digital human frontend has a player. The method includes the following steps: S1. The user sends a question text to the large model backend through the human-computer dialogue frontend; S2. The large model backend performs large model reasoning based on the question text and the knowledge base to generate an answer summary and a complete answer; the complete answer is sent to the human-computer dialogue frontend for display; the answer summary is sent to the digital human backend; S3. The classifier performs classification tag division processing based on the answer summary to generate at least one classification tag; the digital human backend performs retrieval and matching on the preset video library based on the generated classification tag to obtain the target preset video; the target preset video is sent to the human-computer dialogue frontend; S4. The digital human processing engine generates a digital human summary video based on the answer summary, and pushes the digital human summary video to the player; The player sequentially plays the target preset video and the digital human summary video.

2. The method for real-time interaction of digital humans driven by a large model according to claim 1, characterized in that, The preset videos include general-purpose videos and vertical-domain videos; in step S3, the target preset video is one of the general-purpose videos, or one of the vertical-domain videos, or any number of general-purpose videos and any number of vertical-domain videos spliced ​​together.

3. The method for real-time interaction of digital humans driven by a large model according to claim 1, characterized in that, In step S3, the classifier is a semantic classifier; the digital human processing engine predicts the required generation time of the digital human summary video based on the answer summary; the digital human backend generates the target preset video based on the classification label matching in the preset video library, and the total duration of the target preset video is the same as the generation time.

4. The method for real-time interaction of digital humans driven by a large model according to claim 1, characterized in that, In step S2, the generation of the answer summary and the complete answer are processed in parallel; or the large model backend first generates the answer summary and sends it to the digital human backend, then generates the complete answer and sends it to the human-computer dialogue frontend.

5. The method for real-time interaction of digital humans driven by a large model according to claim 1, characterized in that, The digital human backend segments the answer summary into multiple script sequence texts; the digital human processing engine includes multiple digital human processing threads; each digital human processing thread corresponds one-to-one with the script sequence text. The digital human processing thread generates behavioral expression video sequences and audio sequences based on the script sequence text, and performs lip-syncing on the obtained behavioral expression video sequences and audio sequences to generate digital human summary video clips. Multiple digital human summary video clips are played sequentially; multiple digital human processing threads process them in parallel.