Digital human live broadcast interaction method and system based on large model
By using a large model to analyze user interactive content and generate response content in the digital human live interactive system, the problem that existing systems cannot achieve real-time two-way communication and multi-modal interaction is solved, and a high-quality real-time interactive experience is achieved.
Patent Information
- Application Number
- CN202510171658.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-09
AI Technical Summary
The existing digital live broadcast interactive system cannot achieve true real-time two-way communication, and lacks multi-modal interactive experiences between vision and voice, resulting in mechanical and human touch of interactive content, and there are problems such as poor interaction quality and poor real-time performance.
The digital human live broadcast interaction method based on the big model is adopted. The user's interactive content is analyzed through the big model, reply instructions are generated, and matching audio content and image sequences with synchronous mouths are generated according to the instructions. The timing alignment process is performed and the stream is pushed to the live broadcast platform to dynamically adjust the live broadcast content of the digital human anchor.
It realizes high-quality real-time two-way interaction, enhances the user's experience, provides multi-modal interactive experience between vision and voice, and improves the quality and real-timeness of interactive content.
Smart Images

Figure CN119967197A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a digital human live broadcast interaction method and system based on a large model. Background Art
[0002] With the continuous advancement of artificial intelligence technology, digital human anchors, as a new type of interactive medium, have gradually emerged in the live broadcast field. Digital human technology combines a variety of cutting-edge technologies such as computer graphics, artificial intelligence, and natural language processing. It can create virtual anchors with realistic appearance and intelligent interactive capabilities. With the development of large models, digital humans have stronger dialogue generation and understanding capabilities.
[0003] Existing live broadcast interactions of digital humans usually interact with fans through pre-recorded interactive videos, which cannot achieve true real-time two-way communication. Some platforms interact with digital humans through text chat. Although this improves interactivity, it lacks the multimodal interactive experience of vision and voice, and it is difficult to meet users' demand for high-quality interaction. It uses basic natural language processing technology to generate automatic replies, but lacks understanding and intelligent response to fans' real-time video input. The interactive content is relatively mechanical and lacks human touch. There are technical problems such as poor interaction quality and poor real-time performance. Summary of the invention
[0004] In view of the above analysis, the embodiments of the present invention aim to provide a digital human live interactive method and system based on a large model, so as to solve one or more of the above problems existing in the prior art.
[0005] The object of the present invention is achieved in that:
[0006] The first embodiment of the present invention provides a digital human live interactive method based on a large model, comprising:
[0007] Configure digital human anchors according to user intentions;
[0008] Obtain interactive requests initiated by users and transmit interactive content in real time;
[0009] Analyze the interactive content through a large model to generate a reply instruction;
[0010] Generate matching audio content according to the reply instruction, and generate an image sequence with synchronized lip sync according to the audio content;
[0011] The image sequence and audio content are processed for time sequence alignment and then pushed to the live broadcast platform to dynamically adjust the live broadcast content of the digital human anchor.
[0012] Furthermore, configuring the digital human anchor according to the user's intention includes: providing a visual interface for the user to customize the elements of the live broadcast room including the appearance attributes, sound characteristics and behavioral logic of the digital human anchor; creating a digital human anchor according to the user-defined live broadcast room elements, configuring the corresponding live broadcast content and streaming it to the live broadcast platform.
[0013] Furthermore, configuring a digital human anchor according to user intentions also includes: predicting user preferences according to the user's historical behavior data, and recommending a digital human anchor that meets the user's preferences.
[0014] Furthermore, the interaction requests initiated by users are obtained, and the interaction content is transmitted in real time, including: authenticating the identities of each user requesting interaction, eliminating interaction requests from users who have not passed the verification, prioritizing the interaction requests that have passed the verification using distributed queue management, and dynamically adjusting the queue priority based on the user interaction value weight, and collecting and transmitting the interaction content of the highest priority user in the queue in real time, wherein the interaction content includes video interaction, voice interaction, and barrage text interaction.
[0015] Furthermore, the interactive content is analyzed by a large model to generate reply instructions, including: extracting and fusing key features in video interaction, voice interaction and barrage text interaction by a multimodal large model, and identifying the user's emotional state and intention based on the fused key features; and generating reply instructions based on the user's emotional state and intention combined with prompt words.
[0016] Furthermore, generating matching audio content according to the reply instruction and generating an image sequence with synchronized lip movements according to the audio content includes: using a text-to-speech engine to generate audio content according to the reply instruction, and using a phoneme-level lip movement matching algorithm to extract phoneme information corresponding to the audio content; driving a three-dimensional lip shape skeleton model based on the phoneme information, using an optical flow compensation algorithm to smooth the lip movement transition between adjacent frames, and generating an image sequence of digital human anchor facial animation through a neural network rendering engine.
[0017] Furthermore, the timing alignment processing of the image sequence and audio content includes: embedding a synchronization timestamp signal in the audio content; matching the playback rhythm of the audio content and the image sequence through a dynamic time warping algorithm; and setting a buffer threshold to automatically correct the deviation between the audio content and the image sequence.
[0018] The second embodiment of the present invention provides a digital human live interactive system based on a large model, comprising:
[0019] A configuration module, used to configure the digital human anchor according to the user's intention;
[0020] The interactive module is used to obtain the interactive request initiated by the user and transmit the interactive content in real time;
[0021] An analysis module, used for analyzing the interactive content through a large model to generate a reply instruction;
[0022] A response generation module, used to generate matching audio content according to the reply instruction, and to generate an image sequence with synchronized lip movements according to the audio content;
[0023] The streaming module is used to perform time alignment processing on the image sequence and audio content and then stream them to the live broadcast platform to dynamically adjust the live broadcast content of the digital human anchor.
[0024] A third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the large model-based digital human live interactive method described in any embodiment is implemented.
[0025] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for live interactive digital human based on a large model described in any one of the embodiments is implemented.
[0026] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0027] The digital human live interactive method based on a large model provided by the present invention, BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0029] Figure 1 A flowchart of a digital human live interactive method based on a large model provided in Example 1 of the present invention;
[0030] Figure 2 A schematic diagram of a digital human live interactive system based on a large model provided in Example 2 of the present invention;
[0031] Figure 3 This is a schematic diagram of the electronic device architecture provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments disclosed in this disclosure can be combined, separated, interchanged and / or rearranged with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0033] Example 1
[0034] A specific embodiment of the present invention, as Figure 1 As shown, a digital human live interactive method based on a large model is disclosed, comprising the following steps:
[0035] S1. Configure the digital human anchor according to the user's intention;
[0036] In this embodiment, step S1 specifically includes:
[0037] S101, providing a visual interface for users to customize the elements of the live broadcast room including the appearance attributes, voice characteristics and behavior logic of the digital human anchor;
[0038] S102: Create a digital human anchor according to the user-defined live broadcast room elements, configure the corresponding live broadcast content and push it to the live broadcast platform.
[0039] Specifically, a drag-and-drop editor is provided for users, through which they can customize the digital human's appearance attributes (such as skin color and proportions of facial features), sound characteristics (selecting personalized sound models through the sound library or uploading samples to train), and behavioral logic templates (such as setting response rules for specific scenarios). After the user configuration is completed, the system calls the Unity or Unreal Engine rendering engine to generate a digital human 3D model, and pushes the initial live broadcast content to the live broadcast platform through the FFmpeg encoding module.
[0040] In some embodiments, step S1 may also predict user preferences based on the user's historical behavior data and recommend a digital human anchor that meets the user's preferences.
[0041] Exemplarily, the system records the user's historical interaction data (such as the frequency of live broadcasts and preferred interaction types), builds a user portrait through a collaborative filtering algorithm, and recommends digital human anchors with similar styles based on the user portrait.
[0042] S2, obtaining the interaction request initiated by the user and transmitting the interaction content in real time;
[0043] In this embodiment, step S2 specifically includes:
[0044] S201, authenticate the identity of each user who requests interaction, and reject interaction requests from users who fail the authentication;
[0045] S202, using distributed queue management to prioritize verified interaction requests, and dynamically adjusting queue priorities based on user interaction value weights;
[0046] S203, collecting and transmitting in real time the interactive content of the user with the highest priority in the queue, wherein the interactive content includes video interaction, voice interaction and barrage text interaction.
[0047] Specifically, after the user initiates an interaction request, the system calls the face recognition or mobile phone number verification interface to verify the identity, eliminate anonymous or illegal accounts, and the verified requests enter the distributed message queue (such as Kafka), and dynamically adjust the priority according to the user's interaction value weight (such as fan level, historical consumption amount). After the user with the highest priority in the queue is selected, the system establishes a two-way video channel through WebRTC to collect their video, voice and barrage data in real time.
[0048] In some embodiments, through a real-time notification mechanism, such as WebSocket, the digital human host sends interaction invitations to multiple users, and the users can choose whether to interact. For users who accept the interaction invitation, the interaction is carried out in parallel by setting the number of interactions and priority strategy.
[0049] S3, analyzing the interactive content through a large model to generate a reply instruction;
[0050] In this embodiment, step S3 specifically includes:
[0051] S301, extracting and fusing key features from video interaction, voice interaction, and bullet screen text interaction through a multimodal large model, and identifying the user's emotional state and intention based on the fused key features;
[0052] S302: Generate a reply instruction based on the user's emotional state and intention combined with the prompt word.
[0053] Specifically, after receiving the user's interactive content, pre-processing such as noise reduction, resolution adjustment, and format conversion is first performed to adapt to the input requirements of the large model. Then, a large model analysis engine such as GhatGlm is used to extract features such as expressions, actions, and environments in video interactions, as well as features about emotional states and intentions in voice interactions and barrage text interactions. Multiple features are mapped into the same intent vector, and then the intent vector and prompt words are combined to generate reply instructions in response to the user's interactive content, in order to instruct the digital human anchor to adjust the live broadcast content in real time and interact with users.
[0054] S4, generating matching audio content according to the reply instruction, and generating an image sequence with synchronized lip movements according to the audio content;
[0055] In this embodiment, step S4 specifically includes:
[0056] S401, using a text-to-speech engine to generate audio content according to the reply instruction, and using a phoneme-level lip matching algorithm to extract phoneme information corresponding to the audio content;
[0057] S402, driving a three-dimensional lip shape skeleton model based on the phoneme information, using an optical flow compensation algorithm to smooth the lip shape transition between adjacent frames, and generating an image sequence of the digital human anchor's facial animation through a neural network rendering engine.
[0058] Specifically, after the large model generates a reply instruction, the next step is to adjust the audio and video content of the digital human anchor according to the instruction to respond to the user's interactive content. First, the deep learning-based TTS model is used to generate audio content and extract phoneme information, namely the phoneme sequence, from the audio content; then, according to the standard phoneme-lip shape correspondence table, the phoneme sequence is converted into a smooth lip motion curve to drive the three-dimensional lip shape skeleton model, and the optical flow compensation algorithm is used to reduce lip shape jitter, making the lip shape changes smoother and more natural; finally, the processed lip shape data is converted into an image sequence of the digital human anchor's facial animation through the neural network rendering engine.
[0059] S5. After performing time alignment processing on the image sequence and audio content, the image sequence and the audio content are pushed to the live broadcast platform to dynamically adjust the live broadcast content of the digital human anchor.
[0060] In this embodiment, step S5 specifically includes:
[0061] S501, embedding a synchronization timestamp signal in the audio content;
[0062] S502, matching the playing rhythm of the audio content and the image sequence through a dynamic time warping algorithm;
[0063] S503: Set a buffer threshold to automatically correct the deviation between the audio content and the image sequence.
[0064] Specifically, a frame-level timestamp is used to mark the correspondence between the audio content and the image sequence generated in step S4 to facilitate subsequent synchronization processing; a dynamic time warping algorithm is then used to calculate the distance between frames and optimize the time alignment curve to adjust the playback rhythm of the audio content and the image sequence; the audio and video delay is then detected through an algorithm such as the Kalman filter, and the playback rate is dynamically adjusted. When the delay exceeds the buffer threshold, the time stretching technology is used to correct the error, thereby achieving audio and video synchronization; finally, the audio content and image sequence that have undergone timing alignment processing are pushed to the live broadcast platform through streaming media technology.
[0065] In some embodiments, it also includes obtaining user feedback on the live broadcast content adjusted by the digital human anchor in real time, and looping through steps S3-S5 to achieve continuous two-way interaction, and recording the data of each interaction for subsequent analysis and optimization, thereby improving the intelligence and personalization capabilities of the system.
[0066] Compared with the existing technology, the digital human live interactive method based on a large model provided in this embodiment uses a high-performance large model to analyze the transmitted interactive content in real time, deeply understand the user's emotional state and intention, thereby generating high-quality reply content and conducting high-quality interactions with the user. It also enhances the user experience by processing audio and video content through time alignment.
[0067] Example 2
[0068] This embodiment provides a digital human live interactive system based on a large model, such as Figure 2 As shown, including:
[0069] A configuration module, used to configure the digital human anchor according to the user's intention;
[0070] The interactive module is used to obtain the interactive request initiated by the user and transmit the interactive content in real time;
[0071] An analysis module, used for analyzing the interactive content through a large model to generate a reply instruction;
[0072] A response generation module, used to generate matching audio content according to the reply instruction, and to generate an image sequence with synchronized lip movements according to the audio content;
[0073] The streaming module is used to perform time alignment processing on the image sequence and audio content and then stream them to the live broadcast platform to dynamically adjust the live broadcast content of the digital human anchor.
[0074] Example 3
[0075] This embodiment provides an electronic device, such as Figure 3 As shown, it includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the digital human live broadcast interaction method based on a large model as described in any of the above embodiments is implemented.
[0076] Example 4
[0077] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the digital human live broadcast interaction method based on a large model as described in any of the above embodiments is implemented.
[0078] Computer readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0079] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0080] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0081] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A digital human live interactive method based on a large model, characterized in that: include: Configure digital human anchors according to user intentions; Obtain interactive requests initiated by users and transmit interactive content in real time; Analyze the interactive content through a large model to generate a reply instruction; Generate matching audio content according to the reply instruction, and generate an image sequence with synchronized lip sync according to the audio content; The image sequence and audio content are processed for time sequence alignment and then pushed to the live broadcast platform to dynamically adjust the live broadcast content of the digital human anchor.
2. The method for live interactive broadcast of digital human based on large model according to claim 1 is characterized in that: The configuring of the digital human anchor according to the user's intention includes: Provide a visual interface for users to customize the elements of the live broadcast room, including the appearance attributes, sound characteristics and behavior logic of the digital human anchor; Create a digital human anchor based on the user-defined live broadcast room elements, configure the corresponding live broadcast content and push it to the live broadcast platform.
3. The method for live interactive broadcast of digital human based on large model according to claim 2 is characterized in that: The configuration of the digital human anchor according to the user's intention also includes: Predict user preferences based on their historical behavior data and recommend digital human anchors that meet their preferences.
4. The method for live interactive broadcast of digital human based on large model according to claim 1 is characterized in that: Obtain interaction requests initiated by users and transmit interactive content in real time, including: Authenticate each user who requests interaction, eliminate interaction requests from users who have not passed verification, prioritize verified interaction requests using distributed queue management, dynamically adjust queue priority based on user interaction value weight, and collect and transmit interaction content from the highest priority user in the queue in real time, including video interaction, voice interaction, and barrage text interaction.
5. The method for live interactive broadcast of digital human based on large model according to claim 4 is characterized in that: Analyzing the interactive content by a large model to generate a reply instruction includes: Through the multimodal large model, key features of video interaction, voice interaction and barrage text interaction are extracted and integrated, and the user's emotional state and intention are identified based on the integrated key features; Generate reply instructions based on the user's emotional state and intention combined with prompt words.
6. The method for live interactive broadcast of digital human based on large model according to claim 5 is characterized in that: Generating matching audio content according to the reply instruction, and generating an image sequence with synchronized lip sync according to the audio content, comprises: Using a text-to-speech engine to generate audio content according to the reply instruction, and using a phoneme-level lip matching algorithm to extract phoneme information corresponding to the audio content; The three-dimensional lip shape skeleton model is driven based on the phoneme information, an optical flow compensation algorithm is used to smooth the lip shape transition between adjacent frames, and an image sequence of the digital human anchor's facial animation is generated through a neural network rendering engine.
7. The method for live interactive broadcast of digital human based on large model according to claim 1 is characterized in that: The performing time alignment processing on the image sequence and the audio content includes: embedding a synchronization time stamp signal in the audio content; Match the playback rhythm of audio content and image sequences through dynamic time warping algorithm; Set the buffer threshold to automatically correct the deviation between the audio content and the image sequence.
8. A digital human live interactive system based on a large model, characterized in that: The system comprises: A configuration module, used to configure the digital human anchor according to the user's intention; The interactive module is used to obtain the interactive request initiated by the user and transmit the interactive content in real time; An analysis module, used for analyzing the interactive content through a large model to generate a reply instruction; A response generation module, used to generate matching audio content according to the reply instruction, and to generate an image sequence with synchronized lip movements according to the audio content; The streaming module is used to perform time alignment processing on the image sequence and audio content and then stream them to the live broadcast platform to dynamically adjust the live broadcast content of the digital human anchor.
9. An electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the large model-based digital human live broadcast interaction method as described in any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the digital human live interactive method based on a large model as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Dynamic video synthesis method and device based on instruction sequence and instruction editing system
CN120676223A
A method, apparatus, and instruction editing system for dynamic video synthesis based on instruction sequences.
CN120676223B
Agricultural product personalized recommendation digital human live broadcast interaction system based on large model
CN121462784A
Real-time video processing method and system based on artificial intelligence
CN121644932A
Digital human interaction system, method and device, storage medium and program product
CN122199762A