Digital human live broadcast real-time interaction method, system, electronic device and program product
By generating real-time response audio and video and inserting them into keyframes of digital human live streams, the problem of lack of real-time interaction in digital human live streams is solved, improving the live stream effect and user engagement, and promoting the transformation and integrated application of the digital economy.
Patent Information
- Application Number
- CN202511308163.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Digital human live streaming lacks real-time interactive capabilities, which weakens the live streaming effect and makes it difficult to achieve the expected dissemination or conversion goals.
By generating response audio and video in response to user interactions, and using digital human templates to insert audio and video into keyframes, we can ensure instant response and maintain the smoothness of the live stream.
It enhances the interactivity and smoothness of live streaming, strengthens user engagement, promotes digital transformation and commercial value, and expands industrial integration applications.
Smart Images

Figure CN120812308B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital person live broadcast, in particular to a real-time interaction method and system for digital person live broadcast, an electronic device and a program product. BACKGROUND
[0002] Digital economy industry refers to a series of economic activities taking digital knowledge and information as core production factors, taking modern information network as key carrier, and driven by information communication technology to improve efficiency and optimize economic structure, including intelligent transportation, Internet finance, mobile payment, artificial intelligence and other fields. Among them, digital person live broadcast is an important branch of artificial intelligence field.
[0003] At present, digital economy industry is accelerating evolution, and digital person technology is developing rapidly. More and more live broadcast scenarios begin to use digital person for live broadcast. However, the live broadcast content of digital person live broadcast is mostly pre-recorded loop material. When facing real-time interaction information of users, it often lacks immediate response capability, thus weakening the overall effect of live broadcast. SUMMARY
[0004] To solve or partially solve the problems in the related art, the present application provides a real-time interaction method, system, electronic device and program product for digital person live broadcast, which can respond immediately when facing real-time interaction information of users, ensure the overall effect of live broadcast, and improve the efficiency of digital transformation.
[0005] The first aspect of the present application provides a real-time interaction method for digital person live broadcast, comprising:
[0006] In response to the user's interaction behavior, a first digital person audio responding to the interaction behavior is generated;
[0007] According to the first digital person audio and a preset digital person template, a first digital person video responding to the interaction behavior is generated;
[0008] An insertion frame is selected in a key frame of digital person live broadcast content, and the first digital person audio and the first digital person video are inserted at the position of the insertion frame; the key frame is a video frame corresponding to a sentence boundary point in the digital person live broadcast content.
[0009] Further, the method described above further comprises:
[0010] The position of the end-of-sentence punctuation in the digital person live broadcast content is identified, and the position of the end-of-sentence punctuation is determined as the sentence boundary point; the end-of-sentence punctuation includes at least one of period, exclamation mark and line feed symbol.
[0011] Further, in the method described above, the digital person live broadcast content includes a second digital person video; the generation step of the digital person template comprises:
[0012] identifying a lip region of the digital human in the second digital human video;
[0013] masking the lip region to obtain the digital human template.
[0014] Further, in the case that part of the lip region is occluded by an occlusion object in the method described above, the masking the lip region comprises:
[0015] masking the lip region using a first mask and masking the occlusion object using a second mask;
[0016] erasing part of the first mask that coincides with the second mask.
[0017] Further, in the method described above, the generating a first digital human video responding to the interactive behavior according to the first digital human audio and a preset digital human template comprises:
[0018] driving the lip region of the digital human using the first digital human audio to generate a lip region video;
[0019] generating a target video segment according to the digital human template, replacing a first target video segment in the target video segment using the lip region video, and replacing a second target video segment in the target video segment using the second digital human video;
[0020] The first target video segment comprises video frames in the target video segment in which the lip region is not occluded or not completely occluded, and the second target video segment comprises video frames in the target video segment in which the lip region is completely occluded.
[0021] Further, in the method described above, the generating a target video segment according to the digital human template comprises:
[0022] cutting a video segment of a set time length from the digital human template with a target frame as a starting frame; the set time length is half of the total time length of the first digital human video; the position of the target frame in the digital human template is the same as the position of the insertion frame in the second digital human video;
[0023] composing the target video segment according to the order of playing the video frames in the video segment first and then playing them in reverse.
[0024] Further, in the method described above, the digital human live streaming content comprises a second digital human audio and a second digital human video; and the inserting the first digital human audio and the first digital human video at the position of the insertion frame comprises:
[0025] inserting the first digital human audio at a position corresponding to the insertion frame in the second digital human audio, and inserting the first digital human video at a position corresponding to the insertion frame in the second digital human video.
[0026] The second aspect of the present application provides a real-time interaction system for digital human live streaming, comprising:
[0027] a live streaming end and a video synthesis end, wherein the live streaming end and the video synthesis end are communicatively connected;
[0028] The live streaming end is configured to generate a first digital human audio in response to an interactive behavior of a user, and the video synthesis end is configured to generate a first digital human video in response to the interactive behavior according to the first digital human audio and a preset digital human template. The live streaming end is further configured to select an insertion frame in a key frame of digital human live streaming content, and insert the first digital human audio and the first digital human video at a position of the insertion frame. The key frame is a video frame corresponding to a sentence boundary point in the digital human live streaming content.
[0029] The third aspect of the present application provides an electronic device, comprising:
[0030] a processor; and
[0031] a memory having executable code stored thereon, wherein the executable code, when executed by the processor, causes the processor to perform the method as described above.
[0032] The fourth aspect of the present application provides a computer program product, comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the method as described above.
[0033] The technical solution provided by the present application can include the following beneficial results:
[0034] The technical solution of the present application generates a first digital human audio in response to an interactive behavior of a user, generates a first digital human video in response to the interactive behavior according to the first digital human audio and a preset digital human template, selects an insertion frame in a key frame of digital human live streaming content, and inserts the first digital human audio and the first digital human video at a position of the insertion frame. The key frame is a video frame corresponding to a sentence boundary point in the digital human live streaming content. In this way, the first digital human audio and the first digital human video in response to the interactive behavior of the user can be generated, and the first digital human audio and the first digital human video can be inserted into the digital human live streaming content to respond to the interactive content of the user in real time. Meanwhile, the first digital human audio and the first digital human video are inserted at the sentence boundary point of the digital human live streaming content, and will not interrupt the current speech tactfully, thereby improving the fluency of the live streaming and ensuring the live streaming effect.
[0035] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the application. BRIEF DESCRIPTION OF DRAWINGS
[0036] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which like reference characters refer to like parts throughout and in which:
[0037] Figure 1 is a flow diagram of a real-time interaction method of digital person live broadcast shown in embodiments of the present application;
[0038] Figure 2 is a flow diagram of a mask processing when a lip area is partially blocked shown in embodiments of the present application;
[0039] Figure 3 is a flow diagram of using a lip area video to replace a mask part in a first target video segment shown in embodiments of the present application;
[0040] Figure 4 is a structural diagram of a real-time interaction system of digital person live broadcast shown in embodiments of the present application;
[0041] Figure 5 is a structural diagram of an electronic device shown in embodiments of the present application. DETAILED DESCRIPTION
[0042] Embodiments of the present application will be described in more detail by referring to the drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present application will be more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0043] The terminology used in the present application is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in the present application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0044] It should be understood that although the terms "first", "second", "third", etc. can be used herein to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information without departing from the scope of the present application. Therefore, the features defined as "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.
[0045] With the rapid development of digital human technology, its penetration range in the live streaming field is expanding. From the product guide of e-commerce platforms, the new product launch of brand parties, to the course promotion of educational institutions, and the introduction of scenic spots in the tourism industry, digital humans have gradually become "new members" in the live streaming lineup in more and more live streaming scenarios, with characteristics such as no need to rest, uniform image, and customizability.
[0046] However, the current digital human live streaming has obvious short boards in interactivity. Most of the content presented by digital human live streaming is still based on pre-recorded materials, which are often played in a fixed script cycle. When the audience sends real-time comments, asks specific questions, or leaves interactive messages during the live streaming process, the digital human is often difficult to respond in a timely and accurate manner. For example, during the process of e-commerce digital human live streaming, when the audience asks about the specific size of a product, the digital human on the screen may still introduce the general information of the product according to the preset process, and has no response to these personalized interactive content.
[0047] This lack of interactivity greatly weakens the appeal of live streaming, resulting in a significant discount in the overall effectiveness of live streaming and making it difficult to achieve the expected communication or conversion goals.
[0048] Moreover, digital human live streaming is a typical application scenario of digital economy industry, and its core value lies in reconstructing consumer experience and supply chain efficiency through real-time interaction. However, the current digital human live streaming is characterized by delayed response, which leads to broken interaction, and in essence, it is a manifestation that the "user experience" and "process intelligentization" links of industrial digitization have not been fully connected, which weakens the effectiveness of digital transformation.
[0049] To solve the above problems, the present application provides a real-time interaction method, system, electronic device and program product for digital human live streaming, which can respond in a timely manner when facing real-time interactive information of users, ensure the overall effectiveness of live streaming, and at the same time, connect the "user experience" and "process intelligentization" links of industrial digitization, and improve the effectiveness of digital transformation.
[0050] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0051] The embodiment of the present application provides a real-time interaction method for digital person live broadcast, which can be executed by a real-time interaction system for digital person live broadcast. The system can be any device with data and instruction processing functions, for example, various types of user terminals such as notebook computers, tablet computers, desktop computers and mobile devices, or a combination of any two or more of the electronic devices, or a server. Referring to Figure 1 The method comprises the following steps.
[0052] S101, in response to the interaction behavior of the user, generating first digital person audio for responding to the interaction behavior.
[0053] The user refers to an audience of the digital person live broadcast, and the interaction behavior of the user refers to various types of participation behaviors initiated by the user in the live broadcast process and related to the live broadcast content, including entering a live broadcast room, commenting, liking, adding to a shopping cart and the like, which are not limited in the embodiment.
[0054] In the embodiment of the present application, the pre-recorded material is played in a loop during the live broadcast, and after detecting the interaction behavior of the user, first digital person audio for responding to the interaction behavior is generated in response to the interaction behavior.
[0055] In some embodiments, the text content for responding to various types of user behaviors can be configured in advance. For example, if the interaction content is that the user enters the live broadcast room, the text content for responding can be "Welcome to the live broadcast room"; if the interaction content is that the user likes, the text content for responding can be "Thank you for the little love of this friend", and the like, which are not limited in the embodiment. The TTS (Text-To-Speech) audio obtained by converting the text content into speech can also be used as the audio content, and a real person can also read the text content to record a real person's voice as the audio content, which are not limited in the embodiment. It should be noted that the timbre of the audio content needs to be consistent with the timbre of the digital person of the live broadcast.
[0056] The mapping relationship between the audio content and various types of interaction behaviors is established, so that the corresponding audio content can be selected as the response or feedback of the interaction behavior according to the mapping relationship after detecting the interaction behavior.
[0057] In the embodiment, after detecting the interaction behavior of the user, the audio matched with the interaction behavior can be selected from the audio content as the first digital person audio for responding to the interaction behavior according to the mapping relationship.
[0058] In some other embodiments, a large language model (LLM) can also be utilized to generate text content in real time to respond to the interactive behavior of the user. Specifically, after detecting the interactive behavior of the user, the interactive behavior can be input into the large language model in real time to obtain text content output by the large language model to respond to the interactive behavior. The text content generated based on the large language model is then converted into speech to obtain the first digital human audio.
[0059] It should be noted that, in order to make the text content output by the large language model to respond to the interactive behavior more accurate, a prompt word can be input into the large language model at the same time as the interactive behavior, which is used to prompt the large language model to output content meeting the requirements.
[0060] For example, the prompt word can be:
[0061] You are an intelligent interactive assistant in a live broadcast room, and you need to generate text content to respond to the interactive behavior of the user according to the interactive behavior of the user, with the following requirements:
[0062] Language style: colloquial, natural and friendly, in line with the tone of a real anchor;
[0063] Behavior matching: strictly distinguish different interactive types such as comments, entering the live broadcast room, likes, and adding to cart;
[0064] Personalization: replaceable variables are marked with {variable name}, such as {user nickname} and {product name}.
[0065] For example, the interactive behavior includes “entering the live broadcast room”, and the text content to respond to the interactive behavior includes “Welcome {user nickname} to enter the live broadcast room”; the interactive behavior includes “like”, and the text content to respond to the interactive behavior includes “Thank you {user nickname} for your little love”; the interactive behavior includes “adding to cart”, and the text content to respond to the interactive behavior includes “{user nickname} has excellent eyesight! The {product name} added to cart is very cost-effective”; the interactive behavior includes “commenting and asking questions”, and the text content to respond to the interactive behavior includes “{user nickname} asks a good question! I will answer it immediately”.
[0066] The large language model can adopt a mature large language model in the prior art, such as DeepSeek, and Doubao, which are not limited in the present embodiment.
[0067] In some embodiments, after detecting the interaction behavior of the user, it can be determined whether there is audio content matching the current acquired interaction behavior according to the mapping relationship between the audio content and various interaction behaviors pre-established in the above embodiments. If there is, the audio content matching the current acquired interaction behavior can be taken as the first digital human audio; if not, the current acquired interaction behavior can be input into the large language model, and the text content generated by the large language model can be converted into speech to obtain the first digital human audio. Such processing can not only shorten the generation time of the first digital human audio, but also improve the response effect of the interaction behavior.
[0068] In a specific embodiment, the real-time interaction system of the digital human live broadcast includes a live broadcast end, and pre-recorded materials can be played in a loop at the live broadcast end. The step of generating the first digital human audio responding to the interaction behavior of the user can also be completed at the live broadcast end, and the text content, audio content, and mapping relationship between the audio content and various interaction behaviors responding to various user behaviors pre-configured can be stored at the live broadcast end.
[0069] S102, generating the first digital human video responding to the interaction behavior according to the first digital human audio and the preset digital human template.
[0070] The digital human template refers to a template capable of generating a digital human video according to audio. The digital human template can be an image or a video containing at least a digital human face, and the digital human template can generate a video with a mouth shape corresponding to the audio under the driving of the first digital human audio.
[0071] In some embodiments, multi-view images of a digital human modeling object can be acquired, and the modeling object can be modeled using mature two-dimensional (2D) or three-dimensional (3D) modeling technology in the prior art to obtain a digital human template. Then, the lip region of the digital human template is driven by the first digital human audio to generate the first digital human video responding to the interaction behavior.
[0072] In a specific embodiment, the real-time interaction system of the digital human live broadcast further includes a video synthesis end, and the video synthesis end is in communication connection with the live broadcast end. The step of generating the first digital human video responding to the interaction behavior according to the first digital human audio and the preset digital human template can be completed at the video synthesis end, and the preset digital human template is also stored in the video synthesis end. It can be understood that the live broadcast end sends the first digital human audio and other related information to the video synthesis end, and the video synthesis end generates the first digital human video responding to the interaction behavior according to the first digital human audio and the preset digital human template.
[0073] S103. Select an insertion frame in a key frame of the digital human live streaming content, and insert the first digital human audio and the first digital human video at the position of the insertion frame.
[0074] Currently, digital human live streaming is mostly a looped playback of pre-recorded materials, which are the digital human live streaming content described above. It can be understood that the digital human live streaming content of the present embodiment refers to the pre-recorded and looped content in the process of digital human live streaming. The digital human live streaming content can be divided into second digital human audio corresponding to audio content and second digital human video corresponding to video content.
[0075] The key frame described above refers to a video frame corresponding to a sentence boundary point in the digital human live streaming content. The sentence boundary point refers to a boundary point between one sentence and another in the second digital human audio of the digital human live streaming content, which is a conversion node between adjacent semantic units.
[0076] In some embodiments, a sentence boundary point recognition model can be trained to recognize the sentence boundary points in the second digital human audio, and then determine the key frames according to the sentence boundary points.
[0077] Specifically, a large amount of audio can be obtained as training samples, and the timestamps of the sentence boundary points in the audio can be marked as training labels. In the training process of the sentence boundary point recognition model, the training samples are input into the sentence boundary point recognition model to obtain the prediction results output by the sentence boundary point recognition model. By comparing the prediction results output by the sentence boundary point recognition model and the training labels, the loss value of the sentence boundary point recognition model is determined, and the parameters of the sentence boundary point recognition model are adjusted to reduce the loss of the sentence boundary point recognition model. Then, the above training process is repeated until the parameters of the sentence boundary point recognition model meet the training requirements. The sentence boundary point recognition model can use any model as a basic model, which is not limited in the present embodiment.
[0078] The second digital human audio is input into the trained sentence boundary point recognition model described above to obtain the timestamps output by the sentence boundary point recognition model as the sentence boundary points of the second digital human audio. In the digital human live streaming content, the video frame with the same timestamp as the timestamp of the sentence boundary point is determined as the key frame of the digital human live streaming content.
[0079] Further, the insertion frame is selected from the key frames of the digital human live content. In an embodiment of the present application, the insertion frame is selected from the key frames of the digital human live content according to the timestamp of the video frame being played by the digital human live content when the interactive behavior of the user is detected. Specifically, the timestamp of the video frame being played by the digital human live content when the interactive behavior of the user is detected can be taken as the reference time, and then the first key frame after a specific time length from the reference time can be selected as the insertion frame.
[0080] The specific time length described above is the time length for waiting for the generation of the first digital human audio and the first digital human video, and the specific value of the specific time length can be determined according to the time consumption for generating the first digital human audio and the first digital human video, for example, 7S, 3S, etc., which is not limited in the present embodiment.
[0081] For example, if the total time length of the digital human live content is 2.5 hours in total, and the timestamp of the video frame being played by the digital human live content is 01:29:00.000, then 01:29:00.000 can be taken as the reference time, and the specific time length is 500ms, and the first key frame after 01:29:00.500 in the digital human live content can be selected as the insertion frame.
[0082] The first digital human audio and the first digital human video are inserted at the position of the insertion frame.
[0083] In order to ensure audio-video synchronization, the position of the key frame can be encoded, and only the key frame with complete encoding information can be inserted with the first digital human audio and the first digital human video, so as to achieve the purpose of ensuring audio-video synchronization, and to ensure that the first digital human audio and the first digital human video will not be inserted into the wrong position. If the insertion is not performed on the key frame, the incomplete encoding information will cause the failure of frame decoding on both the streaming pushing side and the streaming pulling side, resulting in the failure of insertion of the first digital human audio and the first digital human video.
[0084] In a specific embodiment, the step of selecting the insertion frame from the key frames of the digital human live content and inserting the first digital human audio and the first digital human video at the position of the insertion frame can be completed at the live streaming side. The video synthesis side returns the synthesized first digital human video to the live streaming side, and in some embodiments, the video synthesis side returns the synthesized first digital human video to the live streaming side through a long connection ACK (Acknowledgment) confirmation. The live streaming side selects the insertion frame from the key frames of the digital human live content and inserts the first digital human audio and the first digital human video at the position of the insertion frame.
[0085] In the above embodiments, the first digital human audio responding to the interactive behavior is generated, the first digital human video responding to the interactive behavior is generated according to the first digital human audio and the preset digital human template, the insertion frame is selected in the key frame of the digital human live content, and the first digital human audio and the first digital human video are inserted at the position of the insertion frame, wherein the key frame is a video frame corresponding to a sentence boundary point in the digital human live content. In this way, the first digital human audio and the first digital human video responding to the interactive behavior of the user can be generated, and the first digital human audio and the first digital human video are inserted into the digital human live content to respond to the interactive content of the user in real time. Meanwhile, the first digital human audio and the first digital human video are inserted at the sentence boundary point of the digital human live content, so that the current speech is not interrupted harshly, the smoothness of the live broadcast is improved, and the live broadcast effect is ensured.
[0086] Further, by enhancing the real-time and personalization of live interaction, the user engagement and conversion rate can be improved, the commercial value of live e-commerce, virtual services and other digital consumption formats can be deepened, and the digital transformation of the consumption end can be accelerated. At the same time, the integration and application of digital human technology and traditional industries such as finance, education and tourism are promoted, more real-time interactive service modes are generated, and the industry boundaries and service scenarios of the digital economy are expanded.
[0087] As an optional implementation, the method of the above embodiments can further include the following steps:
[0088] The position of the end-of-sentence punctuation in the digital human live content is identified, and the position of the end-of-sentence punctuation is determined as the sentence boundary point. The end-of-sentence punctuation includes at least one of a period, an exclamation point and a line break symbol.
[0089] The end-of-sentence punctuation indicates the completion of a sentence or a semantic. In the embodiments of the present application, the position of the end-of-sentence punctuation is taken as the sentence boundary point. The end-of-sentence punctuation includes at least one of a period, an exclamation point and a line break symbol.
[0090] In some embodiments, the second digital human audio in the digital human live content is subjected to automatic speech recognition (ASR) processing, the punctuation symbols in the second digital human audio are extracted, and the timestamp corresponding to the end-of-sentence punctuation is taken as the sentence boundary point.
[0091] Then, the key frame can be selected in the digital human live content at the timestamp of the sentence boundary point and subjected to key frame coding to realize visual positioning of the controllable insertion point, and the related information of the key frame is stored in the digital human live content.
[0092] In the above embodiment, the sentence boundary point in the digital human live broadcast content can be quickly located by identifying the end-of-sentence punctuation, so as to select the insertion frame in the key frame corresponding to the sentence boundary point, avoid the abrupt and unnatural feeling caused by interrupting the live broadcast speech, and improve the fluency of the live broadcast.
[0093] In addition to generating the digital human template by the 2D or 3D modeling technology in the prior art as described in the above embodiment, as an optional implementation, the present embodiment provides another generation step of the digital human template, specifically comprising:
[0094] Identifying the lip region of the digital human in the second digital human video; performing mask processing on the lip region to obtain the digital human template.
[0095] The lip region of the digital human in the second digital human video can be identified. Specifically, a face recognition algorithm can be called to identify the face region of the digital human in the second digital human video, and the face region of the digital human is cut out and saved, and the lower half of the face region of the digital human is taken as the lip region of the digital human.
[0096] The lip region is subjected to mask processing to obtain the digital human template. For example, the 68-point face landmark detection algorithm (landmark68) can be used to perform mask processing on the lip region to obtain the digital human template.
[0097] In this way, the digital human template can be obtained only by performing mask processing on the lip region of the digital human in the second digital human video, without modeling, and the processing speed is fast.
[0098] It should be noted that the lip region of the digital human can be occluded, for example, the lip region of the digital human is occluded by a microphone, a hand or other objects.
[0099] Specifically, if the pixel points with a depth less than the depth of the lip region can be extracted based on the depth information in the second digital human video, and it is found that the pixel points completely cover the lip region of the digital human, it means that the lip region of the digital human is completely occluded; if the pixel points with a depth less than the depth of the lip region can be extracted based on the depth information in the second digital human video, and it is found that the pixel points cover part of the lip region of the digital human, it means that the lip region of the digital human is partially occluded; if the pixel points with a depth less than the depth of the lip region cannot be extracted based on the depth information in the second digital human video, and it is found that the pixel points have no intersection with the lip region of the digital human, or the pixel points with a depth less than the depth of the lip region cannot be extracted, it means that the lip region of the digital human is not occluded at all.
[0100] If there is a part of the video frames in the second digital human video, in which the lip area of the digital human is completely unobstructed, the lip area can be normally masked for the part of the video frames;
[0101] If there is a part of the video frames in the second digital human video, in which the lip area of the digital human is completely obstructed, the part of the video frames does not need to be masked.
[0102] If there is a part of the video frames in the second digital human video, in which the lip area of the digital human is partially obstructed, the lip area can be masked for the part of the video frames by the following steps:
[0103] The lip area is masked using the first mask, and the obstruction is masked using the second mask; the part of the first mask that coincides with the second mask is erased.
[0104] Specifically, the lip area can be first masked using the first mask, and then the obstruction can be masked using the second mask. Since the lip area of the digital human is partially obstructed, the first mask and the second mask must have a coinciding part, and the coinciding part can be erased from the first mask. In this way, only the unobstructed area of the lip area of the digital human can be masked, and the obstruction is not affected by the mask.
[0105] As shown in the embodiment of Figure 2 , the lip area of the digital human anchor is partially obstructed by the microphone. The lip area can be first masked using the first mask to obtain a first mask image, the microphone can be masked using the second mask to obtain a second mask image, the part of the first mask that coincides with the second mask can be erased to obtain a third mask image, and the third mask image can be restored to the original image of the digital human anchor to obtain a digital human mask image. In this way, only the unobstructed area of the lip area of the digital human anchor can be masked.
[0106] In the above embodiment, the unobstructed area of the lip area of the digital human anchor can be masked only when the lip area of the digital human is partially obstructed by the obstruction, avoiding the influence of the obstruction on the synthesis of the digital human dynamic video, and avoiding the influence of the synthesis of the digital human dynamic video on the shape of the obstruction.
[0107] As an optional implementation, the steps of the above embodiment for generating the first digital human video of the response interaction behavior according to the first digital human audio and the preset digital human template can specifically include the following steps:
[0108] The lip region of the digital human is driven by the first digital human audio to generate a lip region video; a target video segment is generated according to a digital human template, a mask part in a first target video segment of the target video segment is replaced by the lip region video, and a second target video segment of the target video segment is replaced by a second digital human video; the first target video segment includes video frames in which the lip region is not occluded or completely occluded in the target video segment, and the second target video segment includes video frames in which the lip region is completely occluded in the target video segment.
[0109] In an embodiment of the present application, the lip region of the digital human is driven by the first digital human audio to generate a lip region video in which the lip action of the digital human is consistent with the first digital human audio. For example, feature points of the lip region of the digital human can be obtained, and audio features in the first digital human audio can be extracted. The audio features in the first digital human audio can use mel-frequency cepstral coefficients.
[0110] The feature points of the lip region of the digital human and the audio features are input into a video generation network, which can be a Wav2Lip or Wav2Lip-GAN network, without limitation in the present embodiment. The video generation network changes the shape of the static lip region based on the feature points of the lip region of the digital human and the audio features, so that the action of the lip region of the digital human is visually consistent with the first digital human audio, and then outputs a dynamic video of the lip region that is visually consistent with the first digital human audio, i.e., the lip region video in the above embodiment.
[0111] The target video segment is generated according to the digital human template. In some embodiments, a video frame segment with the same length as the first digital human audio can be cut from the digital human template as the target video segment, without limitation in the present embodiment.
[0112] If the target video segment includes the first target video segment, the mask part in the first target video segment of the target video segment is replaced by the lip region video. The first target video segment includes video frames in which the lip region is not occluded or completely occluded in the target video segment.
[0113] In some embodiments, the mask part of a B video frame in the first target video segment needs to be replaced by an A video frame in the lip region video. In the obtained C video frame, the action of the original mask part is consistent with the A video frame, and the action of the other part is consistent with the B video frame, as shown in Figure 3 .
[0114] It should be noted that when the mask part in the first target video segment is replaced by the lip region video, the replacement can be performed according to the position of the lip region key point. If the range of the lip region in the lip region video exceeds the range of the mask, the exceeding part is erased, as shown in Figure 3 .
[0115] Since the mask of the lip region in the embodiment is only the mask of the unoccluded region in the lip region, when the lip region video is used to replace the mask part in the first target video segment, the occlusion will not be replaced, thereby avoiding incomplete output of the occlusion and ensuring that the entire live picture is real and natural.
[0116] If the target video segment includes a second target video segment, the second target video segment of the target video segment is replaced using the second digital human video. The second target video segment includes a video frame in which the lip region of the digital human is completely occluded in the target video segment.
[0117] Since the lip region of the digital human in the second target video segment is completely occluded, the audience cannot see the action of the anchor's lip region, and therefore the original video, i.e., the second digital human video, can be used to replace the second target video segment to improve processing speed.
[0118] In some embodiments, when the lip region video is used to replace the mask part in the first target video segment, and the second digital human video is used to replace the second target video segment, the replacement needs to be performed according to the timestamps of the video frames.
[0119] Specifically, in the embodiment, the digital human template is obtained by performing mask processing on the second digital human video. The digital human template has the same length, frame rate, and resolution as the second digital human video, and each video frame in the digital human template retains the timestamp in the second digital human video. The frame rate and resolution of the generated first digital human video are completely consistent with those of the second digital human video.
[0120] The target video segment can regenerate the timestamps of the video frames in the same manner as the first digital human video. In this case, the video frames at the same position in the target video segment and the first digital human video have the same timestamp. If the mask part of the video frame with timestamp A in the first target video segment needs to be replaced, the video frame with timestamp A in the lip region video needs to be selected for replacement.
[0121] The target video segment can also retain the original timestamp in the digital human template, i.e., the target video segment retains the timestamp in the second digital human video. When the second digital human video is used to replace the second target video segment, if the video frame with original timestamp B in the second target video segment needs to be replaced, the video frame with timestamp B in the second digital human video needs to be selected for replacement.
[0122] In the above embodiments, in the case where part of the lip region is occluded by the occlusion, the occlusion will not be replaced, thereby avoiding incomplete output of the occlusion and ensuring that the entire live picture is real and natural. Meanwhile, in the case where the lip region is completely occluded, the original video, i.e., the second digital human video, is used to replace the second target video segment to improve processing speed.
[0123] In addition to the steps in the above embodiment, as an optional embodiment, the target video segment can also be generated from the digital human template by the following steps:
[0124] A video segment of a set time length is cut from the digital human template with the target frame as the starting frame; the set time length is half of the total time length of the first digital human video; the position of the target frame in the digital human template is the same as the position of the inserted frame in the second digital human video; and the target video segment is composed according to the order of playing the video frames in the video segment in the forward direction first and then in the reverse direction.
[0125] Specifically, in the embodiment of the present application, the video frame in the digital human template whose position is the same as the position of the inserted frame in the second digital human video is selected as the target frame, that is, the position of the target frame in the digital human template is the same as the position of the inserted frame in the second digital human video.
[0126] A video segment of a set time length is cut from the digital human template with the target frame as the starting frame, wherein the set time length is half of the total time length of the first digital human video, and in the case that the first digital human audio and the first digital human video have the same time length, the set time length is also half of the total time length of the first digital human audio.
[0127] Then the target video segment is composed according to the order of playing the video frames in the video segment in the forward direction first and then in the reverse direction. For example, the video segment of the set time length cut from the target frame includes: video frame 1, video frame 2, video frame 3, video frame 4 and video frame 5, and then the target video segment obtained according to the order of playing the video frames in the video segment in the forward direction first and then in the reverse direction is: video frame 1, video frame 2, video frame 3, video frame 4, video frame 5, video frame 5, video frame 4, video frame 3, video frame 2 and video frame 1.
[0128] It should be noted that in order to ensure that the time stamp does not become chaotic when the lip area video is used to replace the mask part in the first target video segment and the second digital human video is used to replace the second target video segment, the target video segment includes two groups of time stamps: one group of time stamps is the time stamp regenerated in the same way as the target video segment and the first digital human video, which is used when the lip area video replaces the mask part in the first target video segment; the other group of time stamps is the original time stamp of the target video segment in the digital human template, that is, the target video segment retains the time stamp in the second digital human video, which is used when the second digital human video replaces the second target video segment.
[0129] It also needs to be explained that since the target video segment is obtained in the order of first playing forward and then playing backward of the video frames in the video segment, the time stamp of the forward playing part can be kept as the original time stamp in the digital human template, and the part of the backward playing can be automatically continued in time stamp according to the end time of the forward playing part and the frame rate in the digital human template.
[0130] In the above embodiment, the target video segment is obtained in the order of first playing forward and then playing backward of the video frames in the video segment, the last frame of the target video segment is actually the target frame, is also the insertion frame, and belongs to the key frame. In this way, after the first digital human video of the response interaction behavior is played, the digital human is still at the current key frame, and another response video can be quickly inserted at the current position. Since the insertion of the other response video at the current position is also started from the target frame and ended at the target frame, the digital human action jump will not occur. Subsequent digital human live content playing can also be performed. Since the next frame of the digital human live content playing is the video frame adjacent to the insertion frame, the digital human action jump will not occur, and the authenticity and fluency of the live broadcast are improved.
[0131] In addition, it also needs to be explained that in some embodiments, the step of generating the first digital human video of the response interaction behavior according to the first digital human audio and the preset digital human template can be completed at the video synthesis end, and the preset digital human template is also stored in the video synthesis end. In order to ensure that the video synthesis end can generate an accurate target video segment, the live broadcast end needs to send the current live broadcast time stamp, the reserved key frame index and other information to the video synthesis end in addition to the first digital human audio. So that the video synthesis end can determine the insertion frame according to the current live broadcast time stamp, the reserved key frame index and other information, and then determine the target frame according to the position of the insertion frame.
[0132] In some embodiments, the live broadcast end sends the first digital human audio, the current live broadcast time stamp, the reserved key frame index and other information to the video synthesis end through a long connection.
[0133] As an optional implementation, the digital human live broadcast content of the above embodiment includes second digital human audio and second digital human video, and the step of inserting the first digital human audio and the first digital human video at the position of the insertion frame in the above embodiment specifically includes the following steps:
[0134] The first digital human audio is inserted at the position corresponding to the insertion frame in the second digital human audio, and the first digital human video is inserted at the position corresponding to the insertion frame in the second digital human video.
[0135] Specifically, in the embodiments of the present application, the digital person live content can be divided into second digital person audio corresponding to the audio content and second digital person video corresponding to the video content. When inserting the first digital person audio and the first digital person video, the first digital person audio can be inserted at the position corresponding to the frame in the second digital person audio, and the first digital person video can be inserted at the position corresponding to the frame in the second digital person video, so as to avoid the situation of audio and video being out of synchronization.
[0136] The first digital person audio and the first digital person video in the above embodiments are not cached or saved, and after synchronous playing is completed, they are pushed out for streaming, and subsequent digital person live content or other response videos and audios are continuously played.
[0137] Corresponding to the foregoing application function implementation method embodiments, the present application further provides a digital person live real-time interaction system, an electronic device, a computer readable storage medium, a computer program product and corresponding embodiments.
[0138] Figure 4 is a structural schematic diagram of the digital person live real-time interaction system shown in the embodiments of the present application.
[0139] Referring to Figure 4 The system of the above embodiments includes a live end 100 and a video synthesis end 110, and the live end 100 and the video synthesis end 110 are communicatively connected. It should be noted that the live end 100 and the video synthesis end 110 can be connected in a wired manner or in a wireless manner, and the present embodiments do not make any limitation.
[0140] The live end 100 is configured to generate first digital person audio responding to the interactive behavior of the user in response to the interactive behavior of the user; the video synthesis end 110 is configured to generate first digital person video responding to the interactive behavior according to the first digital person audio and a preset digital person template; and the live end 100 is further configured to select an insertion frame in a key frame of digital person live content, and insert the first digital person audio and the first digital person video at the position of the insertion frame; the key frame is a video frame corresponding to a sentence boundary point in the digital person live content.
[0141] Further, the live end 100 of the above embodiments is further configured to identify a position of a sentence ending punctuation in the digital person live content, and determine the position of the sentence ending punctuation as the sentence boundary point; the sentence ending punctuation includes at least one of a period, an exclamation mark and a line feed symbol.
[0142] Further, the digital person live content of the above embodiments includes second digital person video; and the video synthesis end 110 of the above embodiments is further configured to identify a lip region of the digital person in the second digital person video; and perform masking processing on the lip region to obtain the digital person template.
[0143] Further, in the case that part of the lip region is blocked by the occlusion object, the video synthesis end 110 of the above embodiment, when performing the masking processing on the lip region, is specifically configured to:
[0144] mask the lip region using the first mask and mask the occlusion object using the second mask; and erase the part of the first mask that coincides with the second mask.
[0145] Further, the video synthesis end 110 of the above embodiment, when generating the first digital human video of the response interactive behavior according to the first digital human audio and the preset digital human template, is specifically configured to:
[0146] drive the lip region of the digital human using the first digital human audio to generate a lip region video; generate a target video segment according to the digital human template, replace the mask part in a first target video segment of the target video segment using the lip region video, and replace a second target video segment of the target video segment using the second digital human video; the first target video segment includes video frames in which the lip region is not blocked and not completely blocked in the target video segment, and the second target video segment includes video frames in which the lip region is completely blocked in the target video segment.
[0147] Further, the video synthesis end 110 of the above embodiment, when generating the target video segment according to the digital human template, is specifically configured to:
[0148] cut a video segment of a set time length from the digital human template with the target frame as a starting frame; the set time length is half of the total time length of the first digital human video; the position of the target frame in the digital human template is the same as the position of the insertion frame in the second digital human video; and the target video segment is composed in the order of playing the video frames in the video segment first and then playing them in reverse.
[0149] Further, the digital human live streaming content of the above embodiment includes the second digital human audio and the second digital human video; and the live streaming end 100 of the above embodiment, when inserting the first digital human audio and the first digital human video at the position of the insertion frame, is specifically configured to:
[0150] insert the first digital human audio at the position corresponding to the insertion frame in the second digital human audio, and insert the first digital human video at the position corresponding to the insertion frame in the second digital human video.
[0151] In the above embodiment, the live streaming and the interactive content are decoupled, only the short-time interactive content is synthesized in real time, and the dependence on computing resources is significantly reduced. Moreover, the video synthesis end 110 as an independent module can serve multiple live streams and support resource sharing.
[0152] In a specific embodiment, in the case of using an RTX4090 graphics card, the video synthesis end 110 can support real-time responses of 30 live ends 210 digital person live broadcast, and the peak support can reach 50 live ends, which is much better than the cost-benefit ratio of the traditional full-link real-time scheme.
[0153] As to the system in the above embodiments, the specific manner in which the live end 100 and the video synthesis end 110 perform operations has been described in detail in the embodiments related to the method, and will not be described in detail here.
[0154] Figure 5 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present application.
[0155] Referring to FIG. 1, Figure 5 The electronic device includes a memory 200 and a processor 210.
[0156] The electronic device includes a memory 200 and a processor 210.
[0157] The memory 200 is connected with the processor 210, and is configured to store programs.
[0158] The processor 210 is configured to realize part or all of the method described above by running the programs stored in the memory 200.
[0159] Specifically, the electronic device described above can further include a bus, a communication interface 220, an input device 230 and an output device 240.
[0160] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected with each other through the bus.
[0161] The bus can include a path for transmitting information between various components of a computer system.
[0162] The processor 210 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0163] The processor 210 can include a main processor and can further include a baseband chip, a modem, etc.
[0164] The storage 200 can include various types of storage units such as a system memory, a read-only memory (ROM), and a permanent storage device. Among them, the ROM can store static data or instructions required by the processor 210 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose stored instructions and data even after the computer is powered off. In some embodiments, the permanent storage device employs a mass storage device (e.g., a magnetic or optical disk, a flash memory) as a permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (e.g., a floppy disk, an optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device such as a dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor at runtime. In addition, the storage 200 can include a combination of any computer readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), magnetic disks and / or optical disks. In some embodiments, the storage 200 can include a readable and / or writable removable storage device such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., an SD card, a min SD card, a Micro-SD card, etc.), a magnetic floppy disk, etc. The computer readable storage media does not include a carrier wave and an instantaneous electronic signal transmitted through wireless or wired transmission.
[0165] The input device 230 can include a device that receives data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.
[0166] The output device 240 can include a device that allows information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0167] The communication interface 220 can include a device using any transceiver to communicate with other devices or communication networks such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0168] The processor 210 executes programs stored in the storage 200 and calls other devices, which can be used to implement part or all of the above-described methods.
[0169] Further, the method according to the present application can also be implemented as a computer program product, which includes computer program code instructions for executing part or all of the steps in the above method of the present application. Alternatively, the computer program can be stored in a readable storage medium of a computer device or in the cloud; the processor of the computer device reads the computer program from the readable storage medium or the cloud.
[0170] The program code of the program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language, or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.
[0171] The above computer program product can be specifically implemented by means of hardware, software or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium, and in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.
[0172] Alternatively, the present application can also be implemented as a computer readable storage medium (or a non-transitory machine readable storage medium or a machine readable storage medium) having stored thereon executable codes (or computer programs or computer instruction codes) which, when executed by a processor of an electronic device (or a server, etc.), cause the processor to perform part or all of the steps of the above method according to the present application.
[0173] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0174] Having described various embodiments of the application, it is to be understood that the above description is meant not to limit and not to encompass all of the possible embodiments. Many modifications and variations of this application can be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. It is intended that the scope of the application be defined by the scope of the patent and equivalents thereof.
Claims
1. A real-time interactive method for live streaming of digital humans, characterized in that, include: In response to a user's interactive behavior, generate a first digital human audio response to the interactive behavior; Based on the first digital human audio and a preset digital human template, a first digital human video responding to the interactive behavior is generated; Select an insertion frame from the keyframes of the digital human live stream content, and insert the first digital human audio and the first digital human video at the position of the insertion frame; The keyframe is the video frame corresponding to the sentence demarcation point in the live content of the digital human; The digital human live stream content includes a second digital human video; the digital human template generation step includes: identifying the lip area of the digital human in the second digital human video, masking the lip area, and obtaining the digital human template; The step of generating a first digital human video responding to the interactive behavior based on the first digital human audio and a preset digital human template includes: using the first digital human audio to drive the lip region of the digital human to generate a lip region video; generating a target video segment based on the digital human template; if the target video segment includes a first target video segment, then replacing the masked portion in the first target video segment with the lip region video; if the target video segment includes a second target video segment, then replacing the second target video segment with the second digital human video; the first target video segment includes video frames in the target video segment where the lip region is not masked and is not completely masked, and the second target video segment includes video frames in the target video segment where the lip region is completely masked.
2. The method according to claim 1, characterized in that, Also includes: Identify the location of the sentence-ending punctuation mark in the digital human live broadcast content, and determine the location of the sentence-ending punctuation mark as the sentence-breaking point; the sentence-ending punctuation mark includes at least one of a period, an exclamation mark, and a line break.
3. The method according to claim 1, characterized in that, When a portion of the lip area is obscured by an obstruction, the masking process for the lip area includes: The lip area is masked using a first mask, and the obstruction is masked using a second mask. Erase the portion of the first mask that overlaps with the second mask.
4. The method according to claim 1, characterized in that, The step of generating the target video segment based on the digital human template includes: From the digital human template, a video segment of a set duration is extracted with the target frame as the starting frame; the set duration is half of the total duration of the first digital human video; the position of the target frame in the digital human template is the same as the position of the inserted frame in the second digital human video; The target video segment is formed by first playing the video frames forward and then backward in the order they appear in the video segment.
5. The method according to any one of claims 1-4, characterized in that, The live stream content of the digital human includes second digital human audio and second digital human video; inserting the first digital human audio and first digital human video at the position of the inserted frame includes: The first digital human audio is inserted at the position corresponding to the insertion frame in the second digital human audio, and the first digital human video is inserted at the position corresponding to the insertion frame in the second digital human video.
6. A real-time interactive system for live streaming of digital humans, characterized in that, include: The live streaming terminal and the video compositing terminal are communicatively connected. The live streaming terminal is used to respond to the user's interactive behavior and generate a first digital human audio that responds to the interactive behavior; the video synthesis terminal is used to generate a first digital human video that responds to the interactive behavior based on the first digital human audio and a preset digital human template. The live streaming terminal is also used to select an insertion frame from the keyframes of the digital human live streaming content, and insert the first digital human audio and the first digital human video at the position of the insertion frame; the keyframe is the video frame corresponding to the sentence demarcation point in the digital human live streaming content. The live stream content of the digital human includes a second digital human video; the video synthesis terminal is also used to identify the lip area of the digital human in the second digital human video, and to mask the lip area to obtain the digital human template. The step of generating a first digital human video responding to the interactive behavior based on the first digital human audio and a preset digital human template includes: using the first digital human audio to drive the lip area of the digital human to generate a lip area video. A target video segment is generated based on the digital human template. If the target video segment includes a first target video segment, the masked portion in the first target video segment is replaced with the video of the lip region. If the target video segment includes a second target video segment, the second target video segment is replaced with the video of the second digital human. The first target video segment includes video frames in the target video segment where the lip region is not masked and is not completely masked. The second target video segment includes video frames in the target video segment where the lip region is completely masked.
7. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-5.
8. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the method described in any one of claims 1-5.
Citation Information
Patent Citations
Character-driven lip sound synchronous digital human generation method and device, equipment and medium
CN119274534A
Avatar control method, apparatus, and related device
WO2025025564A1