Data transmission processing method and system for digital human-generated video

Through collaborative processing on the server and client, the necessary feature data is transmitted according to the audio and video clip identification sequence, the problem of high hardware resources and transmission bandwidth requirements in digital human video generation is solved, and more efficient video synthesis and playback is achieved.

CN120050452BActive Publication Date: 2025-08-15国家超级计算天津中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510509134.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-15
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing digital video generation methods have high demands in hardware resources and transmission bandwidth, resulting in increased equipment costs and poor video display fluency.

Method used

By determining the video clip identification sequence based on the audio duration and video clips on the server and sending it to the client, combining the client's facial features and audio clips, only the necessary feature data is transmitted to synthesize videos, reducing the data transmission amount and the server processing amount.

Benefits of technology

It reduces the amount of data transmission for digital human video synthesis and playback, improves fluency and hardware resource utilization efficiency, and reduces network transmission costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050452B_ABST
    Figure CN120050452B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing and discloses a data transmission processing method and system for a digital human-generated video. The method is applied to a server and comprises: determining a video segment identification sequence according to the audio duration of a target audio and each action video segment, and sending the video segment identification sequence to a client; determining an audio segment corresponding to each selected segment according to the target audio and the duration of each segment identification in the video segment identification sequence; obtaining a facial feature existence result of the client and determining a feature to be transmitted corresponding to each selected segment according to the existence result, each selected segment and each audio segment; and sending each feature to be transmitted and the audio segment in sequence to the client in the order of the video segment identification sequence, so that the client synthesizes and plays the target video according to the received data and pre-stored data, thereby reducing the amount of data transmission required for synthesizing and playing the digital human video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data transmission and processing method and system for digital human-generated videos. Background Art

[0002] Generating digital human videos is crucial for achieving a vivid display. Currently, the most common method for generating digital human videos is for the server to first synthesize lip movements based on speech, then precisely apply those movements to the video character to create a complete video, which is then sent to the client for display.

[0003] However, the entire process, from lip syncing to the final video, requires significant hardware resources, increasing equipment costs and placing extremely high demands on hardware performance. Furthermore, the large amount of video data synthesized on the server and then transmitted to the client consumes significant bandwidth, affecting the smoothness of the video display. This can lead to frequent freezes, especially in poor network conditions.

[0004] In view of this, the present invention is proposed. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a data transmission processing method and system for digital human generated videos, which reduces the amount of data transmission required for digital human video synthesis and playback, and improves the smoothness of digital human video synthesis and playback.

[0006] An embodiment of the present invention provides a data transmission and processing method for a video generated by a digital human, which is applied to a server. The method includes:

[0007] Determine a video segment identification sequence based on the audio duration of the target audio and each action video segment, and send the video segment identification sequence to the client;

[0008] Determining the audio segment corresponding to each selected segment according to the target audio and the duration of the selected segment corresponding to each segment identifier in the video segment identifier sequence;

[0009] Obtaining a facial feature existence result from the client, and determining a feature to be transmitted corresponding to each selected segment based on the existence result, each selected segment, and an audio segment corresponding to each selected segment;

[0010] Sending the features to be transmitted and the audio segments corresponding to the selected segments to the client in sequence according to the order of the video segment identification sequence, so that the client synthesizes and plays the target video based on the received features to be transmitted, the audio segments, the pre-stored decoder parameters of the lip-sync decoder, the action video segments, the lip-sync position timing sequence corresponding to the action video segments, the silent video segments, and the video segment identification sequence;

[0011] The action video segments and the silent video segments are determined based on the original video segmentation of the target digital human.

[0012] An embodiment of the present invention provides a data transmission and processing method for a video generated by a digital human, which is applied to a client. The method includes:

[0013] Receive a video segment identification sequence sent by a server, and send the existence result of the facial feature to the server, so that the server determines the to-be-transmitted feature corresponding to each selected segment based on the existence result, the selected segment corresponding to each segment identification in the video segment identification sequence, and the audio segment corresponding to each selected segment; wherein the video segment identification sequence is determined by the server based on the audio duration of the target audio and each action video segment, and the audio segment corresponding to each selected segment is determined by the server based on the target audio and the duration of the selected segment corresponding to each segment identification in the video segment identification sequence;

[0014] The system sequentially receives the features to be transmitted of the selected segments corresponding to each segment identifier in the video segment identifier sequence sent by the server, and synthesizes and plays the target video based on the received features to be transmitted, the audio segment, the decoder parameters of the pre-stored lip decoder, the action video segments, the lip position timing sequence corresponding to each action video segment, the silent video segments, and the video segment identifier sequence.

[0015] An embodiment of the present invention provides a data transmission and processing system for a video generated by a digital human, characterized in that it includes: a server and a client; the server executes the steps of the data transmission and processing method for a video generated by a digital human as described in any embodiment; the client executes the steps of the data transmission and processing method for a video generated by a digital human as described in any embodiment.

[0016] The embodiments of the present invention have the following technical effects:

[0017] The server determines a video segment identification sequence based on the audio duration of the target audio and each action video segment, and sends the video segment identification sequence to the client to indicate the playback order of the client. Furthermore, based on the target audio and the duration of the selected segments corresponding to each segment identifier in the video segment identification sequence, the audio segment corresponding to each selected segment is determined, so that the target audio can be split as required for subsequent block transmission. The presence result of the client's facial features is obtained, and based on the presence result, each selected segment, and the audio segment corresponding to each selected segment, the features to be transmitted corresponding to each selected segment are determined to extract the feature data to be transmitted. Finally, the features to be transmitted and the audio segment corresponding to each selected segment are sent to the client in the order of the video segment identification sequence. The client synthesizes and plays the target video based on the received features to be transmitted, the audio segment, the pre-stored decoder parameters of the lip position decoder, the action video segments, the lip position timing sequence corresponding to each action video segment, the silent video segments, and the video segment identification sequence. This reduces the task processing workload of the server and the data transmission volume between the server and the client, and improves the fluency of the subsequent digital human video synthesis and playback on the client. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is a flow chart of a data transmission processing method for a digital human-generated video provided by an embodiment of the present invention;

[0020] Figure 2 Schematic diagram of various feature points corresponding to the face position and the character's lip shape area provided by an embodiment of the present invention;

[0021] Figure 3 This is a flowchart of another data transmission processing method for a digital human-generated video provided by an embodiment of the present invention;

[0022] Figure 4 This is a structural diagram of a data transmission and processing system for a digital human-generated video provided by an embodiment of the present invention;

[0023] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0024] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0025] The data transmission processing method for digital human-generated videos provided in the embodiments of the present invention is primarily applicable to reducing the amount of data transmitted during the generation and playback of digital human videos and improving the smoothness of digital human videos. The data transmission processing for digital human-generated videos provided in the embodiments of the present invention can be performed by the server and / or the client.

[0026] Example 1

[0027] Figure 1 This is a flow chart of a data transmission and processing method for a digital human-generated video provided by an embodiment of the present invention. Figure 1 The data transmission and processing method of the digital human-generated video is applied to the server, specifically including:

[0028] S110 : Determine a video segment identification sequence according to the audio duration of the target audio and each action video segment, and send the video segment identification sequence to the client.

[0029] The target audio is the audio required for the final digital human video to be synthesized, that is, the audio required for the digital human to lip-sync. The audio duration is the duration required for normal playback of the target audio. Each action video segment and each silent video segment is determined based on the original video segmentation of the target digital human. The target digital human is the digital human used in the final synthesized video, and the original video is the video obtained when the target digital human is recorded and filmed. The action video segment is a video segment obtained by splitting the part of the original video where the target digital human speaks, and there is at least one action video segment. The silent video segment can be a video segment obtained by splitting the part of the original video where the target digital human does not speak, or it can be a video segment processed from a certain action video segment into a silent video segment, and there is at least one silent video segment. The video segment identifier sequence is a sequence of segment identifiers of the action video segments required for the subsequent synthesis of the digital human video, wherein the segment identifiers can appear repeatedly.

[0030] Specifically, one action video clip is randomly selected from each action video clip, and its corresponding segment identifier is added to the video segment identifier sequence. The total duration of the action video clips corresponding to each segment identifier in the video segment identifier sequence is calculated to determine whether it is greater than or equal to the audio duration of the target audio. If so, the video segment identifier sequence is determined to be obtained. If not, one action video clip is randomly selected from each action video clip, and its corresponding segment identifier is added to the video segment identifier sequence until the total duration of the action video clips corresponding to each segment identifier in the video segment identifier sequence is greater than or equal to the audio duration of the target audio, thereby obtaining the video segment identifier sequence. The video segment identifier sequence is then sent to the client so that the corresponding action video clips can be retrieved according to the video segment identifier sequence in the client to synthesize and play the target video.

[0031] Based on the above example, the video segment identification sequence can be determined according to the audio duration of the target audio and each action video segment in the following manner:

[0032] Initialize the video segment identification sequence, and use the sum of the segment durations of the selected segments corresponding to the segment identifications in the video segment identification sequence as the existing duration;

[0033] If the existing duration is less than the audio duration of the target audio, randomly select one from the segment identifiers corresponding to the action video segments and add it to the video segment identifier sequence, and return to the step of taking the sum of the segment durations of the selected segments corresponding to the segment identifiers in the video segment identifier sequence as the existing duration, until the existing duration is greater than or equal to the audio duration of the target audio;

[0034] When the existing duration is greater than or equal to the audio duration of the target audio, a video segment identification sequence is obtained.

[0035] The segment identifier is the identifier of each action video segment and is used to distinguish different action video segments. Existing segments are action video segments determined according to the segment identifiers in the video segment identifier sequence. They can appear repeatedly to reflect the randomness and diversity of the digital human's movements. Existing duration is the sum of the durations of the selected segments corresponding to each segment identifier in the video segment identifier sequence. Selected segments are the action video segments corresponding to each segment identifier in the video segment identifier sequence.

[0036] Specifically, initialize the video segment identification sequence to set the video segment identification sequence to be empty. Then, take the sum of the segment durations of the selected segments corresponding to each segment identification in the video segment identification sequence as the existing duration. Determine the size relationship between the existing duration and the audio duration of the target audio. If it is less than, randomly select one from the segment identifications corresponding to each action video segment, add it to the video segment identification sequence, and return to execute the step of taking the sum of the segment durations of the selected segments corresponding to each segment identification in the video segment identification sequence as the existing duration, until the existing duration is greater than or equal to the audio duration of the target audio. It can be understood that, during the selection process, there is no restriction on selecting repeated or non-repeated segment identifications. If it is greater than or equal to, the current video segment identification sequence is taken as the finalized video segment identification sequence.

[0037] Based on the above example, before determining the video segment identification sequence based on the audio duration of the target audio and each action video segment, it is also necessary to determine each action video segment and silent video segment based on the original video of the target digital human, determine the lip position timing sequence and segment identification corresponding to the action video segment, train the lip generation model, and send the required information to the client in advance for storage so that it can be directly called in subsequent use to avoid large amounts of data transmission during the synthesis process. Specifically, it can be:

[0038] Determining at least one action video segment and at least one silent video segment based on the original video of the target digital human and the shortest duration of the video segment;

[0039] For each action video clip, determine the lip position timing sequence and clip identifier corresponding to the action video clip;

[0040] The decoder parameters of the pre-trained lip-sync decoder, each silent video segment, each action video segment, and the lip position timing sequence and segment identifier corresponding to each action video segment are sent to the client.

[0041] The identity encoder, speech encoder, and lip-sync decoder comprise the lip-sync generation model (Wav2Lip). This lip-sync generation model is trained based on the temporal sequence of lip positions corresponding to each action video clip and its corresponding action video clip. The identity encoder extracts facial features from the video, specifically facial features (e.g., facial structure and skin texture). The speech encoder processes the audio signal and extracts speech features. The lip-sync decoder fuses facial and speech features and generates lip-sync animations for the lip-sync portion. The minimum video clip duration is a pre-set minimum duration for each action video clip and each silent video clip. The decoder parameters are the various parameter values required by the lip-sync decoder in the trained lip-sync generation model.

[0042] Specifically, the original video of the target digital human is split based on the minimum duration of the video segments, resulting in at least one action video segment and at least one silent video segment. For each action video segment, the facial position in each video frame is identified and the feature points are determined. The area enclosed by the feature points in the lower half of the facial position is considered the character's lip shape area. The temporal sequence of the feature points on and within this area is determined, and the lip shape position temporal sequence corresponding to the action video segment is constructed, and the corresponding segment identifier is determined. The lip shape generation model is trained based on each action video segment and the lip shape position temporal sequence corresponding to each action video segment. The decoder parameters of the pre-trained lip shape decoder, the lip shape position temporal sequence corresponding to each silent video segment, each action video segment, and the segment identifier corresponding to each action video segment are sent to the client, so that the client can subsequently retrieve the corresponding data from the storage space based on the received video segment identifier sequence and use the corresponding lip shape decoder.

[0043] For example, the schematic diagram of each feature point corresponding to the face position and the character's mouth area is as follows: Figure 2 As shown in the figure, there are 68 feature points corresponding to the face position.

[0044] Based on the above example, at least one action video segment and at least one silent video segment can be determined according to the original video of the target digital human and the shortest duration of the video segment in the following manner:

[0045] Taking the starting video frame in the original video of the target digital human as the first video frame, and determining the position of the human in the first video frame;

[0046] If the time interval between the first video frame and the last video frame in the original video is greater than or equal to the shortest duration of the video segment, the video frame in the original video whose time interval with the first video frame is equal to the shortest duration of the video segment is selected as the candidate video frame, and the position of the person in the candidate video frame is determined;

[0047] In response to a positional deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame being greater than or equal to a preset deviation, if the to-be-selected video frame is not the last video frame in the original video, using the next video frame of the to-be-selected video frame as the to-be-selected video frame, and returning to the step of determining the position of the person in the to-be-selected video frame until the positional deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame is less than the preset deviation;

[0048] In response to a positional deviation between the position of the person in the selected video frame and the position of the person in the first video frame being less than a preset deviation, the selected video frame is used as a second video frame corresponding to the first video frame, a video segment between the first video frame and the second video frame is used as a target video segment, and a speaking situation of the target digital person in the target video segment is determined;

[0049] Using the next video frame of the second video frame as a new first video frame, and returning to the step of determining the position of the person in the first video frame;

[0050] The target digital person's speaking situation is the target video clip of the target digital person speaking as each action video clip;

[0051] If there is at least one target digital person whose speaking condition is that the target digital person is not speaking, then the target video segments whose speaking condition is that the target digital person is not speaking are used as silent video segments;

[0052] If any target digital human speaks, then one is randomly selected from each action video clip as the video clip to be processed, the mouth of the target digital human in the video clip to be processed is processed as a closed state, and the processed video clip to be processed is used as a silent video clip.

[0053] Among them, the first video frame is the first video frame in the target video clip. The second video frame is the last video frame in the target video clip. The character position is the position of the target digital person in the video frame. The selected video frame is the candidate ending video frame corresponding to the first video frame, that is, the candidate second video frame corresponding to the first video frame. The position deviation is the absolute value of the distance between the character position of the first video frame and the character position of the selected video frame. If multiple position points describe the character position, it can be the sum of the absolute values of the distances of each position point. The target video clip is a clip split from the original video, in which the position deviation of the character position between the first video frame and the last video frame is less than the preset deviation. The preset deviation is used to determine whether the character positions basically overlap. The target digital person's speaking situation includes the target digital person speaking and the target digital person not speaking, which can be identified based on the audio corresponding to the target video clip. The video clip to be processed is an action video clip to be processed into a silent video clip.

[0054] Specifically, the starting video frame in the original video of the target digital human is used as the first video frame, and the position of the person in the first video frame is determined. First, it is determined whether the time interval between the first video frame and the last video frame in the original video is greater than or equal to the shortest duration of the video segment. If not, it means that the remaining part of the video cannot be split into the target video segment due to insufficient duration. If so, the video segment can be split, and the video frame in the original video whose time interval with the first video frame is equal to the shortest duration of the video segment is used as the selected video frame, and the position of the person in the selected video frame is identified and determined. Then, the position deviation between the position of the person in the selected video frame and the position of the person in the first video frame is calculated, which can be the sum of the absolute values of the distances between the corresponding position points corresponding to the person's position. It is determined whether the position deviation is greater than or equal to the preset deviation. If so, it is determined that the positions of the characters in the first video frame and the selected video frame do not substantially overlap, and the selected video frame cannot be used as the second video frame corresponding to the first video frame. Therefore, the next video frame after the selected video frame is used as the new selected video frame, and the process of determining the position of the characters in the selected video frame is repeated until the positional deviation between the positions of the characters in the selected video frame and the first video frame is less than a predetermined deviation. If not, it is determined that the positions of the characters in the first video frame and the selected video frame substantially overlap, and the selected video frame can be used as the second video frame corresponding to the first video frame. The video segment between the first and second video frames is used as a target video segment, and the target digital person's speech in the target video segment is determined based on the audio in the target video segment. Furthermore, the next video frame after the second video frame is used as the new first video frame, and the process is repeated, that is, the process of determining the position of the characters in the first video frame is repeated until the original video is split. It can be seen that the target digital person in the first and second video frames of each split target video segment substantially overlaps. After splitting the target video segments, the target video segments in which the target digital person speaks are used as the action video segments. A determination is made as to whether there is at least one target digital person in which the target digital person is not speaking. If so, it indicates that there is a silent video segment. The target video segments in which the target digital person is not speaking are directly used as the silent video segments. If not, it indicates that there is no silent video segment and a silent video segment needs to be produced. One of the action video segments is randomly selected as the video segment to be processed. This selection can be based on a random number. The mouth of the target digital person in the video segment to be processed is identified and the mouth animation is processed to a closed state to simulate the mouth animation of the target digital person not speaking. The processed video segment to be processed is used as the silent video segment.

[0055] Based on the above example, the facial features corresponding to each action video clip can also be sent to the client in advance to reduce the amount of data transmission when synthesizing the target video later. Specifically, it can be:

[0056] For each action video clip, the image frame sequence in the action video clip is input into a pre-trained identity encoder to obtain the facial features corresponding to the action video clip;

[0057] The facial features corresponding to each action video clip are sent to the client.

[0058] The image frame sequence is the remaining portion of the action video clip after removing the audio, and the facial features are the result of feature extraction of the facial region by the identity encoder based on the image frame sequence.

[0059] Specifically, for each action video clip, the audio and video are separated to obtain a sequence of image frames. This sequence of image frames is then fed into a pre-trained identity encoder, which outputs the facial features corresponding to the action video clip. Furthermore, the facial features corresponding to each action video clip are sent to the client, where they are pre-stored.

[0060] It is understandable that since the amount of facial feature data corresponding to each action video clip is not large, it can be transmitted to the client in advance for storage, or the facial features can be determined and transmitted when the features to be transmitted corresponding to the selected clip are subsequently determined and transmitted.

[0061] S120: Determine the audio segment corresponding to each selected segment according to the target audio and the duration of the selected segment corresponding to each segment identifier in the video segment identifier sequence.

[0062] The audio clip is a portion of the target audio that is captured according to the duration of the selected clip.

[0063] Specifically, the target audio is intercepted in sequence according to the duration of the selected segments corresponding to each segment identifier in the video segment identifier sequence, and the audio segments corresponding to each selected segment can be obtained. It can be understood that the audio duration of the last audio segment may be shorter than the duration of the corresponding selected segment.

[0064] Optionally, a more accurate splitting can be performed in combination with the audio information of the target audio, which may include: dividing the target audio into audio segments according to the duration of the selected segments corresponding to each segment identifier in the video segment identifier sequence, wherein the segmentation position of adjacent audio segments is located at the speaking pause, and the duration of the audio segment is equivalent to the duration of the corresponding selected segment (the duration of the audio segment is less than or equal to the duration of the corresponding selected segment). Therefore, a selected segment can correspond to a lip position timing sequence and an audio segment.

[0065] S130: Obtain the existence result of the facial features of the client, and determine the to-be-transmitted features corresponding to each selected segment based on the existence result, each selected segment, and the audio segment corresponding to each selected segment.

[0066] The facial feature presence result describes whether the client has pre-stored facial features for the target digital human's corresponding action video clip. This can be either true or false. The features to be transmitted are the features required for real-time transmission when synthesizing and playing the target digital human's target video according to the target audio.

[0067] Specifically, obtaining the presence result of the client's facial features may include issuing a corresponding facial feature presence recognition request to the client and determining the presence result of the client's facial features based on feedback from the client based on the facial feature presence recognition request. Based on the presence result, each selected segment and the audio segment corresponding to each selected segment are processed. When different presence results are determined, the features to be transmitted corresponding to each selected segment can be understood as indicating whether facial features need to be added to the features to be transmitted.

[0068] Based on the above example, the following method can be used to determine the to-be-transmitted features corresponding to each selected segment based on the existence result, each selected segment, and the audio segment corresponding to each selected segment:

[0069] In response to the existence result being that the image frame sequence in the selected segment is not present, for each selected segment, inputting the image frame sequence in the selected segment into a pre-trained identity encoder to obtain facial features corresponding to the selected segment, inputting the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain speech features corresponding to the selected segment, and using the facial features and the speech features as features to be transmitted corresponding to the selected segment;

[0070] In response to the existence result being existence, for each selected segment, the audio segment corresponding to the selected segment is input into a pre-trained speech encoder to obtain the speech features corresponding to the selected segment, and the speech features are used as the features to be transmitted corresponding to the selected segment.

[0071] The speech feature is the result of feature extraction performed by the speech encoder based on the audio segment corresponding to the selected segment.

[0072] Specifically, if the existence result is not present, the features to be transmitted include the facial features and voice features of each selected segment. For each selected segment, the selected segment is separated into audio and video to obtain a sequence of image frames, and the image frame sequence is input into a pre-trained identity encoder to obtain the facial features corresponding to the selected segment, and the audio segment corresponding to the selected segment is input into a pre-trained voice encoder to obtain the voice features corresponding to the selected segment. Then, the facial features and voice features can be used as the features to be transmitted corresponding to the selected segment. If the existence result is present, the features to be transmitted only include the voice features of each selected segment. The method of obtaining the voice features of each selected segment is the same as the above case and will not be repeated here. Then, the voice features are used as the features to be transmitted corresponding to the selected segment.

[0073] S140. The features to be transmitted and the audio segments corresponding to the selected segments are sent to the client in sequence according to the order of the video segment identification sequence, so that the client can synthesize and play the target video based on the received features to be transmitted, the audio segments, the decoder parameters of the pre-stored lip decoder, the action video segments, the lip position timing sequence corresponding to the action video segments, the silent video segments and the video segment identification sequence.

[0074] The target video is the final digital human video, the audio is the target audio, and the digital human's lip shape corresponds to the target audio.

[0075] Specifically, the features to be transmitted and the audio segments corresponding to each selected segment are sent to the client in groups according to the order of the video segment identification sequence. This allows the client to synthesize and play the video upon receiving a set of data, improving real-time performance. The client can synthesize and play the target video based on the received features to be transmitted and the audio segments, combined with the decoder parameters of the lip-sync decoder pre-stored in the client, the action video segments, the corresponding lip-sync position timing sequence of each action video segment, the silent video segments, and the video segment identification sequence.

[0076] This approach eliminates the need for complex video synthesis operations on the server side. Instead, it directly transmits the lip-feature data stream (the features to be transmitted) to the client, significantly reducing server-side processing tasks and lowering the computing power requirements of the server hardware. Furthermore, since lip-feature data is transmitted rather than the entire video, the data volume is significantly reduced. This lower data volume reduces network transmission costs and improves data transmission efficiency.

[0077] The present invention has the following technical effects: at the server side, a video segment identification sequence is determined according to the audio duration of the target audio and each action video segment, and the video segment identification sequence is sent to the client side to indicate the playback order of the client side, and then, according to the target audio and the duration of the selected segment corresponding to each segment identification in the video segment identification sequence, the audio segment corresponding to each selected segment is determined, so that the target audio can be split according to demand, which is convenient for subsequent block transmission, and the existence result of the client's facial features is obtained, and according to the existence result, each selected segment and the audio segment corresponding to each selected segment, the feature to be transmitted corresponding to each selected segment is determined. Features are extracted to extract the feature data to be transmitted. Finally, the features to be transmitted and the audio clips corresponding to each selected clip are sent to the client in sequence according to the order of the video clip identification sequence, so that the client can synthesize and play the target video according to the received features to be transmitted, audio clips, decoder parameters of the pre-stored lip decoder, each action video clip, the lip position timing sequence corresponding to each action video clip, each silent video clip and the video clip identification sequence, thereby reducing the task processing workload of the server and the data transmission volume between the server and the client, and improving the smoothness of the digital human video synthesis and playback on the subsequent client.

[0078] Example 2

[0079] Figure 3 This is a flow chart of another data transmission and processing method for a digital human-generated video provided by an embodiment of the present invention. Figure 3 The data transmission and processing method of the digital human-generated video is applied to the client, specifically including:

[0080] S210. Receive a video segment identification sequence sent by the server, and send the existence result of the facial feature to the server, so that the server determines the to-be-transmitted feature corresponding to each selected segment based on the existence result, the selected segment corresponding to each segment identification in the video segment identification sequence, and the audio segment corresponding to each selected segment.

[0081] Among them, the video segment identification sequence is determined by the server based on the audio duration of the target audio and each action video segment, and the audio segment corresponding to each selected segment is determined by the server based on the target audio and the duration of the selected segment corresponding to each segment identification in the video segment identification sequence.

[0082] Specifically, the client can receive a video segment identification sequence sent by the server, and can determine from the storage space each action video segment corresponding to the video segment identification sequence, determine whether corresponding facial features are stored for these determined action video segments, obtain a facial feature existence result, and send the facial feature existence result to the server. This allows the server to extract the features to be transmitted based on the facial feature existence result after receiving the facial feature existence result. In other words, the server determines the features to be transmitted corresponding to each selected segment based on the existence result, the selected segments corresponding to each segment identification in the video segment identification sequence, and the audio segments corresponding to each selected segment.

[0083] S220, sequentially receiving the features to be transmitted of the selected segments corresponding to each segment identifier in the video segment identification sequence sent by the server and the audio segments, synthesizing and playing the target video according to the received features to be transmitted, the audio segments, the decoder parameters of the pre-stored lip decoder, the action video segments, the lip position timing sequence corresponding to the action video segments, the silent video segments and the video segment identification sequence.

[0084] Specifically, the features to be transmitted of the selected segments corresponding to each segment identifier in the video segment identification sequence sent by the server and the audio segments are received in sequence. In the order of reception, the received features to be transmitted and the decoder parameters of the pre-stored lip decoder can be used in sequence to synthesize the lip area video segments that need to be superimposed and played. In the order of the received features to be transmitted, the corresponding selected segments and the corresponding lip position timing sequence are determined in the video segment identification sequence. The lip area video segments that need to be superimposed and played are superimposed on the corresponding selected segments according to the determined corresponding lip position timing sequence, and the corresponding audio segments are played synchronously, so as to achieve the synthesis and playback of each part of the video in sequence, and the video of each part is continuously combined into the target video.

[0085] It is understandable that the client can process a set of features to be transmitted of the selected segments upon receiving them, without having to wait for all features to be transmitted of the selected segments to be transmitted to be transmitted, thereby improving the real-time performance of the target video playback.

[0086] Based on the above example, the target video can be synthesized and played based on the received features to be transmitted, the audio clip, the pre-stored decoder parameters of the lip decoder, the action video clips, the lip position timing sequence corresponding to each action video clip, the silent video clips, and the video clip identification sequence in the following manner:

[0087] For each feature to be transmitted, according to the order of the features to be transmitted and the video segment identification sequence, the selected segment corresponding to the feature to be transmitted and the lip position timing sequence corresponding to the feature to be transmitted are determined from the pre-stored action video segments;

[0088] In response to the facial feature existence result being present, the facial feature and the feature to be transmitted are input into a lip decoder corresponding to the decoder parameter to obtain a lip area video clip; in response to the facial feature existence result being absent, the feature to be transmitted is input into a lip decoder corresponding to the decoder parameter to obtain a lip area video clip;

[0089] The playback position is determined according to the timing sequence of the lip position corresponding to the feature to be transmitted, and when the selected segment and audio segment are played synchronously, the lip area video segment is superimposed and played at the playback position of the selected segment according to the timing sequence to synthesize and play the target video of the part corresponding to the feature to be transmitted and the audio segment.

[0090] The order can be the sorted position of the features to be transmitted received during the synthesis and playback of the target video, that is, the number of the received set of features to be transmitted. The lip-sync area video clip is an animation clip of the lip-sync area generated by the lip-sync decoder. The playback position is the video position when the lip-sync area video clip is played overlaid.

[0091] Specifically, for each feature to be transmitted, a determination is made as to which group of features to be transmitted it is received during the current synthesis process. In other words, the order of the features to be transmitted is determined. The corresponding segment identifiers are then determined from the video segment identifier sequence according to this order. Furthermore, the action video segment corresponding to the segment is determined from the pre-stored action video segments as the selected segment corresponding to the feature to be transmitted. Furthermore, the lip position timing sequence corresponding to the selected segment can be determined from the storage space as the lip position timing sequence corresponding to the feature to be transmitted. If the facial feature existence result is positive, facial feature extraction is not required. The facial feature and the features to be transmitted can be input into the lip decoder corresponding to the decoder parameters to obtain a lip area video segment. If the facial feature existence result is negative, it indicates that the features to be transmitted contain all the required features. Therefore, the features to be transmitted are directly input into the lip decoder corresponding to the decoder parameters to obtain a lip area video segment. Based on the lip position timing sequence corresponding to the features to be transmitted, the playback position required for continuously playing the lip area video segment can be determined. When playing the selected clip and audio clip synchronously, the lip area video clip is superimposed and played in time sequence at the playback position of the selected clip, so as to synthesize and play part of the target video on the client, thereby improving the real-time performance of the synthesis and playback of the target video. It can be understood that the synthesized and played part of the target video is the part of the target video corresponding to the currently received features to be transmitted and audio clip.

[0092] For example, upon receiving a video segment identifier sequence, the client can retrieve the corresponding selected segments and lip position sequences, while simultaneously waiting for the corresponding lip feature data stream (features to be transmitted) to be sent from the server. After receiving the lip feature data stream corresponding to the action video, the client determines the corresponding order and, based on the order and the video segment identifier sequence, assigns the order to the selected segment x. The lip feature data stream (or lip feature data stream or pre-stored corresponding facial features) is upsampled by the lip decoder to generate a lip region video segment x'. This lip region video segment x' is then superimposed onto the corresponding selected segment x according to the corresponding lip position sequence and played simultaneously, with the corresponding audio segment playing synchronously during playback.

[0093] It is understandable that, in order to improve the coherence between the synthesized target video parts, smoothing processing may be performed between the target video parts that are continuously synthesized and played.

[0094] Based on the above example, after playing the lip-sync area video clips at the playback position of the selected clip in time sequence, if the features to be transmitted and the audio clip corresponding to the next selected clip have not been received, in order to ensure the continuity of the target video and avoid lag, the pre-stored silent video clips can be used for processing. Specifically, it can be:

[0095] If the features to be transmitted and the audio segment corresponding to the next selected segment are not received, at least one silent video segment is determined from the pre-stored silent video segments, and the at least one determined silent video segment is played in sequence until the features to be transmitted and the audio segment corresponding to the next selected segment are received.

[0096] Specifically, if the features and audio clips to be transmitted corresponding to the next selected segment are not received, a silent video segment can be played first while waiting for the transmission to occur to avoid video freezes. Specifically, at least one silent video segment is determined from the pre-stored silent video segments and played sequentially until the features and audio clips to be transmitted corresponding to the next selected segment are received and the next portion of the target video is synthesized. At this point, playback of the silent video segment is stopped and playback of the synthesized portion of the target video continues.

[0097] For example, after the client finishes playing a selected segment, if the lip feature data stream (features to be transmitted) corresponding to the next selected segment has not been received due to bandwidth limitations or slow server task processing, then the silent video segment can be played to avoid the situation where the digital human video is stuck and waiting.

[0098] While ensuring the quality of digital human videos, the above method effectively reduces hardware resource requirements, saves transmission bandwidth, and improves the smoothness of video generation. It proposes directly streaming the lip features back to the client, which then upsamples and decodes the lip features and displays them on the video layer. At the same time, the client monitors the progress of the service port lip feature data generation in real time. When it detects that the synthesis speed is slow, the client inserts a silent video clip to wait for the server's progress, thereby improving the smoothness of the client's digital human display experience.

[0099] The present invention has the following technical effects: a video segment identification sequence sent by a server is received at a client, and the existence result of the facial features is sent to the server, so that the server determines the to-be-transmitted features corresponding to each selected segment according to the existence result, the selected segment corresponding to each segment identification in the video segment identification sequence, and the audio segment corresponding to each selected segment, and sequentially receives the to-be-transmitted features and audio segments of the selected segment corresponding to each segment identification in the video segment identification sequence sent by the server, and synthesizes and plays the target video according to the received to-be-transmitted features, audio segments, pre-stored decoder parameters of a lip-sync decoder, each action video segment, the lip-sync position timing sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence, thereby achieving the effect of lip-sync synthesis on the server side, reducing the processing task load of the server side, and improving the video processing efficiency.

[0100] Example 3

[0101] Figure 4 FIG is a structural diagram of a data transmission processing system for a digital human-generated video provided by an embodiment of the present invention. Figure 4 As shown, the data transmission and processing system for digital human-generated videos includes: a server 310 and a client 320; the server 310 executes the steps of the data transmission and processing method for digital human-generated videos provided in any embodiment of the present invention; the client 320 executes the steps of the data transmission and processing method for digital human-generated videos provided in any embodiment of the present invention.

[0102] The system of the above embodiment is used to implement the data transmission and processing method of the digital human-generated video corresponding to any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0103] Example 4

[0104] Based on the above examples, both the server and the client can be electronic devices. Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 5 As shown, the electronic device 400 includes one or more processors 401 and a memory 402 .

[0105] The processor 401 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.

[0106] Memory 402 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and processor 401 may execute these program instructions to implement the data transmission and processing method for digital human-generated video according to any embodiment of the present invention described above and / or other desired functions. Various contents, such as initial external parameters and thresholds, may also be stored in the computer-readable storage medium.

[0107] In one example, electronic device 400 may further include an input device 403 and an output device 404, which are interconnected via a bus system and / or other connection mechanisms (not shown). Input device 403 may include, for example, a keyboard, a mouse, etc. Output device 404 may output various information to the outside, including warning information, braking force, etc. Output device 404 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.

[0108] Of course, to simplify, Figure 5 Only some of the components related to the present invention in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 400 may further include any other appropriate components according to specific application scenarios.

[0109] Example 5

[0110] In addition to the above-mentioned methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the data transmission processing method for digital human-generated video provided by any embodiment of the present invention.

[0111] The computer program product may be written in any combination of one or more programming languages to implement the operations of embodiments of the present invention, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0112] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps of the data transmission processing method for digital human-generated video provided by any embodiment of the present invention.

[0113] The computer-readable storage medium may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0114] It should be noted that the terms used in the present invention are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "an", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method or device comprising a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also include elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.

[0115] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A data transmission and processing method for a digital human-generated video, characterized in that: Applied to the server, including: Determine a video segment identification sequence based on the audio duration of the target audio and each action video segment, and send the video segment identification sequence to the client; Determining the audio segment corresponding to each selected segment according to the target audio and the duration of the selected segment corresponding to each segment identifier in the video segment identifier sequence; Get the existence result of the client's facial features; In response to the existence result being that the image frame sequence in each selected segment is not present, inputting the image frame sequence in the selected segment into a pre-trained identity encoder to obtain facial features corresponding to the selected segment, inputting the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain speech features corresponding to the selected segment, and using the facial features and the speech features as features to be transmitted corresponding to the selected segment; In response to the existence result being existence, for each selected segment, inputting the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain speech features corresponding to the selected segment, and using the speech features as features to be transmitted corresponding to the selected segment; Sending the features to be transmitted and the audio segments corresponding to the selected segments to the client in sequence according to the order of the video segment identification sequence, so that the client synthesizes and plays the target video based on the received features to be transmitted, the audio segments, the pre-stored decoder parameters of the lip-sync decoder, the action video segments, the lip-sync position timing sequence corresponding to the action video segments, the silent video segments, and the video segment identification sequence; The action video segments and the silent video segments are determined based on the original video segmentation of the target digital human.

2. The method according to claim 1, characterized in that Before determining the video segment identification sequence according to the audio duration of the target audio and each action video segment, the method further includes: Determining at least one action video segment and at least one silent video segment based on the original video of the target digital human and the shortest duration of the video segment; For each action video clip, determining a lip position timing sequence and a clip identifier corresponding to the action video clip; Sending decoder parameters of a pre-trained lip-sync decoder, each silent video segment, each action video segment, and the corresponding lip position timing sequence and segment identifier of each action video segment to the client; The identity encoder, the speech encoder and the lip-sync decoder constitute a lip-sync generation model, and the lip-sync generation model is trained based on each action video clip and the lip-sync position time sequence corresponding to each action video clip.

3. The method according to claim 2, characterized in that The step of determining at least one action video segment and at least one silent video segment based on the original video of the target digital human and the shortest duration of the video segments includes: Taking the starting video frame in the original video of the target digital human as the first video frame, and determining the position of the human in the first video frame; If the time interval between the first video frame and the last video frame in the original video is greater than or equal to the shortest duration of the video segment, select a video frame in the original video whose time interval with the first video frame is equal to the shortest duration of the video segment as a candidate video frame, and determine the position of the person in the candidate video frame; In response to a positional deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame being greater than or equal to a preset deviation, if the to-be-selected video frame is not the last video frame in the original video, using the next video frame of the to-be-selected video frame as the to-be-selected video frame, and returning to the step of determining the position of the person in the to-be-selected video frame until the positional deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame is less than the preset deviation; In response to a positional deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame being less than the preset deviation, the to-be-selected video frame is used as a second video frame corresponding to the first video frame, a video segment between the first video frame and the second video frame is used as a target video segment, and a speaking situation of the target digital person in the target video segment is determined; Using the next video frame of the second video frame as a new first video frame, and returning to the step of determining the position of the person in the first video frame; The target digital person's speaking situation is the target video clip of the target digital person speaking as each action video clip; If there is at least one target digital person whose speaking condition is that the target digital person is not speaking, then the target video segments whose speaking condition is that the target digital person is not speaking are used as silent video segments; If any target digital human speaks, then one is randomly selected from each action video segment as the video segment to be processed, the mouth of the target digital human in the video segment to be processed is processed to a closed state, and the processed video segment to be processed is used as a silent video segment.

4. The method according to claim 2, characterized in that Also includes: For each action video clip, inputting the image frame sequence in the action video clip into a pre-trained identity encoder to obtain facial features corresponding to the action video clip; The facial features corresponding to each action video clip are sent to the client.

5. The method according to claim 1, wherein The step of determining a video segment identification sequence based on the audio duration of the target audio and each action video segment includes: Initialize a video segment identification sequence, and use the sum of the segment durations of the selected segments corresponding to the segment identifiers in the video segment identification sequence as the existing duration; If the existing duration is less than the audio duration of the target audio, randomly select one from the segment identifiers corresponding to the action video segments, add the segment identifier to the video segment identifier sequence, and return to the step of taking the sum of the segment durations of the selected segments corresponding to the segment identifiers in the video segment identifier sequence as the existing duration, until the existing duration is greater than or equal to the audio duration of the target audio; In a case where the existing duration is greater than or equal to the audio duration of the target audio, the video segment identification sequence is obtained.

6. A data transmission and processing method for a digital human-generated video, characterized in that: Applied to the client, including: Receive a video segment identification sequence sent by a server, and send the existence result of the facial features to the server, so that the server responds to the existence result as non-existence, then, for each selected segment, input the image frame sequence in the selected segment into a pre-trained identity encoder to obtain the facial features corresponding to the selected segment, input the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain the speech features corresponding to the selected segment, and use the facial features and the speech features as the features to be transmitted corresponding to the selected segment; in response to the existence result as existence, then, for each selected segment, input the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain the speech features corresponding to the selected segment, and use the speech features as the features to be transmitted corresponding to the selected segment; wherein, the video segment identification sequence is determined by the server according to the audio duration of the target audio and each action video segment, and the audio segment corresponding to each selected segment is determined by the server according to the target audio and the duration of the selected segment corresponding to each segment identifier in the video segment identification sequence; The system sequentially receives the features to be transmitted of the selected segments corresponding to each segment identifier in the video segment identifier sequence sent by the server, and synthesizes and plays the target video based on the received features to be transmitted, the audio segment, the decoder parameters of the pre-stored lip decoder, the action video segments, the lip position timing sequence corresponding to each action video segment, the silent video segments, and the video segment identifier sequence.

7. The method according to claim 6, characterized in that The method of synthesizing and playing a target video based on the received features to be transmitted, the audio clip, the pre-stored decoder parameters of the lip decoder, the action video clips, the lip position timing sequence corresponding to the action video clips, the silent video clips, and the video clip identification sequence includes: For each feature to be transmitted, according to the order of the features to be transmitted and the video segment identification sequence, determining the selected segment corresponding to the feature to be transmitted and the lip position timing sequence corresponding to the feature to be transmitted from the pre-stored action video segments; In response to the facial feature existence result being present, the facial feature and the feature to be transmitted are input into a lip decoder corresponding to the decoder parameter to obtain a lip area video segment; in response to the facial feature existence result being absent, the feature to be transmitted is input into a lip decoder corresponding to the decoder parameter to obtain a lip area video segment; The playback position is determined according to the timing sequence of the lip position corresponding to the feature to be transmitted, and when the selected segment and the audio segment are played synchronously, the lip area video segment is played in a time-superimposed manner at the playback position of the selected segment to synthesize and play the target video of the portion corresponding to the feature to be transmitted and the audio segment.

8. The method according to claim 7, characterized in that After the lip-sync area video clip is played at the playback position of the selected clip in a time sequence, the method further includes: If the features to be transmitted and the audio segment corresponding to the next selected segment are not received, at least one silent video segment is determined from the pre-stored silent video segments, and the at least one determined silent video segment is played in sequence until the features to be transmitted and the audio segment corresponding to the next selected segment are received.

9. A data transmission and processing system for digital human-generated video, characterized in that: include: Server and client; The server performs the steps of the data transmission processing method for digital human-generated video according to any one of claims 1 to 5; The client executes the steps of the data transmission processing method for digital human-generated video according to any one of claims 6 to 8.

Citation Information

Patent Citations

  • Digital human-driven rendering method and device, electronic equipment and storage medium

    CN116863039A

  • Method and device for increasing response speed of remote digital human

    CN117294905A