Data transmission processing method and system for digital human-generated video

By determining the video clip identification sequence on the server and sending it to the client, the client synthesizes and plays videos based on the received characteristics and audio clips, solving the problem of hardware resources and bandwidth occupation during the generation and playback of digital human videos in the prior art, realizing more efficient data transmission and smooth video playback.

CN120050452AActive Publication Date: 2025-05-27国家超级计算天津中心
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510509134.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-27
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing digital video generation method requires a large amount of hardware resources and transmission bandwidth during synthesis and playback, resulting in high equipment costs, high hardware performance requirements and poor video fluency, especially when network conditions are poor, it is prone to lag.

Method used

By determining the video clip identification sequence on the server and sending it to the client, the client synthesizes and plays the target video based on the received characteristics to be transmitted, audio clips and pre-stored lip decoder parameters, thus achieving the effect of reducing data transmission and improving video fluency.

Benefits of technology

It reduces the server-side task processing and data transmission, improves the fluency of digital human video synthesis and playback, and reduces the hardware resource requirements and transmission bandwidth usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050452A_ABST
    Figure CN120050452A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, and discloses a data transmission processing method and system for a digital human generated video, and the method is applied to a server, and comprises the steps: determining a video clip identification sequence according to the audio duration of a target audio and each action video clip, and transmitting the video clip identification sequence to a client; determining an audio clip corresponding to each selected clip according to the duration of the selected clip corresponding to each clip identifier in the target audio and video clip identifier sequence; obtaining an existence result of facial features of the client, and determining to-be-transmitted features corresponding to the selected segments according to the existence result, the selected segments and the audio segments; and the to-be-transmitted features and the audio clips are sequentially sent to the client according to the sequence of the video clip identification sequence, so that the client synthesizes and plays the target video according to the received data and the pre-stored data, and the effect of reducing the data transmission amount required when the digital human video is synthesized and played is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data transmission processing method and system for generating videos of digital humans. Background Art

[0002] The generation of videos of digital humans is a key link for their vivid display. At present, the common way to generate videos of digital humans is that the server first synthesizes lip shapes according to speech, and then accurately pastes the lip shapes onto the video characters, and then synthesizes a complete video, and sends the video to the client for display.

[0003] However, from lip shape synthesis to the complete synthesis of the final video, the whole process requires a large amount of hardware resources, which not only increases the equipment cost, but also places extremely high requirements on the hardware performance. In addition, a large amount of video data is synthesized on the server side and then transmitted to the client, which also occupies a large amount of transmission bandwidth and affects the smoothness of video display. Especially in the case of poor network conditions, the phenomenon of freezing occurs frequently.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a data transmission processing method and system for generating videos of digital humans, realizing the effect of reducing the amount of data transmission required for the synthesis and playback of digital human videos, and improving the smoothness of the synthesis and playback of digital human videos.

[0006] An embodiment of the present invention provides a data transmission processing method for generating videos of digital humans, which is applied to a server, and the method includes:

[0007] Determine a video segment identification sequence according to the audio duration of the target audio and each action video segment, and send the video segment identification sequence to the client;

[0008] Determine the audio segment corresponding to each selected segment according to the target audio and the duration of the selected segment corresponding to each segment identification in the video segment identification sequence;

[0009] Obtain the existence result of the facial features of the client, and determine the feature to be transmitted corresponding to each selected segment according to the existence result, each selected segment, and the audio segment corresponding to each selected segment;

[0010] Send the to-be-transmitted features corresponding to each selected segment and the audio segments to the client in the order of the video segment identification sequence, so that the client synthesizes and plays the target video according to the received to-be-transmitted features, audio segments, decoder parameters of the pre-stored mouth shape decoder, each action video segment, the mouth shape position time sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence;

[0011] Among them, each action video segment and each silent video segment are determined by splitting the original video of the target digital human.

[0012] An embodiment of the present invention provides a data transmission and processing method for a digital human to generate a video, which is applied to a client. The method includes:

[0013] Receive the video segment identification sequence sent by the server, and send the existence result of the facial feature to the server, so that the server determines the to-be-transmitted features corresponding to each selected segment according to the existence result, the selected segments corresponding to each segment identification in the video segment identification sequence, and the audio segments corresponding to each selected segment; wherein, the video segment identification sequence is determined by the server according to the audio duration of the target audio and each action video segment, and the audio segments corresponding to each selected segment are determined by the server according to the target audio and the duration of the selected segments corresponding to each segment identification in the video segment identification sequence;

[0014] Sequentially receive the to-be-transmitted features and audio segments of the selected segments corresponding to each segment identification in the video segment identification sequence sent by the server, and synthesize and play the target video according to the received to-be-transmitted features, audio segments, decoder parameters of the pre-stored mouth shape decoder, each action video segment, the mouth shape position time sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence.

[0015] An embodiment of the present invention provides a data transmission and processing system for a digital human to generate a video, which is characterized by including: a server and a client; the server executes the steps of the data transmission and processing method for a digital human to generate a video in any embodiment; the client executes the steps of the data transmission and processing method for a digital human to generate a video in any embodiment.

[0016] The embodiment of the present invention has the following technical effects:

[0017] On the server side, according to the audio duration of the target audio and each action video segment, determine the video segment identification sequence, and send the video segment identification sequence to the client to indicate the playing order of the client. Furthermore, according to the target audio and the duration of the selected segments corresponding to each segment identification in the video segment identification sequence, determine the audio segments corresponding to each selected segment, so as to split the target audio as required for subsequent block transmission. Obtain the existence result of the facial features of the client, and according to the existence result, each selected segment, and the audio segments corresponding to each selected segment, determine the features to be transmitted corresponding to each selected segment to extract the feature data to be transmitted. Finally, send the features to be transmitted and the audio segments corresponding to each selected segment to the client in the order of the video segment identification sequence, so that the client synthesizes and plays the target video according to the received features to be transmitted, audio segments, the decoder parameters of the pre-stored lip decoder, each action video segment, the lip position time sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence, achieving the effects of reducing the server-side task processing volume and reducing the data transmission volume between the server and the client, and improving the fluency of digital human video synthesis and playback on the subsequent client. Brief Description of the Drawings

[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of a data transmission and processing method for generating a digital human video provided by an embodiment of the present invention;

[0020] Figure 2 It is a schematic diagram of each feature point corresponding to the face position and the lip region of the person provided by an embodiment of the present invention;

[0021] Figure 3 It is a flowchart of another data transmission and processing method for generating a digital human video provided by an embodiment of the present invention;

[0022] Figure 4 It is a schematic structural diagram of a data transmission and processing system for generating a digital human video provided by an embodiment of the present invention;

[0023] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope protected by the present invention.

[0025] The data transmission processing method for digital human-generated videos provided by the embodiments of the present invention is mainly applicable to the situation of reducing the data transmission volume during the generation and playback of digital human videos and improving the smoothness of digital human videos. The data transmission processing for digital human-generated videos provided by the embodiments of the present invention can be executed by the server and / or the client.

[0026] Embodiment 1

[0027] Figure 1 is a flowchart of a data transmission processing method for digital human-generated videos provided by the embodiments of the present invention. Refer to Figure 1 This data transmission processing method for digital human-generated videos is applied to the server and specifically includes:

[0028] S110. Determine a video segment identification sequence according to the audio duration of the target audio and each action video segment, and send the video segment identification sequence to the client.

[0029] Among them, the target audio is the audio required to synthesize the final digital human video, that is, the audio for which the digital human needs to lip-sync. The audio duration is the duration required to normally play the target audio. Each action video segment and each silent video segment are determined based on the splitting of the original video of the target digital human. The target digital human is the digital human used in the final synthesized video, and the original video is the video obtained when the target digital human is recorded and filmed. The action video segment is the video segment split from the part of the original video where the target digital human is speaking, and the number of action video segments is at least one. The silent video segment can be the video segment split from the part of the original video where the target digital human is not speaking, or the video segment obtained by processing an action video segment into a non-speaking video segment, and the number of silent video segments is at least one. The video segment identification sequence is a sequence composed of the segment identifiers of the action video segments required for subsequent synthesis of the digital human video in order, where the segment identifiers can appear repeatedly.

[0030] Specifically, randomly select one from each action video segment, add its corresponding segment identifier to the video segment identifier sequence, and calculate whether the total duration of the action video segments corresponding to each segment identifier in the video segment identifier sequence is greater than or equal to the audio duration of the target audio. If so, determine that the video segment identifier sequence is obtained; if not, randomly select one from each action video segment, add its corresponding segment identifier to the video segment identifier sequence until the total duration of the action video segments corresponding to each segment identifier in the video segment identifier sequence is greater than or equal to the audio duration of the target audio, and obtain the video segment identifier sequence. Furthermore, send the video segment identifier sequence to the client so that the corresponding action video segments can be retrieved according to the video segment identifier sequence in the client later to synthesize and play the target video.

[0031] Based on the above example, the video segment identifier sequence can be determined according to the audio duration of the target audio and each action video segment in the following way:

[0032] Initialize the video segment identifier sequence, and use the sum of the segment durations of the selected segments corresponding to each segment identifier in the video segment identifier sequence as the existing duration;

[0033] When the existing duration is less than the audio duration of the target audio, randomly select one from the segment identifiers corresponding to each action video segment, add it to the video segment identifier sequence, and return to execute the step of using the sum of the segment durations of the selected segments corresponding to each segment identifier in the video segment identifier sequence as the existing duration until the existing duration is greater than or equal to the audio duration of the target audio;

[0034] When the existing duration is greater than or equal to the audio duration of the target audio, obtain the video segment identifier sequence.

[0035] Among them, the segment identifier is the identifier of each action video segment, used to distinguish different action video segments. The existing segment is the action video segment determined according to each segment identifier in the video segment identifier sequence, and can appear repeatedly to reflect the randomness and diversity of the digital human actions. The existing duration is the sum of the segment durations of the selected segments corresponding to each segment identifier in the video segment identifier sequence. The selected segment is the action video segment corresponding to each segment identifier in the video segment identifier sequence.

[0036] Specifically, initialize the video segment identification sequence to empty it. Furthermore, take the sum of the segment durations of the selected segments corresponding to each segment identification in the video segment identification sequence as the existing duration. Determine the size relationship between the existing duration and the audio duration of the target audio. If it is less, randomly select one from the segment identifications corresponding to each action video segment, add it to the video segment identification sequence, and return to execute the step of taking the sum of the segment durations of the selected segments corresponding to each segment identification in the video segment identification sequence as the existing duration until the existing duration is greater than or equal to the audio duration of the target audio. It can be understood that during the selection process, there is no restriction on selecting duplicate or non-duplicate segment identifications. If it is greater than or equal, use the current video segment identification sequence as the finally determined video segment identification sequence.

[0037] Based on the above example, before determining the video segment identification sequence according to the audio duration of the target audio and each action video segment, it is also necessary to determine each action video segment and each silent video segment according to the original video of the target digital human, determine the timing sequence of mouth shape positions and segment identifications corresponding to each action video segment, train the mouth shape generation model, and send the required information to the client in advance for storage, so as to directly call it during subsequent use and avoid a large amount of data transmission during the synthesis process. Specifically, it can be:

[0038] Determine at least one action video segment and at least one silent video segment according to the original video of the target digital human and the shortest duration of the video segment;

[0039] For each action video segment, determine the timing sequence of mouth shape positions and segment identifications corresponding to the action video segment;

[0040] Send the decoder parameters of the pre-trained mouth shape decoder, each silent video segment, each action video segment, and the timing sequence of mouth shape positions and segment identifications corresponding to each action video segment to the client.

[0041] Among them, the identity encoder, the speech encoder, and the mouth shape decoder constitute the mouth shape generation model (Wav2Lip), and the mouth shape generation model is trained based on each action video segment and the timing sequence of mouth shape positions corresponding to each action video segment respectively. The identity encoder is used to extract the features of the facial images in the video, that is, the facial features (such as those that can reflect the facial structure, skin texture, etc.). The speech encoder is used to process the audio signal and extract the speech features. The mouth shape decoder is used to fuse the facial features and the speech features and generate the mouth shape animation of the mouth shape part. The shortest duration of the video segment is the shortest duration of each action video segment and each silent video segment set in advance. The decoder parameters are various parameter values required for the mouth shape decoder in the trained mouth shape generation model.

[0042] Specifically, the original video of the target digital human is split with the shortest duration of the video segment as the limit, obtaining at least one action video segment and at least one silent video segment. For each action video segment, identify the face positions in each video frame, determine each feature point, and regard the area enclosed by connecting the feature points in the lower half of the face position as the mouth shape area of the person. Determine the time sequence of the feature points above and within this area, construct the time sequence of the mouth shape position corresponding to this action video segment, and determine the corresponding segment identifier. The mouth shape generation model is trained based on each action video segment and the time sequence of the mouth shape position corresponding to each action video segment respectively. Send the decoder parameters of the pre-trained mouth shape decoder, each silent video segment, each action video segment, the time sequence of the mouth shape position corresponding to each action video segment respectively, and the segment identifier to the client, so that the client can subsequently retrieve the corresponding data from the storage space according to the received video segment identifier sequence and use the corresponding mouth shape decoder.

[0043] Exemplarily, the schematic diagrams of each feature point corresponding to the face position and the mouth shape area of the person are as Figure 2 shown. Among them, the number of feature points corresponding to the face position is 68.

[0044] Based on the above example, at least one action video segment and at least one silent video segment can be determined according to the original video of the target digital human and the shortest duration of the video segment in the following way:

[0045] Take the starting video frame in the original video of the target digital human as the first video frame, and determine the person position in the first video frame;

[0046] When the time interval between the first video frame and the last video frame in the original video is greater than or equal to the shortest duration of the video segment, take the video frame in the original video whose time interval from the first video frame is equal to the shortest duration of the video segment as the candidate video frame, and determine the person position in the candidate video frame;

[0047] In response to the position deviation between the person position in the candidate video frame and the person position in the first video frame being greater than or equal to the preset deviation, when the candidate video frame is not the last video frame in the original video, take the next video frame of the candidate video frame as the candidate video frame, and return to execute the step of determining the person position in the candidate video frame until the position deviation between the person position in the candidate video frame and the person position in the first video frame is less than the preset deviation;

[0048] If the position deviation between the position of the person in the candidate video frame and the position of the person in the first video frame is less than a preset deviation, then the candidate video frame is taken as the second video frame corresponding to the first video frame, the video segment between the first video frame and the second video frame is taken as a target video segment, and the speaking situation of the target digital human in the target video segment is determined;

[0049] Take the next video frame of the second video frame as the new first video frame, and return to execute the step of determining the position of the person in the first video frame;

[0050] Take the target video segments in which the speaking situation of each target digital human is that the target digital human is speaking as each action video segment;

[0051] If there is at least one target digital human speaking situation where the target digital human is not speaking, then take the target video segments in which the speaking situation of each target digital human is that the target digital human is not speaking as each silent video segment;

[0052] If the speaking situation of any target digital human is that the target digital human is speaking, then randomly select one from each action video segment as the video segment to be processed, process the mouth of the target digital human in the video segment to be processed into a closed state, and take the processed video segment to be processed as a silent video segment.

[0053] Among them, the first video frame is the first video frame in the target video segment. The second video frame is the last video frame in the target video segment. The person position is the position of the target digital human in the video frame. The candidate video frame is the candidate end video frame corresponding to the first video frame, that is, the candidate second video frame corresponding to the first video frame. The position deviation is the absolute value of the distance between the person position of the first video frame and the person position of the candidate video frame. If multiple position points describe the person position, it can be the sum of the absolute values of the distances of each position point. The target video segment is a segment split from the original video, where the position deviation between the first video frame and the last video frame of the person position is less than the preset deviation. The preset deviation is used to judge whether the person positions are basically coincident. The speaking situation of the target digital human includes that the target digital human is speaking and that the target digital human is not speaking, which can be recognized according to the audio corresponding to the target video segment. The video segment to be processed is an action video segment to be processed into a silent video segment.

[0054] Specifically, the starting video frame in the original video of the target digital human is used as the first video frame, and the position of the person in the first video frame is determined. First, it is judged whether the time interval between the first video frame and the last video frame in the original video is greater than or equal to the shortest duration of the video segment. If not, it means that the remaining part of the video cannot be split into the target video segment due to insufficient duration. If so, video segment splitting can be performed. The video frame in the original video whose time interval from the first video frame is equal to the shortest duration of the video segment is used as the candidate video frame, and the position of the person in the candidate video frame is identified and determined. Furthermore, the position deviation between the position of the person in the candidate video frame and the position of the person in the first video frame can be calculated, such as the sum of the absolute values of the distances between the corresponding position points corresponding to the determined person positions. It is judged whether the position deviation is greater than or equal to the preset deviation. If so, it is determined that the positions of the person in the first video frame and the candidate video frame do not basically coincide, and the candidate video frame cannot be used as the second video frame corresponding to the first video frame. Therefore, the next video frame of the candidate video frame is used as the new candidate video frame, and the step of the position of the person in the candidate video frame is returned and executed until the position deviation between the position of the person in the candidate video frame and the position of the person in the first video frame is less than the preset deviation. If not, it is determined that the positions of the person in the first video frame and the candidate video frame basically coincide, the candidate video frame can be used as the second video frame corresponding to the first video frame, the video segment between the first video frame and the second video frame is used as a target video segment, and the speaking situation of the target digital human in the target video segment is determined according to the audio in the target video segment. Furthermore, the next video frame of the second video frame is used as the new first video frame, and the loop is performed, that is, the step of determining the position of the person in the first video frame is returned and executed until the original video is split completely. Thus, it can be known that the target digital humans in the first video frame and the second video frame of each split target video segment basically coincide. After each target video segment is split, the target video segments with the speaking situation of the target digital human being that the target digital human is speaking are used as each action video segment. It is judged whether there is at least one target digital human speaking situation in each target video segment being that the target digital human is not speaking. If so, it means that there is a silent video segment, and the target video segments with the speaking situation of the target digital human being that the target digital human is not speaking are directly used as each silent video segment; if not, it means that there is currently no silent video segment, and a silent video segment needs to be produced. One of the action video segments is randomly selected as the video segment to be processed, which can be selected according to a random number, etc. The mouth of the target digital human in the video segment to be processed is recognized, and the animation of the mouth is processed into a closed state to simulate the mouth animation of the target digital human not speaking, and the processed video segment to be processed is used as the silent video segment.

[0055] Based on the above examples, the facial features corresponding to each action video segment can also be sent to the client in advance to reduce the data transmission volume during subsequent synthesis of the target video. Specifically, it can be as follows:

[0056] For each action video segment, input the image frame sequence in the action video segment into a pre-trained identity encoder to obtain the facial features corresponding to the action video segment;

[0057] Send the facial features corresponding to each action video segment to the client.

[0058] Among them, the image frame sequence is the part remaining after removing the audio from the action video segment. The facial features are the results of the identity encoder extracting features of the facial region based on the image frame sequence.

[0059] Specifically, for each action video segment, perform audio-visual separation on the action video segment to obtain the image frame sequence therein, input the image frame sequence into a pre-trained identity encoder, and output the facial features corresponding to the action video segment. Moreover, send the facial features corresponding to each action video segment to the client to facilitate pre-storing the facial features corresponding to each action video segment on the client.

[0060] It can be understood that since the data volume of the facial features corresponding to each action video segment is not large, therefore, it can be pre-transmitted to the client for storage, or the facial features can be determined and transmitted when determining and transmitting the features to be transmitted corresponding to the selected segments subsequently.

[0061] S120. Determine the audio segments corresponding to the selected segments according to the target audio and the durations of the selected segments corresponding to the segment identifiers in the video segment identifier sequence.

[0062] Among them, the audio segment is a part of the target audio intercepted according to the duration of the selected segment.

[0063] Specifically, intercept the target audio in sequence according to the durations of the selected segments corresponding to the segment identifiers in the video segment identifier sequence, and the audio segments corresponding to the selected segments can be obtained. It can be understood that the audio duration of the last audio segment may be less than the duration of the corresponding selected segment.

[0064] Optionally, more accurate splitting can be performed in combination with the audio information of the target audio, which can include: splitting the target audio into audio segments according to the durations of the selected segments corresponding to the segment identifiers in the video segment identifier sequence. Among them, the splitting positions of adjacent audio segments are at the speech pauses, and the duration of the audio segment is equivalent to the duration of the corresponding selected segment (the duration of the audio segment is less than or equal to the duration of the corresponding selected segment). Therefore, one selected segment can correspond to an oral position time sequence and an audio segment.

[0065] S130. Obtain the existence result of the facial features of the client, and determine the features to be transmitted corresponding to each selected segment according to the existence result, each selected segment, and the audio segment corresponding to each selected segment.

[0066] Among them, the existence result of the facial features is used to describe whether the facial features of the action video segment corresponding to the target digital human are pre-stored in the client, which can be existent or non-existent. The features to be transmitted are the features that need to be transmitted in real time when synthesizing and playing the target video of the target digital human according to the target audio.

[0067] Specifically, to obtain the existence result of the facial features of the client, it can be to send a corresponding facial feature existence recognition request to the client, and determine the existence result of the facial features of the client according to the feedback information fed back by the client based on the facial feature existence recognition request. According to the existence result, process each selected segment and the audio segment corresponding to each selected segment. When determining different existence results, the features to be transmitted corresponding to each selected segment can be understood as whether to add facial features to the features to be transmitted.

[0068] Based on the above example, the following method can be used to determine the features to be transmitted corresponding to each selected segment according to the existence result, each selected segment, and the audio segment corresponding to each selected segment:

[0069] In response to the existence result being non-existent, for each selected segment, input the image frame sequence in the selected segment into a pre-trained identity encoder to obtain the facial features corresponding to the selected segment, input the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain the speech features corresponding to the selected segment, and use the facial features and the speech features as the features to be transmitted corresponding to the selected segment;

[0070] In response to the existence result being existent, for each selected segment, input the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain the speech features corresponding to the selected segment, and use the speech features as the features to be transmitted corresponding to the selected segment.

[0071] Among them, the speech features are the result of feature extraction by the speech encoder based on the audio segment corresponding to the selected segment.

[0072] Specifically, if the result is non-existence, the features to be transmitted include the facial features and voice features of each selected segment. For each selected segment, perform audio-visual separation on the selected segment to obtain the image frame sequence therein, input the image frame sequence into a pre-trained identity encoder to obtain the facial features corresponding to the selected segment, and input the audio segment corresponding to the selected segment into a pre-trained voice encoder to obtain the voice features corresponding to the selected segment. Furthermore, the facial features and voice features can be used as the features to be transmitted corresponding to the selected segment. If the result is existence, the features to be transmitted only need to include the voice features of each selected segment. The method for obtaining the voice features of each selected segment is the same as the above situation and will not be elaborated here. Furthermore, the voice features are used as the features to be transmitted corresponding to the selected segment.

[0073] S140. Send the features to be transmitted corresponding to each selected segment and the audio segments to the client in sequence according to the order of the video segment identification sequence, so that the client can synthesize and play the target video based on the received features to be transmitted, audio segments, decoder parameters of the pre-stored mouth shape decoder, each action video segment, the mouth shape position time sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence.

[0074] Among them, the target video is the final digital human video, the audio therein is the target audio, and the mouth shape of the digital human corresponds to the target audio.

[0075] Specifically, the features to be transmitted corresponding to each selected segment and the audio segments are sent to the client in groups in sequence according to the order of the video segment identification sequence, so that the client can synthesize and play the video when receiving a group of data, improving the real-time performance. The client can synthesize and play the target video based on the received features to be transmitted and audio segments, combined with the decoder parameters of the pre-stored mouth shape decoder in the client, each action video segment, the mouth shape position time sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence.

[0076] The above method no longer performs complex video synthesis operations on the server side, but directly transmits the mouth shape feature data stream (features to be transmitted) to the client, greatly reducing the server-side processing tasks and the requirements for the server-side hardware computing power. Moreover, since what is transmitted is the mouth shape feature data rather than the complete video, the data volume is greatly reduced. During the network transmission process, the lower data volume requirement reduces the network transmission cost and improves the data transmission efficiency.

[0077] The present invention has the following technical effects: At the server side, according to the audio duration of the target audio and each action video segment, a video segment identification sequence is determined and sent to the client to indicate the playing order of the client. Furthermore, according to the target audio and the duration of the selected segments corresponding to each segment identification in the video segment identification sequence, the audio segments corresponding to each selected segment are determined to facilitate the splitting of the target audio as required for subsequent block transmission. The existence result of the facial features of the client is obtained, and according to the existence result, each selected segment, and the audio segments corresponding to each selected segment, the features to be transmitted corresponding to each selected segment are determined to extract the feature data to be transmitted. Finally, the features to be transmitted and the audio segments corresponding to each selected segment are sequentially sent to the client in the order of the video segment identification sequence, so that the client synthesizes and plays the target video according to the received features to be transmitted, audio segments, decoder parameters of the pre-stored lip decoder, each action video segment, the lip position time sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence, achieving the effects of reducing the server task processing volume and reducing the data transmission volume between the server and the client, and improving the smoothness of digital human video synthesis and playing on the subsequent client.

[0078] Embodiment 2

[0079] Figure 3 It is a flowchart of another data transmission and processing method for generating a digital human video provided by an embodiment of the present invention. Refer to Figure 3 , this data transmission and processing method for generating a digital human video is applied to the client and specifically includes:

[0080] S210. Receive the video segment identification sequence sent by the server, and send the existence result of the facial features to the server, so that the server determines the features to be transmitted corresponding to each selected segment according to the existence result, each selected segment corresponding to each segment identification in the video segment identification sequence, and the audio segments corresponding to each selected segment.

[0081] Among them, the video segment identification sequence is determined by the server according to the audio duration of the target audio and each action video segment, and the audio segments corresponding to each selected segment are determined by the server according to the target audio and the duration of the selected segments corresponding to each segment identification in the video segment identification sequence.

[0082] Specifically, the client can receive the sequence of video segment identifiers sent by the server, and can determine each action video segment corresponding to the sequence of video segment identifiers from the storage space, and determine whether corresponding facial features are stored for these determined action video segments, obtain the existence result of the facial features, and send the existence result of the facial features to the server. In this way, after receiving the existence result of the facial features, the server can extract the features to be transmitted according to this existence result, that is, the server determines the features to be transmitted corresponding to each selected segment according to the existence result, the selected segments corresponding to the segment identifiers in the sequence of video segment identifiers, and the audio segments corresponding to each selected segment.

[0083] S220. Sequentially receive the features to be transmitted and the audio segments corresponding to the selected segments corresponding to the segment identifiers in the sequence of video segment identifiers sent by the server, and synthesize and play the target video according to the received features to be transmitted, audio segments, decoder parameters of the pre-stored lip decoder, each action video segment, the lip position time sequence corresponding to each action video segment, each silent video segment, and the sequence of video segment identifiers.

[0084] Specifically, sequentially receive the features to be transmitted and the audio segments corresponding to the selected segments corresponding to the segment identifiers in the sequence of video segment identifiers sent by the server. According to the received order, the received features to be transmitted and the decoder parameters of the pre-stored lip decoder can be used to synthesize the lip region video segments to be superimposed and played in sequence. According to the order of the received features to be transmitted, determine the corresponding selected segments and the corresponding lip position time sequence in the sequence of video segment identifiers, superimpose the lip region video segments to be superimposed and played on the corresponding selected segments according to the determined corresponding lip position time sequence, and synchronously play the corresponding audio segments, so as to synthesize and play each part of the video in sequence, and the video of each part is continuously combined into the target video.

[0085] It can be understood that the client can process when receiving the features to be transmitted of a group of selected segments, without waiting for all the features to be transmitted of the selected segments to be transmitted, improving the real-time performance of the target video playback.

[0086] Based on the above example, the target video can be synthesized and played in the following way according to the received features to be transmitted, audio segments, decoder parameters of the pre-stored lip decoder, each action video segment, the lip position time sequence corresponding to each action video segment, each silent video segment, and the sequence of video segment identifiers:

[0087] For each feature to be transmitted, according to the order of the features to be transmitted and the video segment identification sequence, determine the selected segment corresponding to the feature to be transmitted and the lip position time sequence corresponding to the feature to be transmitted from each pre-stored action video segment;

[0088] In response to the presence result of the facial feature being present, input the facial feature and the feature to be transmitted into the lip decoder corresponding to the decoder parameters to obtain the lip region video segment; in response to the presence result of the facial feature being absent, input the feature to be transmitted into the lip decoder corresponding to the decoder parameters to obtain the lip region video segment;

[0089] Determine the playback position according to the lip position time sequence corresponding to the feature to be transmitted, and when synchronously playing the selected segment and the audio segment, superimpose and play the lip region video segment at the playback position of the selected segment in sequence to synthesize and play the target video corresponding to the part of the feature to be transmitted and the audio segment.

[0090] Among them, the order can be the sorting position of the feature to be transmitted received during the process of synthesizing and playing the target video, that is, which group of features to be transmitted is received. The lip region video segment is an animation segment of the lip region generated based on the lip decoder. The playback position is the video position when superimposing and playing the lip region video segment.

[0091] Specifically, for each feature to be transmitted, determine which group of features to be transmitted is received during the current synthesis process, that is, determine the order of the feature to be transmitted, determine the corresponding segment identifier in the video segment identification sequence according to this order, and then determine the action video segment corresponding to this segment from each pre-stored action video segment as the selected segment corresponding to the feature to be transmitted. And the lip position time sequence corresponding to the selected segment can be determined from the storage space as the lip position time sequence corresponding to the feature to be transmitted. If the presence result of the facial feature is present, there is no need to extract the facial feature, and the facial feature and the feature to be transmitted can be input into the lip decoder corresponding to the decoder parameters to obtain the lip region video segment. If the presence result of the facial feature is absent, it means that the feature to be transmitted contains all the features required to be used. Therefore, directly input the feature to be transmitted into the lip decoder corresponding to the decoder parameters to obtain the lip region video segment. According to the lip position time sequence corresponding to the feature to be transmitted, the playback position required for continuously playing the lip region video segment can be determined. When synchronously playing the selected segment and the audio segment, superimpose and play the lip region video segment at the playback position of the selected segment in sequence to synthesize and play a part of the target video on the client side, improving the real-time performance of the synthesis and playback of the target video. It can be understood that the part of the target video synthesized and played is the part of the target video corresponding to the currently received feature to be transmitted and the audio segment.

[0092] Exemplarily, the client receives a video segment identification sequence, and can retrieve the corresponding selected segments and lip position sequence, while waiting for the corresponding lip feature data stream (features to be transmitted) sent by the receiving server. After the lip feature data stream corresponding to this action video is received, the corresponding order is determined, and the order is matched to the selected segment x according to the order and the video segment identification sequence. The lip feature data stream (or the lip feature data stream or the pre-corresponding stored facial features) is upsampled by the lip decoder to generate a lip region video segment x', and then the lip region video segment x' is superimposed on the corresponding selected segment x according to the corresponding lip position sequence and played simultaneously, and the corresponding audio segment is played synchronously during playback.

[0093] It can be understood that in order to improve the coherence between the synthesized target videos of various parts, smoothing processing can be performed between the target videos of the parts that are continuously synthesized and played.

[0094] Based on the above example, after the lip-sync area video clip is played in time sequence at the playback position of the selected clip, if the to-be-transmitted feature and audio clip corresponding to the next selected clip have not been received, in order to ensure the continuity of the target video and avoid freezes, the pre-stored silent video clip can be used for processing, which can be:

[0095] If the features to be transmitted and the audio segment corresponding to the next selected segment are not received, at least one silent video segment is determined from the pre-stored silent video segments, and the at least one determined silent video segment is played in sequence until the features to be transmitted and the audio segment corresponding to the next selected segment are received.

[0096] Specifically, if the to-be-transmitted feature and audio segment corresponding to the next selected segment are not received, in order to avoid video freeze, the silent video segment can be played first while waiting for reception. That is, at least one silent video segment is determined from the pre-stored silent video segments, and the determined at least one silent video segment is played in sequence until the to-be-transmitted feature and audio segment corresponding to the next selected segment are received and the next partial target video is synthesized, then the silent video segment is stopped from being played, and the synthesized partial target video is continued to be played.

[0097] For example, after the client has played a selected segment, if the lip feature data stream (features to be transmitted) corresponding to the next selected segment has not been received due to bandwidth limitations or slow server task processing, then the silent video segment can be played in succession to avoid the situation where the digital human video is stuck and waiting.

[0098] On the premise of ensuring the video quality of the digital human, the above method effectively reduces the hardware resource requirements, saves the transmission bandwidth, and improves the fluency of video generation. It is proposed to directly stream the lip feature back to the client, and the client up-samples and decodes the lip feature and then displays it on the video layer. At the same time, the client monitors the generation progress of the lip feature data of the service port in real time. When it detects that the synthesis speed is slow, the client inserts a silent video segment and waits for the progress of the server to improve the fluency experience of the digital human display on the client side.

[0099] The present invention has the following technical effects: The client receives the video segment identification sequence sent by the server and sends the existence result of the facial feature to the server, so that the server determines the features to be transmitted corresponding to each selected segment according to the existence result, the selected segments corresponding to each segment identification in the video segment identification sequence, and the audio segments corresponding to each selected segment. Then the client sequentially receives the features to be transmitted and the audio segments of the selected segments corresponding to each segment identification in the video segment identification sequence sent by the server. According to the received features to be transmitted, audio segments, the decoder parameters of the pre-stored lip decoder, each action video segment, the lip position time sequence corresponding to each action video segment, each silent video segment, and the video segment identification sequence, the target video is synthesized and played, achieving the effect of lip synthesis on the server side, reducing the processing task volume on the server side, and improving the video processing efficiency.

[0100] Embodiment III

[0101] Figure 4 It is a schematic structural diagram of a data transmission and processing system for generating a digital human video provided by an embodiment of the present invention. As Figure 4 shown, the data transmission and processing system for generating a digital human video includes: a server 310 and a client 320; the server 310 executes the steps of the data transmission and processing method for generating a digital human video provided by any embodiment of the present invention; the client 320 executes the steps of the data transmission and processing method for generating a digital human video provided by any embodiment of the present invention.

[0102] The system of the above embodiment is used to implement the corresponding data transmission and processing method for generating a digital human video in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0103] Embodiment IV

[0104] On the basis of the above example, both the server and the client can be electronic devices. Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 5 shown, the electronic device 400 includes one or more processors 401 and a memory 402.

[0105] The processor 401 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 400 to perform desired functions.

[0106] The memory 402 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 401 can run the program instructions to implement the data transmission processing method for generating digital human videos in any embodiment of the present invention described above and / or other desired functions. Various contents such as initial extrinsic parameters, thresholds, etc. can also be stored in the computer-readable storage media.

[0107] In one example, the electronic device 400 can further include: an input device 403 and an output device 404, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 403 can include, for example, a keyboard, a mouse, etc. The output device 404 can output various information to the outside, including warning prompt information, braking force, etc. The output device 404 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0108] Of course, for simplicity, Figure 5 only some of the components related to the present invention in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 400 can further include any other appropriate components.

[0109] Embodiment Five

[0110] In addition to the above methods and devices, an embodiment of the present invention can also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps of the data transmission processing method for generating digital human videos provided in any embodiment of the present invention.

[0111] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0112] In addition, an embodiment of the present invention can also be a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps of the data transmission processing method for generating digital life videos provided by any embodiment of the present invention.

[0113] The computer-readable storage medium can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0114] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" do not specifically refer to the singular and may also include the plural. The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, or device comprising the element.

[0115] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A data transmission and processing method for a digital human generated video, characterized in that: Applied to the server, including: Determine a video segment identification sequence according to the audio duration of the target audio and each action video segment, and send the video segment identification sequence to the client; Determine the audio segment corresponding to each selected segment according to the target audio and the duration of the selected segment corresponding to each segment identifier in the video segment identifier sequence; Obtaining the existence result of the facial features of the client, and determining the to-be-transmitted features corresponding to each selected segment according to the existence result, each selected segment, and the audio segment corresponding to each selected segment; The features to be transmitted and the audio segments corresponding to the selected segments are sequentially sent to the client in the order of the video segment identification sequence, so that the client synthesizes and plays the target video according to the received features to be transmitted, the audio segments, the decoder parameters of the pre-stored lip decoder, the action video segments, the lip position timing sequence corresponding to the action video segments, the silent video segments and the video segment identification sequence; The action video segments and the silent video segments are determined based on the original video segmentation of the target digital human.

2. The method according to claim 1, characterized in that Before determining the video segment identification sequence according to the audio duration of the target audio and each action video segment, the method further includes: Determine at least one action video segment and at least one silent video segment according to the original video of the target digital human and the shortest duration of the video segment; For each action video clip, determining a lip position timing sequence and a clip identifier corresponding to the action video clip; Sending decoder parameters of a pre-trained lip decoder, each silent video segment, each action video segment, and the lip position timing sequence and segment identifier corresponding to each action video segment to the client; The identity encoder, the speech encoder and the lip-sync decoder constitute a lip-sync generation model, and the lip-sync generation model is trained based on each action video clip and the lip-sync position time sequence corresponding to each action video clip.

3. The method according to claim 2, characterized in that The step of determining at least one action video segment and at least one silent video segment according to the original video of the target digital human and the shortest duration of the video segments includes: Taking the starting video frame in the original video of the target digital human as the first video frame, and determining the position of the human in the first video frame; In the case where the time interval between the first video frame and the last video frame in the original video is greater than or equal to the shortest duration of the video segment, a video frame in the original video whose time interval with the first video frame is equal to the shortest duration of the video segment is used as a candidate video frame, and a position of a person in the candidate video frame is determined; In response to the position deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame being greater than or equal to a preset deviation, if the to-be-selected video frame is not the last video frame in the original video, taking the next video frame of the to-be-selected video frame as the to-be-selected video frame, and returning to the step of determining the position of the person in the to-be-selected video frame until the position deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame is less than the preset deviation; In response to the position deviation between the position of the person in the to-be-selected video frame and the position of the person in the first video frame being less than the preset deviation, the to-be-selected video frame is used as a second video frame corresponding to the first video frame, a video segment between the first video frame and the second video frame is used as a target video segment, and a speaking situation of the target digital person in the target video segment is determined; Using the next video frame of the second video frame as a new first video frame, and returning to the step of determining the position of the person in the first video frame; The target digital person's speaking situation is a target video clip of the target digital person speaking as each action video clip; If there is at least one target digital person whose speaking condition is that the target digital person does not speak, then the target video segments whose speaking conditions are that the target digital person does not speak are used as the silent video segments; If any target digital person speaks, then one is randomly selected from each action video segment as the video segment to be processed, the mouth of the target digital person in the video segment to be processed is processed as a closed state, and the processed video segment to be processed is used as a silent video segment.

4. The method according to claim 2, characterized in that: Also includes: For each action video clip, inputting the image frame sequence in the action video clip into a pre-trained identity encoder to obtain facial features corresponding to the action video clip; The facial features corresponding to each action video clip are sent to the client.

5. The method according to claim 1, characterized in that The determining, according to the existence result, the selected segments and the audio segments corresponding to the selected segments, the to-be-transmitted features corresponding to the selected segments includes: In response to the existence result being that there is no existence, for each selected segment, inputting the image frame sequence in the selected segment into a pre-trained identity encoder to obtain facial features corresponding to the selected segment, inputting the audio segment corresponding to the selected segment into a pre-trained speech encoder to obtain speech features corresponding to the selected segment, and using the facial features and the speech features as features to be transmitted corresponding to the selected segment; In response to the existence result being existence, for each selected segment, the audio segment corresponding to the selected segment is input into a pre-trained speech encoder to obtain the speech features corresponding to the selected segment, and the speech features are used as the features to be transmitted corresponding to the selected segment.

6. The method according to claim 1, characterized in that The step of determining the video segment identification sequence according to the audio duration of the target audio and each action video segment includes: Initialize a video segment identification sequence, and use the sum of the segment durations of the selected segments corresponding to the segment identifications in the video segment identification sequence as the existing duration; In the case where the existing duration is less than the audio duration of the target audio, randomly select one from the segment identifiers corresponding to the action video segments, add it to the video segment identifier sequence, and return to execute the step of taking the sum of the segment durations of the selected segments corresponding to the segment identifiers in the video segment identifier sequence as the existing duration, until the existing duration is greater than or equal to the audio duration of the target audio; In a case where the existing duration is greater than or equal to the audio duration of the target audio, the video segment identification sequence is obtained.

7. A data transmission and processing method for a digital human generated video, characterized in that: Applied to the client, including: Receive a video segment identification sequence sent by a server, and send the existence result of the facial features to the server, so that the server determines the to-be-transmitted features corresponding to each selected segment according to the existence result, the selected segment corresponding to each segment identification in the video segment identification sequence, and the audio segment corresponding to each selected segment; wherein the video segment identification sequence is determined by the server according to the audio duration of the target audio and each action video segment, and the audio segment corresponding to each selected segment is determined by the server according to the target audio and the duration of the selected segment corresponding to each segment identification in the video segment identification sequence; The device sequentially receives the features to be transmitted of the selected segments corresponding to each segment identifier in the video segment identifier sequence sent by the server and the audio segment, and synthesizes and plays the target video according to the received features to be transmitted, the audio segment, the decoder parameters of the pre-stored lip decoder, the action video segments, the lip position timing sequence corresponding to the action video segments, the silent video segments and the video segment identifier sequence.

8. The method according to claim 7, characterized in that The method of synthesizing and playing the target video according to the received features to be transmitted, the audio clips, the decoder parameters of the pre-stored lip decoder, the action video clips, the lip position timing sequence corresponding to the action video clips, the silent video clips and the video clip identification sequence comprises: For each feature to be transmitted, according to the order of the features to be transmitted and the video segment identification sequence, from each pre-stored action video segment, determine the selected segment corresponding to the feature to be transmitted and the lip position timing sequence corresponding to the feature to be transmitted; In response to the existence result of the facial feature being present, the facial feature and the feature to be transmitted are input into the lip decoder corresponding to the decoder parameter to obtain a lip area video clip; in response to the existence result of the facial feature being absent, the feature to be transmitted is input into the lip decoder corresponding to the decoder parameter to obtain a lip area video clip; The playback position is determined according to the lip position timing sequence corresponding to the feature to be transmitted, and when the selected segment and the audio segment are played synchronously, the lip area video segment is played in a time-sequenced overlay at the playback position of the selected segment to synthesize and play the target video of the portion corresponding to the feature to be transmitted and the audio segment.

9. The method according to claim 8, characterized in that After the lip-shaped area video segment is played in a time sequence at the playback position of the selected segment, the method further includes: If the features to be transmitted and the audio segment corresponding to the next selected segment are not received, at least one silent video segment is determined from the pre-stored silent video segments, and the at least one determined silent video segment is played in sequence until the features to be transmitted and the audio segment corresponding to the next selected segment are received.

10. A data transmission and processing system for a digital human-generated video, characterized in that: include: Server and client; The server performs the steps of the data transmission processing method for digital human generated video according to any one of claims 1 to 6; The client executes the steps of the data transmission processing method for digital human generated video as described in any one of claims 7 to 9.

Citation Information

Patent Citations

  • Digital human-driven rendering method and device, electronic equipment and storage medium

    CN116863039A

  • Method and system for generating digital human video based on video material

    CN117119123A

  • Method and device for increasing response speed of remote digital human

    CN117294905A

  • Broadcast video splitting method and device, storage medium and electronic equipment

    CN118264848A

  • Segmented rendering digital human video interaction method and device, terminal and storage medium

    CN118301413A