Digital human generation method, device and system, electronic equipment and storage medium
By extracting features on the server and sending them to the client in real time, combining the face mask feature sequence, collaborative calculation between the server and the client is realized, solving the problem of high cost in digital life generation, achieving fast and accurate digital life generation, and reducing the calculation amount and bandwidth consumption.
Patent Information
- Application Number
- CN202510212683.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing digital life generation solution, the cost of server-side deployment and video streaming traffic are high, and the calculation volume is large, which affects user experience and promotion applications, especially in scenarios where cost-sensitive or unstable network environments are.
By extracting feature on the server, publishing audio and text feature sequences in real time to the client, and combining the face mask feature sequences sent offline in advance, the client performs digital human generation, realizing collaborative calculations between the server and the client, reducing calculation amount and bandwidth consumption.
It realizes fast and accurate digital generation, reduces the cost of server-side deployment and video streaming traffic, and improves computing efficiency and user experience.
Smart Images

Figure CN120339469A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a digital human generation method, device, system, electronic device, and storage medium. Background Art
[0002] With the rapid development of technology, intelligent interaction technology has penetrated into all walks of life, greatly improving work efficiency and user experience. Among them, image generation and digital human technology, as one of the key technologies in the field of intelligent interaction, have shown great application potential in fields such as intelligent customer service and live streaming with goods due to their natural, convenient, and efficient characteristics.
[0003] Currently, the solutions for digital human generation generally follow a server - dominated process, that is, first perform data processing on the original data on the server, and then perform image generation on this basis; finally, encode the generated images and transmit them to the client through video streaming technology. The client is responsible for receiving and decoding these video streams to present complete video frames for users to watch.
[0004] However, this current solution has a large amount of calculation, the model structure for image generation is complex, and most of the calculations are performed on the server, which requires high computing resources on the server, thereby increasing the deployment cost. Moreover, transmitting video images in the form of video streaming not only consumes a large amount of network bandwidth resources, increases operating costs, but also may introduce delays during transmission, affecting user experience. These factors jointly limit the wide application of this technology, especially in application scenarios that are sensitive to costs or have unstable network environments. Summary of the Invention
[0005] The present invention provides a digital human generation method, device, system, electronic device, and storage medium to solve the problem of high push - flow cost and server deployment cost existing in the digital human generation solution in the prior art, realize the reasonable utilization of computing power, and effectively reduce the generation cost.
[0006] The present invention provides a digital human generation method, which is applied to the server and includes: Obtain target audio and / or target text for digital human generation in real - time; Extract features based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; Send the audio feature sequence and / or the text feature sequence to the client in real - time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, and obtains and displays a target face image sequence; The face mask feature sequence is determined by the server in advance based on the original face image sequence and sent offline to the client.
[0007] According to a digital human generation method provided by the present invention, the step of sending the audio feature sequence and / or the text feature sequence to the client in real time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, and obtains and displays a target face image sequence, includes: Send the audio feature sequence and / or the text feature sequence to the client in real time, so that the client performs feature fusion based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and generates a digital human based on the fused feature sequence obtained by the feature fusion, and obtains and displays a target face image sequence.
[0008] According to a digital human generation method provided by the present invention, an encoder in an encoding and decoding model for digital human generation is deployed on the server, and a decoder in the encoding and decoding model is deployed on the client; The encoder is used to extract features from the face mask image sequence to obtain a face mask feature sequence; the face mask image sequence is determined based on the original face image sequence; The decoder is used to generate a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and obtain a target face image sequence.
[0009] According to a digital human generation method provided by the present invention, the face mask feature sequence is determined based on the following steps: Perform face detection on each original face image in the original face image sequence to obtain the original face region in each original face image; Based on each original face region, perform image masking on each original face image in the original face image sequence to obtain a face mask image sequence; Extract features from the face mask image sequence to obtain a face mask feature sequence, and send the face mask feature sequence offline to the client, so that the client stores the face mask feature sequence and establishes an index; When generating a digital human, the client can obtain the face mask feature sequence through the index.
[0010] The present invention also provides a digital human generation method, which is applied to a client and includes: Receive the audio feature sequence and / or the text feature sequence sent by the server in real time; Generate a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and obtain and display a target face image sequence; The face mask feature sequence is obtained by the server through feature extraction based on the original face image sequence in advance, and is sent offline to the client; The audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction based on the target audio and / or target text obtained in real time.
[0011] According to a digital human generation method provided by the present invention, the generating a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and obtaining and displaying a target face image sequence includes: Perform feature fusion based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence to obtain a fused feature sequence; Generate a digital human based on the fused feature sequence, and obtain and display a target face image sequence.
[0012] According to a digital human generation method provided by the present invention, the generating a digital human based on the fused feature sequence, and obtaining and displaying a target face image sequence includes: Generate a target face region sequence based on the fused feature sequence; Perform image fusion based on the target face region sequence and the original face image sequence to obtain and display a target face image sequence.
[0013] The present invention also provides a digital human generation device, which is applied to a server and includes: An acquisition unit, configured to acquire in real time a target audio and / or a target text for digital human generation; An extraction unit, configured to perform feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; A sending unit, configured to send the audio feature sequence and / or the text feature sequence to the client in real time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and obtains and displays a target face image sequence; The face mask feature sequence is determined by the server in advance based on the original face image sequence, and is sent offline to the client.
[0014] The present invention also provides a digital human generation device, which is applied to a client and includes: A receiving unit, configured to receive in real time an audio feature sequence and / or a text feature sequence sent by a server; A generating unit, configured to generate a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; Wherein, the face mask feature sequence is obtained by the server through feature extraction on an original face image sequence in advance and sent to the client offline; The audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction on a target audio and / or a target text obtained in real time.
[0015] The present invention further provides a digital human generation system, including a server and a client, which are communicatively connected; The server is configured to obtain in real time a target audio and / or a target text for digital human generation, perform feature extraction on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence, and send the audio feature sequence and / or the text feature sequence to the client in real time; The client is configured to receive in real time the audio feature sequence and / or the text feature sequence sent by the server, and generate a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; The face mask feature sequence is determined by the server in advance based on an original face image sequence and sent to the client offline.
[0016] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, the digital human generation method described in any one of the above is implemented.
[0017] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the digital human generation method described in any one of the above is implemented.
[0018] The digital human generation method, device, system, electronic device and storage medium provided by the present invention are such that the server first extracts features from the target audio and / or target text obtained in real time to obtain an audio feature sequence and / or a text feature sequence, and then sends the audio feature sequence and / or the text feature sequence to the client in real time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence and a face mask feature sequence, and obtains and displays a target face image sequence; the face mask feature sequence is determined by the server in advance based on the original face image sequence and sent to the client offline. By effectively cutting and partitioning the overall calculation process, the collaborative digital human generation between the client and the server is realized, which can not only maximize the computing power of the client and the server, thereby solving the problems of high server deployment cost and high video streaming traffic cost, realizing fast and accurate digital human generation, but also reducing the calculation amount, reducing bandwidth consumption, and reducing the digital human generation cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 is one of the flow diagrams of the digital human generation method provided by the present invention; Figure 2 is a schematic diagram of the encoding and decoding model provided by the present invention; Figure 3 is another flow diagram of the digital human generation method provided by the present invention; Figure 4 is the overall flow chart of the digital human generation method provided by the present invention; Figure 5 is one of the structural diagrams of the digital human generation device provided by the present invention; Figure 6 is another structural diagram of the digital human generation device provided by the present invention; Figure 7 is the structural diagram of the digital human generation system provided by the present invention; Figure 8 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0022] As key technologies in the new generation of intelligent interaction, image generation and digital human technology have shown great application potential in fields such as intelligent customer service, live streaming with goods, and short video production due to their natural, convenient, efficient, and high-precision characteristics. They can not only significantly reduce labor costs but also effectively improve the efficiency and quality of content production.
[0023] Currently, the mainstream methods for digital human generation generally follow a server-dominated process: First, the original data is processed on the server to extract key feature sequences; subsequently, combining these feature sequences with face image sequences, a complex generation model is used for image generation; finally, the generated images are video-encoded and transmitted to the client through video streaming technology. The client is responsible for receiving and decoding these video streams.
[0024] However, this current solution has many defects. On the one hand, due to the large overall computational volume, complex generation model structure, and most of the computational processes being carried out on the server, there are high requirements for the server's computing resources, thus driving up the deployment cost. On the other hand, as a transmission means, video streaming not only consumes a large amount of network bandwidth resources, resulting in increased operating costs, but may also introduce delays during transmission, affecting the user experience. These factors jointly limit the large-scale promotion and application of this technology, especially in application scenarios that are sensitive to costs or have unstable network environments.
[0025] In response to this, the present invention provides a digital human generation method that coordinates the client and the server, aiming to effectively cut and partition the overall computational process, maximize the computing power of the client and the server, thereby solving the problems of high server deployment costs and high video streaming traffic costs, achieving fast and accurate digital human generation, while reducing the computational volume, bandwidth consumption, and digital human generation costs.
[0026] Figure 1 is one of the flow diagrams of the digital human generation method provided by the present invention. As Figure 1 shown, the method includes: Step 110, obtaining the target audio and / or target text for digital human generation in real time; Step 120, performing feature extraction based on the target audio and / or target text to obtain an audio feature sequence and / or a text feature sequence; Step 130: Real-time send the audio feature sequence and / or the text feature sequence to the client, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and obtains and displays a target face image sequence; Wherein, the face mask feature sequence is pre-determined by the server based on the original face image sequence and sent to the client offline.
[0027] Specifically, in the current digital human generation solution, due to the large amount of calculation and the complex structure of the generation model, most of the calculations are performed on the server side, resulting in a high deployment cost of the server side; moreover, after the image generation is completed, the video screen is transmitted through the push stream method, and the client plays the video through pulling and decoding. The process of video push stream consumes a large amount of traffic, resulting in an increase in traffic bandwidth cost, making it difficult to carry out large-scale promotion and use.
[0028] In view of this, in the embodiments of the present invention, it is proposed that the tasks can be divided based on the computing power of the server side and the client side to effectively cut the overall calculation process, so as to maximize the utilization of the computing power of each side, and further reduce the deployment cost of the server side and the video push stream cost, fundamentally solve the problem of high traffic bandwidth cost and the synthesis response problem brought by video push stream, and realize the reduction of the overall calculation amount and calculation complexity, as well as the improvement of efficiency. Specifically, it can be to first process the text and / or audio for digital human generation on the server side to generate corresponding feature sequences; then, based on the face mask feature sequence, and the text feature sequence and / or the audio feature sequence, use the computing power of the client side to generate a digital human to obtain and display a target face image sequence.
[0029] It can be understood that in the actual application process, before digital human generation, it is usually necessary to first determine the data for digital human generation and driving. This data can be text data, audio data, or text data and audio data. The embodiments of the present invention do not make specific limitations on this. That is, when the user needs to perform digital human generation, the user can input the text data and / or audio data for driving the digital human generation, and the server side can obtain this data in real time, that is, the target text and / or target audio for digital human generation can be obtained in real time.
[0030] Immediately afterwards, the obtained target audio and / or target text can be processed in real time to obtain their corresponding feature sequences, that is, the audio feature sequence and / or the text feature sequence. Here, specifically, in the server side, the target audio and / or target text obtained in real time can be feature-extracted to extract the information helpful for the generation of the face image of the digital human and encoded into features, so as to obtain the audio feature sequence and / or the text feature sequence.
[0031] It should be noted that the process of feature extraction to obtain the corresponding feature sequence may be to first perform frame segmentation on the target audio and / or target text to obtain multiple audio frames and / or text segments, and then perform feature extraction on them to obtain the corresponding audio feature sequence and / or text feature sequence; or it may be to first perform feature extraction on the target audio and / or target text to extract global features, obtain the corresponding audio features and / or text features, and then perform feature segmentation on them to obtain the corresponding audio feature sequence and / or text feature sequence. The embodiments of the present invention do not make specific limitations on this.
[0032] However, considering the scheme of first feature extraction and then segmentation, since the segmentation is performed on the global features, the obtained feature sequence retains more global information but loses local details. While the scheme of first frame segmentation and then feature extraction focuses more on local features because the feature extraction is performed on each frame separately, and can more finely reflect the local changes of the audio and / or text. Specifically, in the digital human generation process driven by the target audio and / or target text in the embodiments of the present invention, the local details in the audio and / or text are emphasized. Therefore, preferably, in the embodiments of the present invention, the target audio and / or target text may be first frame segmented, and then the frame-segmented audio and / or text may be feature extracted to obtain the corresponding audio feature sequence and / or text feature sequence.
[0033] In the embodiments of the present invention, performing feature extraction on the target audio and / or target text in the server can better capture the local features of the audio, short-time changes in the signal, as well as the local semantics and detailed information of the text, so that the extracted audio feature sequence and / or text feature sequence can more comprehensively and finely reflect the phoneme changes and / or context information in the input information, etc., avoid problems such as discontinuous context and loss of audio details, and can more finely capture the local features of the audio and / or text, thereby improving the accuracy and efficiency of processing.
[0034] For example, when generating lip movements according to the audio so that the lips can present movements corresponding to the audio content, for the input audio, first frame segmentation and then feature extraction can better capture the local features in the audio signal, such as pitch, timbre, speech rate, etc. Since the audio signal usually changes significantly within a short time (such as within a syllable or a word), through frame segmentation, the audio signal is divided into multiple short segments, and the features within each segment are relatively stable, and the obtained audio feature sequence is more refined, thereby improving the matching accuracy between the lip movements and the audio content. In addition, first frame segmentation can ensure that each lip movement matches the corresponding audio frame features, thus achieving a more accurate synchronization effect.
[0035] After that, the server can immediately send the extracted audio feature sequence and / or text feature sequence to the client in real time to perform the subsequent digital human generation process on the client side, thereby achieving an effective segmentation of the task. As a result, what the client receives is no longer a generated image sequence, nor does it need to perform video encoding and streaming, thus greatly reducing the traffic cost.
[0036] Specifically, when the audio feature sequence and / or text feature sequence are obtained, the server can directly send this audio feature sequence and / or text feature sequence to the client in real time, requesting the client to perform digital human generation based on this to obtain the final face image sequence and display it.
[0037] Correspondingly, after the server sends it in real time, the client can immediately receive this audio feature sequence and / or text feature sequence. Then, based on this audio feature sequence and / or text feature sequence, and the face mask feature sequence received in advance, digital human generation can be performed to obtain the final face image sequence, that is, the target face image sequence, and display it to present a complete video for the user to watch. Here, specifically, based on the face mask feature sequence that characterizes the digital human's image and body posture, digital human generation can be driven by the audio feature sequence and / or text feature sequence to generate the face area that is masked / missing in the image characterized by the face mask feature sequence, and make the final face image of the digital human match both the target audio and / or target text, and also match the digital human's image and body posture, without being jerky or rigid, etc., thereby achieving fast and accurate digital human generation and obtaining a vivid and natural target face image sequence for the user to watch.
[0038] Among them, the face mask feature sequence is determined by the server in advance based on the original face image sequence and sent to the client. However, it should be noted that different from the real-time reception and processing of the target audio and / or target text, the processing of the original face image sequence here is an offline process. In the server, through processing at different times (offline and online processing), the processing pressure on the server can be further reduced on the basis of task segmentation, avoiding the problem of tight computing power resources and excessive processing pressure caused by processing in the same period, reducing the requirement for computing power, and reducing the deployment cost of the server.
[0039] Here, specifically, the server can pre-obtain / receive the original face image sequence in advance, then determine the face regions in each original face image in this sequence, mask this region, and then perform feature extraction on the masked image to extract information characterizing the figure, posture, etc. of the person, and encode it into features, so as to obtain the face mask features corresponding to the original face image, and thus a face mask feature sequence can be constructed. Among them, the original face image can be obtained by shooting with a network camera, or obtained by network search and download, or separated from a video. The embodiments of the present invention do not make specific limitations on this.
[0040] Compared with the traditional solution where calculations are performed on the server side to obtain the generated image and perform video encoding and streaming, which has the problems of high server deployment cost and high push-stream traffic bandwidth cost, and with the development of science and technology and the improvement of the computing power of mobile devices, considering the solution of migrating all calculations on the server side to the client side, which brings the problem of a significant increase in device power consumption. In the embodiments of the present invention, through effective task splitting, effective operation coordination between the two ends and the maximum utilization of computing power are achieved. Not only is the server deployment cost reduced, but also no significant power consumption pressure is brought to the client side. While reducing the computing power requirements, the amount of calculation, computational complexity and cost are reduced.
[0041] For the digital human generation method provided by the present invention, the server first performs feature extraction on the real-time acquired target audio and / or target text to obtain an audio feature sequence and / or a text feature sequence, and then sends the audio feature sequence and / or the text feature sequence to the client in real time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, obtains a target face image sequence and displays it; the face mask feature sequence is pre-determined by the server based on the original face image sequence and sent to the client offline. Through the effective cutting and chunking of the overall calculation process, digital human generation with client-server collaboration is realized, which can not only maximize the computing power of the client and the server, thus solving the problems of high server deployment cost and high video push-stream traffic cost, and realizing fast and accurate digital human generation, but also reduce the amount of calculation, reduce bandwidth consumption, and reduce the cost of digital human generation.
[0042] Based on the above embodiments, step 130 includes: Send the audio feature sequence and / or the text feature sequence to the client in real time, so that the client performs feature fusion based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and performs digital human generation based on the fused feature sequence obtained by the feature fusion, obtains a target face image sequence and displays it.
[0043] Specifically, the process of real-time sending the audio feature sequence and / or text feature sequence to the client so that the client can generate a digital human accordingly is actually that the server first sends the audio feature sequence and / or text feature sequence to the client in real time. Then, upon receiving this audio feature sequence and / or text feature sequence, the client calls the pre-received face mask feature sequence to generate a digital human, obtains the target face image sequence and displays it to present a complete video frame for the user to watch.
[0044] In detail, in the client, information fusion is first performed at the feature level, that is, face feature embedding is performed based on the real-time received feature sequence to achieve information fusion and obtain the fused feature sequence. Specifically, the audio feature sequence and / or text feature sequence, as well as the face mask feature sequence, are subjected to feature fusion to obtain the fused feature sequence.
[0045] Immediately afterwards, digital human generation can be performed according to this fused feature sequence to obtain the target face image sequence. Here, specifically, in the client, this fused feature sequence is decoded to restore it to an image, thereby obtaining the target face image sequence, and this target face image sequence can be displayed for the user to watch.
[0046] Based on the above embodiments, an encoder in the encoding and decoding model for digital human generation is deployed on the server, and a decoder in the encoding and decoding model is deployed on the client; The encoder is used to extract features from the face mask image sequence to obtain the face mask feature sequence; the face mask image sequence is determined based on the original face image sequence; The decoder is used to perform digital human generation based on the audio feature sequence and / or text feature sequence, as well as the face mask feature sequence to obtain the target face image sequence.
[0047] Considering that in the current digital human generation solution, the model network structure for image generation is complex, which exacerbates the computing power problem and cost problem on the server side. In the embodiments of the present invention, it is proposed that the model structure can be improved to achieve two-end deployment, thereby reducing the computing power requirements of the server, alleviating the cost pressure, reducing the amount of calculation and computational complexity, and improving the generation efficiency.
[0048] In detail, in the embodiments of the present invention, the model for digital human generation is split, so that the encoding part for feature extraction is deployed on the server, and the decoding part for generating the face image of the digital human is deployed on the client. That is, an encoder in the encoding and decoding model for digital human generation is deployed on the server, and a decoder in the encoding and decoding model is deployed on the client; Figure 2 is a schematic diagram of the encoding and decoding model provided by the present invention, as Figure 2As shown in the figure, the encoder deployed on the server can encode each face mask image in the face mask image sequence to obtain a face mask feature sequence. Among them, the face mask image sequence is determined by the server in advance based on the original face image sequence through face detection and image masking. The decoder can then decode the fused feature sequence obtained by fusing the audio feature sequence and / or the text feature sequence, and the face mask feature sequence to realize digital human generation, thereby obtaining the target face image sequence.
[0049] It should be noted that the encoding and decoding model here can be pre-trained through training data. Among them, the encoder and the decoder can be trained together or separately. The training method can be supervised, unsupervised, or other methods. The embodiments of the present invention do not make specific limitations on this, as long as it can ensure that the model finally has good feature extraction ability and accurate digital human generation ability.
[0050] In the embodiments of the present invention, by improving the model structure, the model originally deployed on the server can be deployed in two parts to perform tasks separately, so that the encoder in the server only encodes the original face images, and the decoder in the client only performs feature decoding, thereby greatly reducing the number of model parameters, effectively reducing the complexity of the model, and reducing the deployment cost.
[0051] Based on the above embodiments, the face mask feature sequence is determined based on the following steps: Perform face detection on each original face image in the original face image sequence to obtain the original face regions in each original face image; Based on each original face region, perform image masking on each original face image in the original face image sequence to obtain a face mask image sequence; Extract features from the face mask image sequence to obtain a face mask feature sequence, and offline distribute the face mask feature sequence to the client so that the client stores the face mask feature sequence and establishes an index; The client can obtain the face mask feature sequence through the index when performing digital human generation.
[0052] Specifically, the process of the server processing the original face image sequence to obtain the face mask feature sequence is an offline processing process, and this process specifically includes: The server first obtains the original face image sequence. Here, the original face images can be obtained by shooting with a network camera, or by network search, download, or separated from a video. The embodiments of the present invention do not make specific limitations on this.
[0053] Subsequently, face detection can be performed on each original face image in the original face image sequence to detect the location of the face and frame it with a detection box, thereby obtaining the original face region in each original face image.
[0054] After that, based on each original face region, image masking can be performed on each original face image in the original face image sequence to mask the original face region in each original face image, thereby obtaining the remaining original face images, that is, face mask images. Then, feature extraction can be performed on each face mask image in the face mask image sequence to obtain a face mask feature sequence. Since the original face image sequence is usually preset and has a limited sequence length, the sequence length of the face mask image sequence determined based on it is also limited. In the embodiments of the present invention, performing feature extraction on the limited face mask image sequence only requires one inference, which is simpler and more efficient. Then, the extracted face mask feature sequence can be sent offline to the client for the client to store and establish an index, so that the face mask feature sequence can be quickly obtained according to the index during subsequent digital human generation, and feature embedding and digital human generation can be performed, realizing the efficient cooperation between the server and the client.
[0055] The present invention also provides a digital human generation device. Figure 3 It is the second flow schematic diagram of the digital human generation method provided by the present invention. As Figure 3 shown, this method is applied to the client, and this method includes: Step 310, receiving the audio feature sequence and / or text feature sequence sent by the server in real time; Step 320, generating a digital human based on the audio feature sequence and / or text feature sequence, and the face mask feature sequence to obtain a target face image sequence and display it; Among them, the face mask feature sequence is obtained by the server based on feature extraction of the original face image sequence in advance and sent offline to the client; The audio feature sequence and / or text feature sequence is obtained by the server based on the target audio and / or target text obtained in real time through feature extraction.
[0056] Specifically, considering that in the current digital human generation solution, due to the large amount of calculation and the complex structure of the generation model, most of the calculations are performed on the server, resulting in a high deployment cost of the server; moreover, after the image generation is completed, the video screen is transmitted through push streaming, and the client plays the video through pull streaming and decoding. The process of video push streaming consumes a large amount of traffic, resulting in an increase in traffic bandwidth cost and making it difficult to be widely promoted and used on a large scale.
[0057] In view of this, in the embodiments of the present invention, it is proposed that task chunking can be performed based on the computing power of the server and the client to effectively cut the overall computing process, thereby maximizing the utilization of the computing power of each end, and further reducing the deployment cost of the server and the video streaming cost, fundamentally solving the problem of high traffic bandwidth cost and the synthesis response problem brought by video streaming, achieving a decrease in the overall computing volume and computing complexity, and an improvement in efficiency.
[0058] It can be understood that in the actual application process, before digital human generation, it is usually necessary to first determine the data for digital human generation and driving. This data can be text data, audio data, or both text data and audio data. The embodiments of the present invention do not make specific limitations on this. That is, when a user needs to perform digital human generation, they can input the text data and / or audio data for driving the digital human generation, and the server can obtain this data in real time, that is, the target text and / or target audio for digital human generation can be obtained in real time.
[0059] Immediately afterwards, the obtained target audio and / or target text can be processed in real time to obtain their corresponding feature sequences, that is, audio feature sequences and / or text feature sequences. Here, specifically, in the server, the target audio and / or target text obtained in real time can be feature-extracted to extract the information helpful for the generation of the face image of the digital human and encoded into features, thereby obtaining audio feature sequences and / or text feature sequences.
[0060] It should be noted that the process of feature extraction to obtain the corresponding feature sequences here can be to first perform frame segmentation on the target audio and / or target text to obtain multiple audio frames and / or text segments, and then perform feature extraction on them to obtain the corresponding audio feature sequences and / or text feature sequences; it can also be to first perform feature extraction on the target audio and / or target text to extract global features to obtain the corresponding audio features and / or text features, and then perform feature segmentation on them to obtain the corresponding audio feature sequences and / or text feature sequences. The embodiments of the present invention do not make specific limitations on this.
[0061] However, in the solution of first performing feature extraction and then segmentation, since the segmentation is performed on the global features, the obtained feature sequence retains more global information but loses local details. On the other hand, in the solution of first diving the frames and then performing feature extraction, since the feature extraction is performed separately for each frame, the obtained feature sequence focuses more on local features and can more finely reflect the local changes of the audio and / or text. Specifically, in the digital human generation process driven by the target audio and / or target text in the embodiments of the present invention, the local details in the audio and / or text are emphasized. Therefore, preferably, in the embodiments of the present invention, the target audio and / or target text can be first frame-divided, and then the feature extraction is performed on the frame-divided audio and / or text, so as to obtain the corresponding audio feature sequence and / or text feature sequence.
[0062] In the embodiments of the present invention, performing feature extraction on the target audio and / or target text in the server can better capture the local features of the audio, the short-time changes in the signal, as well as the local semantics and detailed information of the text, so that the extracted audio feature sequence and / or text feature sequence can more comprehensively and finely reflect the phoneme changes and / or context information, etc. in the input information, avoid the problems of incoherent context and lost audio details, and can more finely capture the local features of the audio and / or text, thereby improving the accuracy and efficiency of processing.
[0063] For example, when generating lip movements according to the audio so that the lips can present movements corresponding to the audio content, for the input audio, first dividing the frames and then performing feature extraction can better capture the local features in the audio signal, such as pitch, timbre, speech rate, etc. Since the audio signal usually changes significantly within a short period of time (such as within a syllable or a word), through frame division, the audio signal is divided into multiple short segments, and the features within each segment are relatively stable, and the obtained audio feature sequence is more refined, thereby improving the matching accuracy between the lip movements and the audio content. In addition, first performing frame division can ensure that each lip movement matches the corresponding audio frame features, thereby achieving a more accurate synchronization effect.
[0064] After that, the server can directly send the extracted audio feature sequence and / or text feature sequence to the client in real time, so as to perform the subsequent digital human generation process on the client side, thereby realizing the effective segmentation of the task, so that what the client obtains is no longer the generated image sequence, nor does it need to perform video encoding and streaming, thereby greatly reducing the traffic cost.
[0065] Specifically, here, when obtaining the audio feature sequence and / or text feature sequence, the server can directly send this audio feature sequence and / or text feature sequence to the client in real time, so as to request the client to perform digital human generation based on this, obtain the final face image sequence and display it.
[0066] Correspondingly, after the server-side issues it in real time, the client can receive this audio feature sequence and / or text feature sequence in real time. Then, based on this audio feature sequence and / or text feature sequence, and the face mask feature sequence received in advance, digital human generation can be performed to obtain the final face image sequence, that is, the target face image sequence, and this is displayed to present a complete video for the user to watch. Here, specifically, based on the face mask feature sequence representing the digital human's image and posture, digital human generation can be driven by the audio feature sequence and / or text feature sequence to generate the face area masked / missing in the image represented by the face mask feature sequence, and make the face image of the finally obtained digital human match both the target audio and / or target text, and also match the digital human's image and posture, without being jerky or rigid, etc., so as to achieve fast and accurate digital human generation and obtain a vivid and natural target face image sequence for the user to watch.
[0067] Among them, the face mask feature sequence is obtained by the server-side through feature extraction from the original face image sequence and sent to the client. However, it should be noted that different from the real-time reception and processing of the target audio and / or target text, the processing of the original face image sequence here is an offline process. In the server-side, through processing at different time periods (offline and online processing), the processing pressure on the server-side can be further reduced on the basis of task segmentation, avoiding the problem of tight computing power resources and large processing pressure caused by processing in the same time period, reducing the requirement for computing power, and reducing the deployment cost of the server-side.
[0068] Here, specifically, the server-side can pre-obtain / receive the original face image sequence in advance, then determine the face area in each original face image in this sequence, mask this area, and then perform feature extraction on the masked image to extract information representing the human image, posture, etc., and encode it into features, so as to obtain the face mask feature corresponding to the original face image, and thus the face mask feature sequence can be constructed. Among them, the original face image can be obtained by shooting with a network camera, or obtained by network search and download, or separated from a video. The embodiments of the present invention do not make specific limitations on this.
[0069] Compared with the traditional solution where calculations are performed on the server side to obtain the generated image and perform video encoding and streaming, which has problems such as high server deployment costs and high push-stream traffic bandwidth costs, and with the development of science and technology and the improvement of the computing power of mobile devices, considering the solution of migrating all calculations on the server side to the client side, which brings the problem of a significant increase in device power consumption. In the embodiments of the present invention, through effective task splitting, effective job collaboration between the two ends is achieved, as well as the maximization of the utilization of computing power. This not only reduces the server deployment cost but also does not bring a large power consumption pressure to the client side. While reducing the computing power requirements, it realizes a decrease in the amount of calculation, computational complexity, and cost.
[0070] The digital human generation method provided by the present invention is such that the server first extracts features from the real-time acquired target audio and / or target text, and then real-time sends the obtained audio feature sequence and / or text feature sequence to the client. The client can then receive the audio feature sequence and / or text feature sequence in real time, and based on the audio feature sequence and / or text feature sequence, as well as the face mask feature sequence, perform digital human generation to obtain the target face image sequence and display it; the face mask feature sequence is obtained by the server pre-extracting features from the original face image sequence and offline sending it to the client. Through the effective cutting and chunking of the overall calculation process, digital human generation with collaboration between the client and the server is achieved, which can not only maximize the computing power of the client and the server, thus solving the problems of high server deployment costs and high video push-stream traffic costs and realizing fast and accurate digital human generation, but also reduce the amount of calculation, bandwidth consumption, and digital human generation cost.
[0071] Based on the above embodiments, step 320 includes: Perform feature fusion based on the audio feature sequence and / or text feature sequence, as well as the face mask feature sequence, to obtain a fused feature sequence; Perform digital human generation based on the fused feature sequence to obtain the target face image sequence and display it.
[0072] Specifically, the process of performing digital human generation according to the audio feature sequence and / or text feature sequence, as well as the face mask feature sequence, is actually that the server first real-time sends the audio feature sequence and / or text feature sequence to the client. After receiving the audio feature sequence and / or text feature sequence real-time sent by the server, the client calls the pre-received face mask feature sequence to perform digital human generation, obtain the target face image sequence, and display it to present a complete video frame for the user to watch.
[0073] Specifically, in the client, information fusion is first performed at the feature level, that is, face feature embedding is performed based on the received feature sequence in real time, so as to achieve information fusion and obtain a fused feature sequence. Specifically, feature fusion is performed on the audio feature sequence and / or text feature sequence, as well as the face mask feature sequence, so as to obtain a fused feature sequence.
[0074] Immediately afterwards, digital human generation can be performed according to this fused feature sequence, so as to obtain a target face image sequence. Here, specifically, in the client, this fused feature sequence can be decoded to restore it to an image, so as to obtain a target face image sequence, and this target face image sequence can be displayed for the user to view.
[0075] Based on the above embodiments, performing digital human generation based on the fused feature sequence, obtaining and displaying the target face image sequence, includes: Performing digital human generation based on the fused feature sequence to obtain a target face region sequence; Based on the target face region sequence and the original face image sequence, image fusion is performed to obtain and display the target face image sequence.
[0076] Specifically, the above process of performing digital human generation according to the fused feature sequence, obtaining and displaying the target face image sequence can specifically include: The client first decodes the fused feature sequence to generate the face image of the digital human, so as to obtain the generated face image sequence. Since the generated face image sequence only contains the face region at this time and lacks other appearance information, such as hairstyle, clothes, etc., it can be called the target face region sequence.
[0077] After obtaining this target face region sequence, it is also necessary to combine the original face image sequence to determine the final target face image sequence. That is, each target face region in the generated target face region sequence can be merged with the corresponding original face image in the original face image sequence, and the final image, that is, the target face image sequence, can be obtained through image fusion, and this target face image sequence can be displayed, or this target face image sequence, as well as the target audio and / or target text, can be jointly rendered and displayed to achieve the effect of video playback.
[0078] Figure 4 is the overall flowchart of the digital human generation method provided by the present invention, as Figure 4 shown, this method specifically includes two links: offline processing and online real-time processing, which are respectively: Offline processing: First, perform face detection on each original face image in the original face image sequence on the server side to determine the original face area in each original face image according to the detection frame. Then, an image mask can be performed on the original face area in each original face image to obtain a face mask image sequence, and feature extraction can be performed on this face mask image sequence to obtain a face mask feature sequence. This face mask feature sequence is sent offline to the client side so that the client side can store and index this face mask feature sequence, so that the face mask feature sequence can be quickly obtained according to the index during the subsequent digital human generation process.
[0079] Online real-time processing: The server side first performs feature extraction on the target audio and / or target text for digital human generation, and sends the extracted audio feature sequence and / or text feature sequence to the client side in real time; the client side can then receive this audio feature sequence and / or text feature sequence in real time, and can quickly obtain the pre-stored face mask feature sequence according to the index, so that digital human generation can be performed according to this audio feature sequence and / or text feature sequence, and the face mask feature sequence, to obtain a target face area sequence, and perform image fusion on it with the original face image sequence to obtain a target face image sequence, and this target face image sequence can be displayed.
[0080] Next, the digital human generation device provided by the present invention will be described. The digital human generation device described below can be correspondingly referred to with the digital human generation method described above.
[0081] Figure 5 is one of the structural schematic diagrams of the digital human generation device provided by the present invention. As Figure 5 shown, this device is applied to the server side, and this device includes: An acquisition unit 510, configured to acquire the target audio and / or target text for digital human generation in real time; An extraction unit 520, configured to perform feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; A sending unit 530, configured to send the audio feature sequence and / or the text feature sequence to the client side in real time, so that the client side performs digital human generation based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, to obtain a target face image sequence and display it; The face mask feature sequence is determined by the server side in advance based on the original face image sequence and sent offline to the client side.
[0082] The digital human generation device provided by the present invention enables the server to first extract features from the acquired target audio and / or target text in real time to obtain an audio feature sequence and / or a text feature sequence, and then send the audio feature sequence and / or the text feature sequence to the client in real time, so that the client can generate a digital human based on the audio feature sequence and / or the text feature sequence and a face mask feature sequence, and obtain and display a target face image sequence; the face mask feature sequence is pre-determined by the server based on an original face image sequence and sent to the client offline. Through effective cutting and chunking of the overall calculation process, collaborative digital human generation between the client and the server is realized, which can not only maximize the computing power of the client and the server, thereby solving the problems of high server deployment costs and high video streaming traffic costs, achieving fast and accurate digital human generation, but also reducing the amount of calculation, reducing bandwidth consumption, and reducing the cost of digital human generation.
[0083] Based on the above embodiment, the sending unit 530 is configured to: Send the audio feature sequence and / or the text feature sequence to the client in real time, so that the client can perform feature fusion based on the audio feature sequence and / or the text feature sequence and the face mask feature sequence, and perform digital human generation based on the fusion feature sequence obtained from the feature fusion to obtain and display a target face image sequence.
[0084] Based on the above embodiment, an encoder in an encoding and decoding model for digital human generation is deployed on the server, and a decoder in the encoding and decoding model is deployed on the client; The encoder is configured to extract features from a face mask image sequence to obtain a face mask feature sequence; the face mask image sequence is determined based on the original face image sequence; The decoder is configured to perform digital human generation based on the audio feature sequence and / or the text feature sequence and the face mask feature sequence to obtain a target face image sequence.
[0085] Based on the above embodiment, the device further includes a face processing unit, configured to: Perform face detection on each original face image in the original face image sequence to obtain the original face region in each original face image; Based on each original face region, perform image masking on each original face image in the original face image sequence to obtain a face mask image sequence; Extract features from the face mask image sequence to obtain a face mask feature sequence, and send the face mask feature sequence to the client offline, so that the client can store the face mask feature sequence and establish an index; When generating a digital human, the client can obtain the face mask feature sequence through indexing.
[0086] Figure 6 It is the second structural schematic diagram of the digital human generation device provided by the present invention. As Figure 6 shown, this device is applied to the client and includes: A receiving unit 610, configured to receive the audio feature sequence and / or text feature sequence sent by the server in real time; A generating unit 620, configured to generate a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, obtain a target face image sequence and display it; Wherein, the face mask feature sequence is obtained by the server through feature extraction based on the original face image sequence in advance and sent to the client offline; The audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction based on the target audio and / or target text obtained in real time.
[0087] For the digital human generation device provided by the present invention, the server first performs feature extraction on the target audio and / or target text obtained in real time, and sends the obtained audio feature sequence and / or text feature sequence to the client in real time. The client can then receive the audio feature sequence and / or text feature sequence in real time, and generate a digital human based on the audio feature sequence and / or text feature sequence, and the face mask feature sequence, obtain a target face image sequence and display it; the face mask feature sequence is obtained by the server through feature extraction based on the original face image sequence in advance and sent to the client offline. Through the effective cutting and partitioning of the overall calculation process, the collaborative digital human generation between the client and the server is realized, which can not only maximize the computing power of the client and the server, thus solving the problems of high server deployment cost and high video streaming traffic cost, realizing fast and accurate digital human generation, but also reducing the amount of calculation, reducing bandwidth consumption, and reducing the cost of digital human generation.
[0088] Based on the above embodiment, the generating unit 620 is used for: Perform feature fusion based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence to obtain a fusion feature sequence; Generate a digital human based on the fusion feature sequence, obtain a target face image sequence and display it.
[0089] Based on the above embodiment, the generating unit 620 is used for: Generate a digital human based on the fusion feature sequence to obtain a target face region sequence; Perform image fusion based on the target face region sequence and the original face image sequence to obtain and display a target face image sequence.
[0090] Figure 7 is a schematic structural diagram of the digital human generation system provided by the present invention. As Figure 7 shown, the system includes a server 710 and a client 720, and the server 710 and the client 720 are communicatively connected; The server 710 is configured to obtain in real time a target audio and / or a target text for digital human generation, perform feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence, and transmit the audio feature sequence and / or the text feature sequence to the client 720 in real time; The client 720 is configured to receive in real time the audio feature sequence and / or the text feature sequence transmitted by the server 710, and perform digital human generation based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence to obtain and display a target face image sequence; The face mask feature sequence is determined by the server 710 in advance based on the original face image sequence and is offline transmitted to the client 720.
[0091] The digital human generation system provided by the present invention includes a server and a client, and the server and the client are communicatively connected; the server is configured to perform feature extraction based on the target audio and / or the target text obtained in real time to obtain an audio feature sequence and / or a text feature sequence, and transmit the audio feature sequence and / or the text feature sequence to the client in real time, and the client is configured to receive the audio feature sequence and / or the text feature sequence in real time, and perform digital human generation based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence to obtain and display a target face image sequence; the face mask feature sequence is determined by the server in advance based on the original face image sequence and is offline transmitted to the client. By effectively cutting and partitioning the overall calculation process, collaborative digital human generation between the client and the server is realized, which can not only maximize the computing power of the client and the server, thereby solving the problems of high server deployment cost and high video streaming traffic cost, realizing fast and accurate digital human generation, but also reducing the amount of calculation, reducing bandwidth consumption, and reducing the cost of digital human generation.
[0092] Figure 8 Illustrates a schematic physical structure diagram of an electronic device. As Figure 8As shown in the figure, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute the digital human generation method. When this method is applied to the server, it includes: obtaining in real time the target audio and / or target text for digital human generation; performing feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; and sending the audio feature sequence and / or the text feature sequence to the client in real time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; the face mask feature sequence is pre-determined by the server based on the original face image sequence and sent to the client offline. When this method is applied to the client, it includes: receiving in real time the audio feature sequence and / or the text feature sequence sent by the server; generating a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; where the face mask feature sequence is obtained by the server through feature extraction based on the original face image sequence and sent to the client offline; the audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction based on the target audio and / or target text obtained in real time.
[0093] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0094] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the digital human generation method provided by each of the above methods. When this method is applied to a server, it includes: obtaining in real time a target audio and / or a target text for digital human generation; performing feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; and sending the audio feature sequence and / or the text feature sequence to a client in real time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; the face mask feature sequence is determined in advance by the server based on an original face image sequence and sent offline to the client. When this method is applied to a client, it includes: receiving in real time the audio feature sequence and / or the text feature sequence sent by the server; generating a digital human based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, to obtain a target face image sequence and display it; wherein, the face mask feature sequence is obtained by the server through feature extraction based on an original face image sequence in advance and sent offline to the client; the audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction based on the target audio and / or the target text obtained in real time.
[0095] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, having a computer program stored thereon, which when executed by a processor, implements the digital human generation method provided by each of the above methods. When this method is applied to the server, it includes: obtaining in real time the target audio and / or target text for digital human generation; performing feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; and sending the audio feature sequence and / or the text feature sequence to the client in real time, so that the client generates a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; the face mask feature sequence is pre-determined by the server based on the original face image sequence and sent to the client offline. When this method is applied to the client, it includes: receiving in real time the audio feature sequence and / or the text feature sequence sent by the server; generating a digital human based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; wherein, the face mask feature sequence is obtained by the server through feature extraction based on the original face image sequence and sent to the client offline; the audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction based on the target audio and / or target text obtained in real time.
[0096] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating digital life, characterized in that, Applied to the server side, including: Obtaining the target audio and / or target text for digital human generation in real time; Performing feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; Real-time sending the audio feature sequence and / or the text feature sequence to the client, so that the client performs digital human generation based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; The face mask feature sequence is pre-determined by the server based on the original face image sequence and sent offline to the client.
2. The digital life generation method according to claim 1, wherein The step of real-time sending the audio feature sequence and / or the text feature sequence to the client, so that the client performs digital human generation based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it, includes: Real-time sending the audio feature sequence and / or the text feature sequence to the client, so that the client performs feature fusion based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, and performs digital human generation based on the fused feature sequence obtained by feature fusion, to obtain a target face image sequence and display it.
3. The digital life generation method according to claim 1, wherein An encoder in an encoding and decoding model for digital human generation is deployed on the server side, and a decoder in the encoding and decoding model is deployed on the client side; The encoder is used to perform feature extraction on the face mask image sequence to obtain a face mask feature sequence; The face mask image sequence is determined based on the original face image sequence; The decoder is used to perform digital human generation based on the audio feature sequence and / or the text feature sequence, and the face mask feature sequence, to obtain a target face image sequence.
4. The digital life generation method according to any one of claims 1 to 3, characterized in that The face mask feature sequence is determined based on the following steps: Performing face detection on each original face image in the original face image sequence to obtain the original face region in each original face image; Based on each original face region, performing image masking on each original face image in the original face image sequence to obtain a face mask image sequence; Performing feature extraction on the face mask image sequence to obtain a face mask feature sequence, and sending the face mask feature sequence offline to the client, so that the client stores the face mask feature sequence and establishes an index; When performing digital human generation, the client can obtain the face mask feature sequence through the index.
5. A method for generating a digital life, characterized in that, Applied to the client side, including: Receiving the audio feature sequence and / or the text feature sequence sent by the server in real time; Performing digital human generation based on the audio feature sequence and / or the text feature sequence, and a face mask feature sequence, to obtain a target face image sequence and display it; Wherein, the face mask feature sequence is obtained by the server through pre-performing feature extraction on the original face image sequence and sent offline to the client; The audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction based on the real-time acquired target audio and / or target text.
6. The digital life generation method according to claim 5, wherein The digital human generation based on the audio feature sequence and / or the text feature sequence, as well as the face mask feature sequence, to obtain and display a target face image sequence includes: Performing feature fusion based on the audio feature sequence and / or the text feature sequence, as well as the face mask feature sequence, to obtain a fused feature sequence; Performing digital human generation based on the fused feature sequence to obtain and display a target face image sequence.
7. The digital life generation method according to claim 6, wherein, The performing digital human generation based on the fused feature sequence to obtain and display a target face image sequence includes: Performing digital human generation based on the fused feature sequence to obtain a target face region sequence; Performing image fusion based on the target face region sequence and the original face image sequence to obtain and display a target face image sequence.
8. A digital life generation device, characterized in that, Applied to the server, it includes: An acquisition unit for real-time acquiring a target audio and / or a target text for digital human generation; An extraction unit for performing feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence; A sending unit for real-time sending the audio feature sequence and / or the text feature sequence to the client, so that the client performs digital human generation based on the audio feature sequence and / or the text feature sequence, as well as the face mask feature sequence, to obtain and display a target face image sequence; The face mask feature sequence is determined by the server in advance based on the original face image sequence and sent offline to the client.
9. A digital life generation device, characterized in that, Applied to the client, it includes: A receiving unit for real-time receiving the audio feature sequence and / or the text feature sequence sent by the server; A generating unit for performing digital human generation based on the audio feature sequence and / or the text feature sequence, as well as the face mask feature sequence, to obtain and display a target face image sequence; Wherein, the face mask feature sequence is obtained by the server through feature extraction based on the original face image sequence in advance and sent offline to the client; The audio feature sequence and / or the text feature sequence are obtained by the server through feature extraction based on the real-time acquired target audio and / or target text.
10. A digital life generation system, characterized in that, It includes a server and a client, and the server and the client are communicatively connected; The server is used for real-time acquiring a target audio and / or a target text for digital human generation, performing feature extraction based on the target audio and / or the target text to obtain an audio feature sequence and / or a text feature sequence, and real-time sending the audio feature sequence and / or the text feature sequence to the client; The client is used for real-time receiving the audio feature sequence and / or the text feature sequence sent by the server, and performing digital human generation based on the audio feature sequence and / or the text feature sequence, as well as the face mask feature sequence, to obtain and display a target face image sequence. The face mask feature sequence is determined by the server in advance based on the original face image sequence and sent offline to the client.
11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the digital human generation method according to any one of claims 1 to 7.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the digital human generation method according to any one of claims 1 to 7.