Digital human video generation method, device, storage medium, and program product
By migrating the digital human inference model from the server to the mobile terminal and using audio data to generate images of continuously changing mouth shapes, the problems of slow inference speed and video playback delay caused by excessive server load were solved, thus improving the playback quality and mouth shape adaptation of digital human videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 58 CHANG LIFE (BEIJING) INFORMATION TECH CO LTD
- Filing Date
- 2024-10-09
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, the server may experience excessive load when executing multiple inference tasks simultaneously, leading to untimely responses, slow inference speed, and latency and stuttering issues when playing digital human videos in poor network environments.
The digital human reasoning model is migrated from the server to the mobile terminal for reasoning. Audio data is used to generate facial region images with continuously changing mouth shapes, which drive the digital human's action scenes and play audio data to generate target digital human videos.
It improved the reasoning speed of digital human models, reduced video playback latency and stuttering, enhanced the matching of lip movements with audio content, and improved video playback quality.
Smart Images

Figure CN119136018B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a method, device, storage medium and program product for generating digital human videos. Background Technology
[0002] With the continuous development of artificial intelligence, AI digital human technology has been widely applied. Currently, most solutions for generating digital humans rely on a server to complete the inference task, and then transmit the generated AI digital human video to a mobile terminal.
[0003] However, this approach has the following problems: When the server needs to execute multiple inference tasks simultaneously, the server may experience slow response times due to excessive load, resulting in slow inference speed. Additionally, after the server generates the digital human video, it needs to transmit it to the mobile terminal via the network. In poor network conditions, this may cause delays and stuttering when playing the digital human video on the mobile terminal. Summary of the Invention
[0004] This application provides a digital human video generation method, apparatus, storage medium, and program product to improve model inference speed and playback quality of digital human videos.
[0005] This application provides a method for generating digital human videos, applied to a mobile terminal running a target digital human inference model. The method includes: acquiring audio data for generating digital human voice, an image resource package for generating digital human image, and the frame rate of the digital human video to be generated. The image resource package includes multiple original digital human base maps, original facial region images in each digital human base map, and position information of each original facial region image in its corresponding original digital human base map; converting the audio data into target audio feature information required by the target digital human inference model; and segmenting the target audio feature information according to the duration and frame rate of the audio data to obtain multiple audio feature data. The process involves: inputting multiple audio feature fragments into the target digital human inference model; mapping these fragments to multiple mouth deformation features; generating multiple target facial region images with continuously changing mouth shapes based on these features; replacing the original facial region images in the original digital human base image with multiple target facial region images based on their position information within the corresponding original digital human base image, thus obtaining multiple target digital human images with continuously changing mouth shapes; and generating digital human motion graphics based on these images, driving the digital human's spoken audio data within the motion graphics to obtain the target digital human video.
[0006] The application also provides an electronic device, including: a memory and a processor; the memory for storing a computer program; and the processor, coupled to the memory, for executing the computer program to implement the steps in the above-described digital human-based live streaming method.
[0007] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps in the live streaming method based on digital humans.
[0008] This application also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, enable the processor to implement the steps in the above-described live streaming method for digital humans.
[0009] In this embodiment, the original inference environment of the target digital human inference model is changed from server-side inference to mobile terminal inference. This avoids the problem of slow inference speed caused by excessive server load when multiple inference tasks need to be executed simultaneously, thus improving the inference speed of the target digital human model. Furthermore, the target digital human inference model uses multiple audio feature segments corresponding to audio data to infer multiple target facial region images with continuously changing mouth shapes. Based on these multiple images of the target digital human with continuously changing mouth shapes, a digital human motion frame is generated, driving the digital human to speak audio data in the motion frame to obtain the target digital human video. This not only improves the adaptation between mouth shape and audio content but also reduces latency and stuttering issues when playing the digital human video, thus improving the playback quality of the target digital human video. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0011] Figure 1 A flowchart illustrating a digital human video generation method provided as an exemplary embodiment of this application;
[0012] Figure 2 A schematic flowchart illustrating a model conversion process provided for an exemplary embodiment of this application;
[0013] Figure 3 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0016] To address the technical issues of slow server-side inference speed and latency and stuttering when playing digital human videos on mobile terminals, this application embodiment modifies the original inference environment of the target digital human inference model, changing it from server-side inference to mobile terminal inference. This avoids the problem of slow inference speed caused by server overload when multiple inference tasks need to be executed simultaneously, thus improving the inference speed of the target digital human model. Furthermore, the target digital human inference model uses multiple audio feature segments corresponding to audio data to infer multiple target facial region images with continuously changing mouth shapes. Based on these multiple images of the target digital human with continuously changing mouth shapes, a digital human motion frame is generated, driving the digital human to speak audio data in the motion frame to obtain the target digital human video. This not only improves the adaptation between mouth shape and audio content but also reduces latency and stuttering issues when playing the digital human video, improving the playback quality of the target digital human video.
[0017] The following describes in detail a solution provided by an embodiment of this application, with reference to the accompanying drawings.
[0018] Figure 1 A flowchart illustrating a digital human video generation method provided as an exemplary embodiment of this application. Figure 1 As shown, this method is applied to a mobile terminal that runs a digital human reasoning model. For ease of description and distinction, the digital human reasoning model is referred to as the target digital human reasoning model. The method includes:
[0019] 101. Obtain the audio data used to generate the digital human's voice, the image resource package used to generate the digital human's image, and the frame rate of the digital human video to be generated. The image resource package contains multiple original digital human base maps, original facial region images in each digital human base map, and the position information of each original facial region image in the corresponding original digital human base map.
[0020] 102. Convert the audio data into target audio feature information required by the target digital human inference model, and perform segmentation on the target audio feature information according to the duration and frame rate of the audio data to obtain multiple audio feature segments.
[0021] 103. Input multiple audio feature segments into the target digital human inference model, map the multiple audio feature segments into multiple mouth deformation features, and generate multiple target facial region images with continuously changing mouth shape based on the multiple mouth deformation features.
[0022] 104. Based on the position information of each original facial region image in the corresponding original digital human base image, replace the original facial region images in multiple original digital human base images with multiple target facial region images to obtain multiple target digital human images with continuously changing mouth shapes.
[0023] 105. Based on multiple images of a target digital human whose mouth shape changes continuously, generate a digital human motion picture, and drive the digital human's spoken audio data in the digital human motion picture to obtain a target digital human video.
[0024] In this application embodiment, the implementation form of the mobile terminal is not limited. The mobile terminal can be a smart handheld device, such as a smartphone, tablet computer, laptop computer, or desktop computer, etc.; or, the mobile terminal can also be a smart wearable device, such as a smartwatch, smart bracelet, etc.; or, the mobile terminal can also be various smart home appliances with display screens, such as smart TVs, smart large screens, or smart robots, etc.
[0025] In this embodiment, audio data refers to the audio data containing speech content. This audio data can be pre-recorded or generated in advance or in real-time by the server or mobile terminal based on text information. The timing of audio data generation depends on the specific application scenario, and this embodiment does not impose specific limitations on this. For example, the real-time flexible dialogue audio data in an AI interview room can be audio data used for responses generated in real-time based on the text information corresponding to the input content of the interlocutors, while the audio data corresponding to fixed phrases in some dialogue scenarios can be pre-recorded or generated in advance based on text information.
[0026] To determine the deformation of the mouth in response to audio data, this embodiment introduces a digital human inference model. This model is at least an AI model used to infer mouth shape based on audio data. Specifically, it can predict and generate mouth shape animations synchronized with the input audio data. This can be used to generate realistic animations of digital humans or virtual characters, allowing the character to move its mouth naturally according to the audio data, thus improving the accuracy of mouth deformation recognition. This embodiment does not limit the type of digital human inference model; for example, it can be Wav2Lip, ER-NeRF, VideoRetalking, MuseTalk, EchoMimic, etc. Among them, Wav2Lip is an open-source lip-sync tool that generates matching lip animations based on input speech. It is applicable to any face and any language and can generate realistic lip movements. ER-NeRF is a real-time Talking Portrait Synthesis model that synchronizes given audio with the lip movements of a person in a video. It should be noted that specific videos need to be trained to generate the corresponding digital human inference model. VideoRetalking can synchronize the lip movements of a person in a video with the input sound. This model includes face video generation, audio-driven lip-sync, and facial enhancement. MuseTalk is a real-time audio-driven lip-sync model that supports multiple language inputs and is particularly suitable for high-definition facial areas. EchoMimic can give static images vivid speech and expressions. It uses a deep learning model to combine audio and facial landmarks to create highly realistic dynamic portrait videos. These models analyze the input speech signal to generate corresponding lip animations, so that the lip movements of the digital human character when speaking match the actual speech content. This improves the accuracy of lip deformation recognition and enhances the realism of the character and the immersion of the viewer. For ease of distinction and description, the digital human reasoning model in actual application is referred to as the target digital human reasoning model. In this embodiment, the focus is on the Wav2Lip model.
[0027] It should be noted that the target digital human inference model is obtained by converting the initial digital human inference model into a different format. The initial digital human inference model is developed on the server using a target model development framework suitable for the server. Specifically, when the mobile terminal requires the target digital human model, it sends a model conversion request message to the server. This message includes the mobile terminal's system type, available resource information, and the model functions required by the mobile terminal, requesting the server to convert the initial digital human inference model into the target digital human inference model based on the system type, available resource information, and model functions. Further, the mobile terminal receives a model conversion response message from the server, which instructs it to retrieve the model file of the target digital human inference model from a remote persistent storage space. The model file is then retrieved from the remote persistent storage space and stored in the local persistent storage space. When using the target digital human inference model, the model file is loaded from the local persistent storage space to run the target digital human inference model.
[0028] It should also be noted that the development process of the initial digital human model may or may not include a training process. If the development process does not include a training process, the training process can be completed on the server before model format conversion, on the mobile terminal before model format conversion, or on the mobile terminal after model format conversion. The development, training, and model format conversion processes of the initial digital human inference model can be found in the relevant descriptions of the following embodiments, and will not be repeated here.
[0029] In this embodiment, the specific implementation of the target digital human inference model running on the mobile terminal is not limited. For example, the mobile terminal can directly run the target digital human inference model; or, the target digital human inference model can run on an application on the mobile terminal. For ease of description and distinction, this application is referred to as the target application. Furthermore, this embodiment does not limit the implementation form of the target application. The target application can be a standalone app, a webpage, or a mini-program that depends on an app for operation.
[0030] In this embodiment, to obtain a digital human video, audio data for generating the digital human's voice and an image resource package for generating the digital human's image can be acquired first. This allows for inference of lip shape changes corresponding to the digital human image based on the audio data, and further, based on the audio data and lip shape changes, a digital human video can be obtained. The image resource package may contain only multiple original digital human base images, or it may simultaneously contain multiple original digital human base images, original facial region images from each base image, and position information of each original facial region image within its corresponding base image. This embodiment does not limit this. Multiple original digital human base images may contain multiple sets of original digital human base images. Each set of original digital human base images can be multiple consecutive digital human images extracted from any digital human video, and the lip shape in these consecutive digital human images changes continuously.
[0031] The original audio data format may differ from the format required by the target digital human inference model. Therefore, after acquiring the audio data, to facilitate the target digital human inference model's reasoning about lip movements corresponding to the audio data, the original audio data format can be converted to the format required by the target digital human inference model. This involves converting the initial audio feature information of the audio data into the target audio feature information required by the target digital human inference model. It should be noted that different digital human inference models may require the same or different types of audio feature information. Taking the target digital human inference model as a Wav2Lip model as an example, the original audio data format is .wav, while the Wav2Lip model requires Mel-spectral data format. Therefore, the audio data needs to be converted from .wav format to Mel-spectral data format, that is, the initial audio feature information of the audio data is converted into Mel-spectral data required by the Wav2Lip model. Mel-spectral data corresponds to a Mel-spectral graph.
[0032] In this embodiment, converting audio data into audio feature information required by the target digital human inference model includes: extracting Fourier transform data from the audio data; generating Mel spectrum data based on the Fourier transform data; and cropping, padding, or resizing the Mel spectrum data according to the requirements of the target digital human inference model for the input data to obtain the target audio feature information. Cropping refers to removing edge noise from the Mel spectrum data or retaining the required spectral portion. Padding refers to adding zero values or other specific values to the edges of the data; typically, when the size of the Mel spectrum data does not meet the model's input requirements, padding can be used to adjust the data size to meet the requirements. Resizing refers to resampling or interpolation to change the resolution or size of the Mel spectrum data to adapt it to a specific input format. The target audio feature information refers to the Mel spectrum feature information corresponding to the Mel spectrum data processed based on the requirements of the target digital human inference model's input data.
[0033] In one optional embodiment, Fourier transform data is extracted from the audio data, and spectral data is generated based on the Fourier transform data. This includes: first, loading the audio data using an appropriate library, such as using the Librosa library to load .wav format audio data; preprocessing the audio data, including but not limited to normalizing volume and removing noise, to improve the accuracy of subsequent processing; converting the audio data into an audio signal, which corresponds to key feature information; performing a Fast Fourier Transform (FFT) on the audio signal to convert it from the time domain to the frequency domain; and using a Mel filter to convert the audio signal in the frequency domain to the Mel scale, generating a Mel spectrogram, which corresponds to Mel spectral data. Further, a Discrete Cosine Transform (DCT) can be performed on the Mel spectrogram to extract Mel Frequency Cepstral Coefficients (MFCCs). The Mel Frequency Cepstral Coefficients are a characteristic representation of the Mel spectral data, i.e., the Mel spectral coefficients are the Mel spectral feature information corresponding to the Mel spectral data.
[0034] Alternatively, in order to improve the accuracy of the inference of the target digital human inference model, the target audio feature information can be segmented, and the target digital human inference model can be used to infer each segment of audio feature information separately to obtain multiple digital human images with continuously changing mouth shapes.
[0035] In one optional embodiment, the target feature information can be segmented based on the duration of the audio data and the frame rate of the digital human video to be generated. The frame rate refers to the number of static images (frames) that can be displayed per second during video playback. Frame rate is an important indicator of video playback smoothness and is usually measured in "Frames Per Second" (FPS). For example, if the frame rate of the digital human to be generated is 25fps, then 25 images will be played within 1 second. Based on this, the segmentation operation of the target audio feature information includes: obtaining the frame rate of the digital human video to be generated; and segmenting the target feature information according to the duration of the audio data and the frame rate of the digital human video to be generated, resulting in multiple audio feature segments.
[0036] When determining the frame rate for the digital human video to be generated, factors such as the intended use of the video, the viewers' devices, the playback platform, and the desired visual effects need to be considered. For example, if the digital human video is intended for online playback and needs to display smooth facial expressions and synchronized lip movements, a frame rate of 30fps or higher can be selected; if the digital human video is intended for film or art projects, 24fps can be selected to suit the viewing experience. Furthermore, technical limitations and costs must be considered. Higher frame rates typically mean processing more data during editing and rendering, which may increase the demand for computing resources and costs. Simultaneously, the size of high frame rate digital human video files will also increase, potentially affecting storage and transmission efficiency. In practical applications, when used in cloud-based intelligent digital human services, the video frame rate is generally no less than 25fps and no more than 60fps; when used for silicon-based digital human cloning, the video frame rate is generally 30fps or 60fps, thus ensuring smoothness and detail while also considering the technical and cost limitations of practical applications.
[0037] Optionally, based on the duration and frame rate of the audio data, the target audio feature information is segmented to obtain multiple audio feature segments. This includes: performing a quotient operation by using the audio data duration as the divisor and the frame rate as the divisor to obtain the video frame duration, where the video frame duration represents the duration of each video frame; and using the video frame duration as the segmentation unit to segment the target audio feature information to obtain multiple audio feature segments. For example, if the audio data duration is 50 seconds and the frame rate of the digital human video to be generated is 25 fps, then...
[0038] Furthermore, after obtaining multiple audio feature segments, these segments can be input into the target digital human inference model to map them into multiple mouth deformation features. Based on these mouth deformation features, multiple target facial region images with continuously changing mouth shapes can be generated.
[0039] In this embodiment, multiple audio feature segments are input into a target digital human inference model, and the multiple audio feature segments are mapped into multiple mouth deformation features. Based on the multiple mouth deformation features, multiple target facial region images with continuously changing mouth shapes are generated. The process includes: inputting multiple audio feature segments into the target digital human inference model; inferring changes in the mouth shape of the digital human based on the multiple audio feature segments to obtain multiple mouth deformation features; converting the multiple mouth deformation features into multiple pixel data and positional distribution information of each pixel in the facial region; and generating multiple target facial region images based on the multiple pixel data and positional distribution information of each pixel in the facial region corresponding to the multiple mouth deformation features.
[0040] In one optional embodiment, the target digital human inference model includes an audio feature processing network layer, which infers mouth shape changes based on multiple audio feature segments to obtain multiple mouth deformation features. This includes inputting multiple audio feature segments into the audio feature processing network layer to generate mouth deformation features corresponding to each audio feature segment. Each mouth deformation feature is a multi-dimensional feature vector, which includes at least a mouth shape parameter, an opening / closing parameter, and a lip position parameter. The mouth shape parameter represents the current mouth shape, the opening / closing parameter represents the current mouth opening / closing degree, and the lip position parameter represents the current mouth position within the face region.
[0041] Optionally, the target digital human inference model may further include: a visual feature processing network layer, which converts multiple mouth deformation features into multiple pixel data and positional distribution information of each pixel in the face region, including: inputting multiple mouth deformation features into the video feature processing network layer, and converting each mouth deformation feature into multiple pixel data and positional distribution information of each pixel in the face region based on the mouth shape parameters, opening and closing parameters and lip position parameters corresponding to each mouth deformation feature.
[0042] Optionally, the target digital human inference model also includes a generator. After obtaining multiple pixel data and the positional distribution information of each pixel, the multiple pixel data and the positional distribution information of each pixel can be input into the generator. Based on the multiple pixel data and the positional distribution information of each pixel, multiple target facial region images with continuously changing mouth shapes are obtained.
[0043] As can be seen from the above embodiments, the image resource package contains multiple original digital human base images, original facial region images in each digital human base image, and position information of each original facial region image in the corresponding original digital human base image. After obtaining multiple target facial region images with continuously changing mouth shapes, the original facial region images in the multiple original digital human base images can be replaced with the multiple target facial region images according to the position information of each original facial region image in the corresponding original digital human base image, so as to obtain multiple target digital human images with continuously changing mouth shapes.
[0044] Furthermore, after obtaining multiple images of the target digital human with continuously changing mouth shapes, a video of the target digital human can be obtained based on these images. This embodiment does not limit the specific implementation method of obtaining a video of the target digital human based on multiple images of the target digital human with continuously changing mouth shapes; the implementation method may vary depending on the application scenario.
[0045] In one optional embodiment, obtaining a target digital human video based on multiple target digital human images showing continuous changes in mouth shape includes: generating digital human motion frames based on the multiple target digital human images showing continuous changes in mouth shape, and driving a digital human in the digital human motion frames to verbally play the audio data to obtain the target digital human video. In this target digital human video, the digital human motion frames and the verbally played audio are two independent parts, suitable for real-time dialogue scenarios, such as real-time question-and-answer sessions in AI live streaming rooms / interview rooms.
[0046] In another optional embodiment, a target digital human video is obtained based on multiple target digital human images with continuously changing mouth shapes, including: generating the target digital human video based on the multiple target digital human images with continuously changing mouth shapes and audio data used to generate the digital human's voice. In this target digital human video, the digital human's motion visuals and spoken audio are a unified whole, suitable for non-real-time playback scenarios, such as non-real-time live-streamed digital human videos. In this embodiment, correspondingly, the generator in the target digital human inference model can also obtain the target digital human video based on multiple target digital human images with continuously changing mouth shapes. Further, the target digital human inference model can also include a lip-sync discriminator for evaluating the synchronization accuracy between the generated video frames and audio data. If the evaluation result does not meet the synchronization accuracy threshold, the target digital human inference model can be retrained based on the synchronization accuracy standard, and the retrained target digital human inference model can be used to generate the target digital human video based on the multiple target digital human images with continuously changing mouth shapes and audio data used to generate the digital human's voice. Furthermore, the target digital human inference model can also include a visual quality discriminator to evaluate whether the generated video frames look natural and without obvious generation artifacts. Similarly, if the evaluation results do not meet the synchronization accuracy threshold, the target digital human inference model can be retrained based on the synchronization accuracy standard. The retrained target digital human inference model can then be used to generate a target digital human video based on multiple target digital human images with continuously changing mouth shapes and audio data used to generate the digital human's voice.
[0047] In the above embodiments of this application, the original inference environment of the target digital human inference model is changed from inference execution on the server side to inference on the mobile terminal. This avoids the problem that the server may be slow due to excessive load when multiple inference tasks need to be executed at the same time, thus improving the inference speed of the target digital human model. In addition, the target digital human inference model uses multiple audio feature segments corresponding to audio data to infer multiple target facial region images with continuously changing mouth shapes. Based on multiple target digital human images with continuously changing mouth shapes, a digital human motion screen is generated, which drives the digital human to speak audio data in the digital human motion screen to obtain the target digital human video. This not only improves the adaptation between mouth shape and audio content, but also reduces the latency and stuttering problems when playing the digital human video, thus improving the playback quality of the target digital human video.
[0048] The following is combined with Figure 2 The conversion process of the target digital human model is described in detail.
[0049] S1. Receive a model conversion request message sent by the mobile terminal. The model conversion request message includes at least the system type of the mobile terminal, available resource information, and the model functions required by the mobile terminal.
[0050] S2. From the digital human reasoning model developed based on the target model development framework, determine the initial digital human reasoning model that is compatible with the system type and required model functions of the mobile terminal. The target model development framework refers to the model development framework applicable to the server.
[0051] S3. If the available resource information of the mobile terminal meets the resource loading conditions required by the data packet, determine the target format conversion tool according to the original model format supported by the initial digital human inference model and the model-format conversion tool mapping table. The model-format conversion tool mapping table stores the format conversion tools corresponding to digital human inference models that support different model formats.
[0052] S4. Call the target format conversion tool to convert the initial digital human inference model to obtain a target digital human inference model that supports the target model format. The target digital human inference model is suitable for running on mobile terminals.
[0053] In this application embodiment, the system type includes, but is not limited to, Android, iOS, or HarmonyOS. Available resource information includes, but is not limited to,: CPU information, such as model and speed, number of cores and threads, and cache size; memory (RAM) information, such as total memory capacity, used and available memory, memory speed, and memory timings; storage information, such as hard disk or solid-state drive capacity, available space, read / write speed, and interface type (e.g., SATA, NVMe); graphics processing unit (GPU) information, such as model, video memory size, and GPU utilization; network information, such as network adapter type, IP address, network bandwidth, and network connection status; and hardware status, such as device health status and fault information.
[0054] In this embodiment of the application, determining an initial digital human reasoning model adapted to the system type and required model functions of a mobile terminal from digital human reasoning models developed based on the target model development framework includes: selecting a first candidate digital human reasoning model adapted to the system type of the mobile terminal from the digital human reasoning models developed based on the target model development framework, wherein different digital human reasoning models are configured with their adapted operating system lists; determining a second candidate digital human reasoning model adapted to the required model functions of the mobile terminal from the functional description information of the first candidate digital human reasoning model; if there is only one second candidate digital human reasoning model, using the second candidate digital human reasoning model as the initial digital human reasoning model; if there are multiple second candidate digital human reasoning models, selecting the one with the smallest model file as the initial digital human reasoning model.
[0055] Optionally, the model file of the initial digital human inference model is obtained; based on the size of the model file, it is determined whether the bandwidth information and / or memory information in the available resource information of the mobile terminal meet the minimum bandwidth value and / or minimum memory amount required in the resource loading conditions; if so, it is determined that the available resource information of the mobile terminal meets the resource loading conditions.
[0056] In this embodiment, a target format conversion tool is invoked to convert the initial digital human inference model to a target digital human inference model that supports the target model format. This includes: invoking the target format conversion tool to parse the model file of the initial digital human inference model to obtain the network structure and weight parameter set of the initial digital human inference model. The network structure describes the multiple network layers included in the initial digital human inference model and the connection relationships between the multiple network layers. The weight parameter set includes the original weight parameters of each network layer. The weight parameters represent the degree of influence of the input data on the model output. Specifically, the weights determine the degree to which the input data should be amplified or reduced before being passed to the next layer. Furthermore, by simulating or performing control flow analysis on the initial digital human inference model, a static computation graph of the initial digital human inference model is obtained. The static computation graph includes the processing logic from input data to output data for each network layer and the data flow dependencies between network layers. Based on the target model format, the network structure, weight parameter set, and static computation graph of the initial digital human inference model are reorganized to obtain the target digital human inference model.
[0057] In this embodiment of the application, the model file of the initial digital human inference model is parsed to obtain the network structure and weight parameter set of the initial digital human inference model. This includes: parsing a sub-file with a specific format in the model file to obtain the weight parameter set of the initial digital human inference model; and defining the model structure of the initial digital human inference model based on the forward propagation function in the model file to obtain the network structure of the initial digital human inference model.
[0058] In this embodiment of the application, a static computation graph of the initial digital human inference model is obtained by simulating the initial digital human inference model. This includes: inputting preset exemplary input data into the initial digital human inference model to simulate the initial digital human inference model; during the simulation, tracking and recording the computational operations involved in each network layer of the initial digital human inference model and the dependencies between the computational operations; and constructing a static computation graph based on the computational operations involved in each network layer and the dependencies between the computational operations.
[0059] In this embodiment, control flow analysis is performed on the initial digital human inference model to obtain its static computation graph. This includes: compiling the forward propagation function in the model file of the initial digital human inference model to transform the control flow statements and computational operations in the forward propagation function into an initial computation graph; analyzing the dependencies between computational operations and adding these dependencies to the initial computation graph to obtain the static computation graph. The forward propagation function refers to...
[0060] In this embodiment, the network structure, weight parameter set, and static computation graph of the initial digital human inference model are reorganized according to the target model format to obtain the target digital human inference model. This includes: converting operators in each network layer of the network structure into target operators in the target model format based on an operator conversion table; the operators are used to implement computational operations involved in the network layers; the operator conversion table stores the correspondence between the original operators in the original model format and the target operators in the target model format; and converting the original weight parameters of each network layer in the weight parameter set into target weight parameters in the target model format based on the weight parameter conversion relationship. The parameter transformation relationship stores the correspondence between the original weight parameters in the original model format and the target weight parameters in the target model format; based on the static computation graph, optimization operations are performed on the network layers, the data flow dependencies between network layers, and / or the target operators in the network layers to obtain the optimized network layers, the data flow dependencies between network layers, and / or the target operators in the network layers, which serve as the optimization objects; based on the optimization objects, the target weight parameters, and the unoptimized objects, the target digital human inference model is generated. The unoptimized objects refer to the network layers, the data flow dependencies between network layers, and the unoptimized objects in the target operators of the network layers.
[0061] In this embodiment, based on the operator conversion table, operators in each network layer of the network structure are converted into target operators in the target model format. This includes: traversing the original operators in each network layer of the network structure; for the currently traversed original operator, querying the operator conversion table; if the currently traversed original operator is found in the operator conversion table, replacing the currently traversed original operator with its corresponding target operator; if the currently traversed original operator is not found in the operator conversion table, replacing the currently traversed original operator with a custom operator, wherein the custom operator supports the target model format.
[0062] In this embodiment, the initial digital human inference model is a neural network model for generating multiple facial region images with continuously changing mouth shapes based on audio feature information. The method further includes: acquiring sample audio data, multiple labeled facial region images with continuously changing mouth shapes corresponding to the sample audio data, and sample frame rate; converting the sample audio data into sample audio feature information required by the basic digital human inference model, and segmenting the sample audio feature information according to the duration and sample frame rate of the sample audio data to obtain multiple sample audio feature segments; inputting the multiple sample audio feature segments into the basic digital human inference model, mapping the multiple sample audio feature segments into multiple sample mouth deformation features, and generating multiple sample facial region images with continuously changing mouth shapes based on the multiple sample mouth deformation features; calculating the loss function between the multiple sample facial region images and the multiple labeled facial region images, and adjusting the model parameters of the basic digital human inference model until the loss function reaches the set model training termination condition if the loss function does not reach the set model training termination condition, so as to obtain the initial digital human inference model.
[0063] To facilitate understanding, the conversion process of the initial digital human reasoning model will be described in detail below with specific examples.
[0064] The conversion process of the digital human inference model mainly includes the following steps: Model input and parsing: Extracting inference logic by loading weights and defining the model structure; Computation graph extraction: Converting the PyTorch dynamic graph into a static computation graph through tracing or scripting; Operator mapping and weight transfer: Mapping PyTorch operators to the target framework and converting weights; Computation graph optimization: Optimizing the computation graph to improve inference efficiency; Generating the target format model: Generating the final model file in target formats such as ONNX and TensorFlow; Verification and adjustment: Verifying whether the converted model is consistent with the original model and making necessary adjustments. A more detailed process is as follows:
[0065] 1. Model Input and Parsing: When an inference model (e.g., a PyTorch .pth file) is input into a conversion tool, the first step is to parse and read the model's network structure and weight parameters. In PyTorch, this means loading the model and using a predefined architecture to load the saved weights. Specifically: using `torch.load()` to load the .pth file loads the model's weight parameters. Next, the model's computation graph is constructed using a predefined network architecture (usually implemented through the subclass `torch.nn.Module`). At this stage, the model's inference logic (i.e., the model's forward propagation function `forward()`) is extracted, recording the operations and data flow of each layer of the model.
[0066] 2. Extraction of the model's computational graph (Tracing or Scripting): To convert the model to the target format, the model's computational graph must be converted from a dynamic graph framework (such as PyTorch) to a static graph. This can be done in two ways:
[0067] A. Tracing mode
[0068] In tracing mode, the transformation tool tracks the sequence of operations performed by the model during forward propagation by feeding it a set of example inputs. This method captures the model's computational graph, but if the model has branching logic or dynamic operations, tracing may not fully capture these dynamic behaviors. For example, control flow operations (such as if-else statements) may be ignored. Specifically, the transformation tool passes example inputs to the model, tracks all tensor operations involved, records the computational operations at each layer (such as convolution, fully connected layers, pooling, etc.) and their corresponding weights, and builds a complete computational graph that describes how the input data is transformed into output data through each layer.
[0069] B. Scripting Mode
[0070] In Scripting mode, the model's Python code is transformed into a static graph, enabling the handling of more complex dynamic control flows. This approach is more suitable for models with complex dynamic logic. The transformation tool analyzes each operation in the `forward()` function and generates a static graph. Specifically, the tool compiles the model's `forward()` function, transforming its control flow (such as `if` and `for` statements) and tensor operations into a static computation graph; it also analyzes the dependencies between operations to generate the model's complete computation graph.
[0071] 3. Conversion between operators and weights
[0072] The generated computation graph contains the model's hierarchical structure, the types of operations at each layer (e.g., convolution, pooling, activation functions), and the data flow dependencies between them. Next, the tool needs to convert these operators and weights into the target format.
[0073] A. Operator Mapping
[0074] Different deep learning frameworks (such as PyTorch, TensorFlow, and ONNX) have their own operator sets. During the conversion process, the tool maps PyTorch operators to operators of the target framework according to the requirements of the target format. For example, PyTorch's `torch.nn.Conv2d` will be mapped to the `Conv` operator in ONNX; PyTorch's ReLU activation function will be mapped to the ReLU operator in ONNX. This mapping is achieved through an operator transformation table. The tool checks the operation type of each layer and looks up the corresponding operator in the target format. If some operations are not supported in the target format, manual adjustments or the use of custom operators may be necessary.
[0075] B. Weight Transfer
[0076] While performing operator transformations, the tool also needs to extract the weight parameters (such as convolution kernels and biases) of each layer from the .pth file and convert them to the target format. For example, weights in PyTorch are usually torch.Tensor objects, while in ONNX, weights are converted into a standard multidimensional array format. These weights need to be saved according to the format and requirements of the target framework. The weight transfer process ensures that the model can use the same trained parameters for inference in the target framework.
[0077] 4. Optimization of the computation graph
[0078] After generating the computation graph in the target format, the conversion tool typically performs a series of optimizations to improve model inference efficiency. Common optimizations include: layer merging: for example, merging convolutional layers and batch normalization (BatchNorm) layers into a single operation to reduce computational cost; constant folding: for some static operations, such as addition and multiplication, the tool may pre-calculate the results during compilation to reduce the computational burden during inference; and redundant node removal: if the model contains unused layers or operations, the tool will remove them to optimize model size and speed.
[0079] 5. Generate the target format model
[0080] After completing operator mapping and weight transfer, the tool generates the final model file according to the target format standard. For example, when converting to ONNX, the tool generates an .onnx file containing the network structure and weights, which can run in ONNX-compatible inference engines (such as ONNX Runtime); when converting to TensorFlow, it may generate a .pb file or save it in TensorFlow SavedModel format.
[0081] 6. Model Validation and Adjustment
[0082] After the conversion is complete, the final step is to verify the consistency of the model. Tools typically verify successful conversion in the following ways: running inference verification: performing inference on both the original and converted models using the same input data to check for consistent output; adjusting incompatible layers: if some layers are not fully supported in the target format, the tool may generate warnings and allow developers to make manual adjustments. For example, certain custom layers may need to be rewritten or alternative implementations provided.
[0083] For example, converting a .pth model to .onnx format mainly involves changing the format of the model's network structure and weights, but the core working principle and structure of the model remain largely unchanged. The following is a detailed explanation:
[0084] In PyTorch, .pth files typically contain two parts: model weight parameters: these are all the parameters of the trained model, such as the weights and biases of each layer; and network structure: this is the code logic that defines how the model receives input, processes it through different layers, and produces output. When converting to .onnx, both the model's network structure and weights are converted to .onnx format. This process preserves the network's computational graph and the trained weights. For the weight parameter conversion, the model's weight parameters are extracted from the .pth file and stored in the .onnx file. These parameters are stored in a format compatible with ONNX. For the network structure conversion, the model's structure is also converted from PyTorch's computational graph representation to ONNX format. This includes the types of each layer (e.g., convolutional layers, fully connected layers, activation functions, etc.) and the connections between them.
[0085] It should be noted that the network structure usually remains unchanged during the conversion process, but there are some situations that require attention:
[0086] 1. Layers not fully compatible with PyTorch and ONNX: Certain special layers or operations implemented in PyTorch may not be directly convertible to ONNX format. These layers may be replaced with equivalent layers supported by ONNX. For example, some specific custom layers or operations, such as certain uncommon dynamic operations, may require manual adjustments or custom implementations during conversion.
[0087] 2. The difference between static and dynamic graphs: PyTorch is a deep learning framework based on dynamic graphs, meaning the network structure is dynamically built at runtime; while ONNX is a static graph, meaning the network structure needs to be fixed during transformation, and all dynamic operations need to be converted into a static computation graph. This transformation may affect some dynamic structures, leading to the need to rewrite the network structure into a compatible static version.
[0088] 3. Changes in Working Principle: The basic working principle of the model usually remains unchanged. However, if incompatible network layers are encountered during conversion (as mentioned above regarding dynamic operations), some adjustments need to be made to ensure correct inference during ONNX runtime (such as ONNX Runtime or other frameworks supporting ONNX). 4. Main Network Layers of the Converted Model: Convolutional Layers (Conv layers): Generally unchanged, still used for feature extraction. Fully Connected Layers: Consistent with PyTorch, still used for classification or regression. Activation Functions: Most common activation functions (such as ReLU, Sigmoid) are retained unchanged. If the network involves dynamic structures, it may be implemented using some static alternatives in ONNX after conversion, with slight adjustments to the working principle.
[0089] 4. Conversion Considerations: Model Input / Output: PyTorch may allow multiple input formats, while ONNX requires explicit input / output tensor shapes. This may require fixing the input size during conversion or explicitly specifying dynamic dimensions. Export Operations: Typically, this involves calling PyTorch torch.onnx.export().
[0090] The method is used for transformation. When exporting, you must ensure that the correct example input is provided so that the computational graph of the model can be exported correctly.
[0091] 5. It's also important to note that when converting a Wav2Lip .pth model to .onnx format, the main changes are to the network structure and weight parameters; the core network structure remains largely unchanged. If the model contains layers or operations incompatible with PyTorch and ONNX, some static substitutions may be necessary. In more special cases, manual adjustments may be required based on the specific model and operations to ensure the network structure behaves consistently in ONNX compared to PyTorch.
[0092] This application provides a model conversion process. Based on a target model development framework, an initial digital human inference model is determined that is compatible with the system type and required model functions of the mobile terminal. The target model development framework refers to a model development framework suitable for the server. When the available resource information of the mobile terminal meets the resource loading conditions required for the data packet, a target format conversion tool is determined according to the original model format supported by the initial digital human inference model and a model-format conversion tool mapping table. The model-format conversion tool mapping table stores format conversion tools corresponding to digital human inference models that support different model formats. The target format conversion tool is called to perform model format conversion processing on the initial digital human inference model to obtain a target digital human inference model that supports the target model format and is suitable for running on the mobile terminal. This can improve the inference speed of the digital human model and reduce subsequent latency and stuttering issues in the digital human video, thereby improving the playback quality of the target digital human video.
[0093] Figure 3 This is a schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this application. For example... Figure 3 As shown, it includes: a memory 30a and a processor 30b; the memory 30a is used to store computer programs; the processor 30b, coupled to the memory 30a, is used to execute the computer programs to achieve the following steps:
[0094] The process involves acquiring audio data for generating the digital human's voice, an image resource package for generating the digital human's image, and the frame rate of the digital human video to be generated. The image resource package contains multiple original digital human base maps, original facial region images within each base map, and the positional information of each original facial region image within its corresponding base map. The audio data is then converted into target audio feature information required by the target digital human inference model. Based on the duration and frame rate of the audio data, the target audio feature information is segmented to obtain multiple audio feature segments. These multiple audio feature segments are then input into the target digital human inference model. Multiple audio feature segments are mapped to multiple mouth deformation features. Based on these mouth deformation features, multiple target facial region images with continuously changing mouth shapes are generated. According to the position information of each original facial region image in the corresponding original digital human base image, the original facial region images in the multiple original digital human base images are replaced with multiple target facial region images to obtain multiple target digital human images with continuously changing mouth shapes. Based on the multiple target digital human images with continuously changing mouth shapes, a digital human motion screen is generated, which drives the digital human's spoken audio data in the digital human motion screen to obtain the target digital human video.
[0095] In this embodiment of the application, when the processor 30b converts audio data into audio feature information required by the target digital human inference model, it is specifically used to: extract Fourier transform data from the audio data; generate Mel spectrum data based on the Fourier transform data; and crop, pad, or resize the Mel spectrum data according to the requirements of the target digital human inference model for the input data to obtain the target audio feature information.
[0096] In this embodiment, multiple audio feature segments are input into the target digital human inference model, and the multiple audio feature segments are mapped into multiple mouth deformation features. When the processor 30b generates multiple target facial region images with continuously changing mouth shapes based on the multiple mouth deformation features, it is specifically used to: input multiple audio feature segments into the target digital human inference model, infer the mouth shape changes of the digital human based on the multiple audio feature segments, and obtain multiple mouth deformation features; convert the multiple mouth deformation features into multiple pixel data and position distribution information of each pixel in the facial region; and generate multiple target facial region images based on the multiple pixel data and position distribution information of each pixel in the facial region corresponding to the multiple mouth deformation features.
[0097] In this embodiment of the application, when the processor 30b infers the changes in the mouth shape of the digital human based on multiple audio feature segments to obtain multiple mouth deformation features, it is specifically used to: generate mouth deformation features corresponding to each audio feature segment, wherein the mouth deformation features are multi-dimensional feature vectors, and the multi-dimensional feature vectors include at least mouth shape parameters, opening and closing parameters, and lip position parameters; and convert the multiple mouth deformation features into multiple pixel data and position distribution information of each pixel in the face region, including: converting each mouth deformation feature into multiple pixel data and position distribution information of each pixel in the face region based on the mouth shape parameters, opening and closing parameters, and lip position parameters corresponding to each mouth deformation feature.
[0098] In this embodiment of the application, when the processor 30b performs segmentation operations on the target audio feature information according to the duration and frame rate of the audio data to obtain multiple audio feature segments, it specifically performs the following: using the duration of the audio data as the divisor and the frame rate as the divisor to perform a quotient operation to obtain the video frame duration; using the video frame duration as the segmentation unit, it performs segmentation operations on the target audio feature information to obtain multiple audio feature segments.
[0099] In this embodiment of the application, the target digital human reasoning model is obtained by converting the initial digital human reasoning model into a model format. The initial digital human reasoning model is developed on the server side using a target model development framework suitable for the server side.
[0100] Optionally, the processor 30b is further configured to send a model conversion request message to the server, the model conversion request message including the system type of the mobile terminal, available resource information, and the model functions required by the mobile terminal, to request the server to convert the initial digital human inference model into a target digital human inference model based on the system type, available resource information, and model functions; receive a model conversion response message returned by the server, the model conversion response message being used to notify the mobile terminal to obtain the model file of the target digital human inference model from the remote persistent storage space; obtain the model file of the target digital human inference model from the remote persistent storage space and store it in the local persistent storage space, and when using the target digital human inference model, load the model file from the local persistent storage space to run the target digital human inference model.
[0101] Furthermore, such as Figure 3 As shown, the server also includes other components such as a communication component 30c, a display 30d, a power supply component 30e, and an audio component 30f. Figure 3 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 3 The components shown.
[0102] The detailed implementation methods and beneficial effects of the electronic devices provided in this application have been described in detail in the foregoing embodiments, and will not be elaborated further here.
[0103] Exemplary embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps in the above-described method embodiments.
[0104] An exemplary embodiment of this application also provides a computer program product comprising a computer program / instructions that, when executed by a processor, enable the processor to perform the steps described in the above method embodiments.
[0105] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0106] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.
[0107] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0108] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0109] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0110] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program code.
[0111] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0112] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0114] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.
[0115] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0116] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0117] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0118] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for generating digital human videos, characterized in that, Applied to a mobile terminal, wherein the mobile terminal runs a target digital human reasoning model, the method includes: The process involves acquiring audio data for generating the digital human's voice, an image resource package for generating the digital human's image, and the frame rate of the digital human video to be generated. The image resource package includes multiple original digital human base maps, original facial region images in each digital human base map, and position information of each original facial region image in the corresponding original digital human base map. The audio data is converted into target audio feature information required by the target digital human inference model, and the target audio feature information is segmented according to the duration of the audio data and the frame rate to obtain multiple audio feature segments; wherein, the segmentation operation includes: dynamically calculating the duration of the video frame based on the frame rate as the segmentation unit, and each audio feature segment corresponds to one video frame, so as to realize dynamic real-time synchronization of audio and video actions on the mobile terminal; The plurality of audio feature segments are input into the target digital human reasoning model, the target digital human reasoning model comprising: An audio feature processing network layer is used to generate mouth deformation features corresponding to each audio feature segment; the mouth deformation features are multi-dimensional feature vectors, which include at least mouth shape parameters, opening and closing parameters, and lip position parameters. The visual feature processing network layer is used to convert each mouth deformation feature into multiple pixel data and position distribution information of each pixel in the face region based on the mouth shape parameters, opening and closing parameters and lip position parameters corresponding to each mouth deformation feature. A generator is used to generate multiple target facial region images based on multiple pixel data and the positional distribution information of each pixel in the facial region corresponding to the multiple mouth deformation features. Based on the position information of each original facial region image in the corresponding original digital human base image, the original facial region images in the multiple original digital human base images are replaced with the multiple target facial region images to obtain multiple target digital human images with continuously changing mouth shapes. Based on the multiple target digital human images with continuously changing mouth shapes, a digital human motion scene is generated, and the digital human in the motion scene is driven to play the audio data to obtain the target digital human video.
2. The method according to claim 1, characterized in that, Converting the audio data into the audio feature information required for the target digital human reasoning model includes: Fourier transform data is extracted from the audio data; Mel spectrum data is generated based on the Fourier transform data; according to the requirements of the target digital human inference model for input data, the Mel spectrum data is cropped, padded, or resized to obtain the target audio feature information.
3. The method according to claim 1, characterized in that, Based on the duration and frame rate of the audio data, the target audio feature information is segmented to obtain multiple audio feature segments, including: The duration of the audio data is used as the dividend and the frame rate is used as the divisor to perform a quotient operation to obtain the video frame duration. Using the video frame duration as a segmentation unit, the target audio feature information is segmented to obtain multiple audio feature segments.
4. The method according to any one of claims 1-3, characterized in that, The target digital human reasoning model is obtained by converting the initial digital human reasoning model into a different model format. The initial digital human reasoning model is developed on the server side using a target model development framework suitable for the server side.
5. The method according to claim 4, characterized in that, Also includes: Send a model conversion request message to the server. The model conversion request message includes the system type of the mobile terminal, available resource information, and the model functions required by the mobile terminal, so as to request the server to convert the initial digital human reasoning model into the target digital human reasoning model based on the system type, available resource information, and model functions. The system receives a model conversion response message returned by the server, which is used to notify the mobile terminal to retrieve the model file of the target digital human reasoning model from the remote persistent storage space. The model file of the target digital human reasoning model is obtained from the remote persistent storage space and stored in the local persistent storage space. When using the target digital human reasoning model, the model file is loaded from the local persistent storage space to run the target digital human reasoning model.
6. An electronic device, characterized in that, include: Memory and processor; The memory is used to store a computer program; the processor, coupled to the memory, is used to execute the computer program to implement the steps of the method according to any one of claims 1-5.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-5.
8. A computer program product, characterized in that, The computer program product includes a computer program / instruction that, when executed by a processor, causes the processor to perform the steps of any one of the methods of claims 1-5.