Video generation method, device, electronic device, and storage medium
By extracting the template lip movement parameters from real lip movement videos, and generating and rendering the videos with voice features, the problem of missing or distortion of lip movement in virtual images is solved, and a more realistic and natural virtual image is achieved, improving the human-computer interaction effect.
Patent Information
- Application Number
- CN202210097018.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The prior art is difficult to generate virtual images with a strong sense of reality, especially in voice data-driven videos, where lip movements are prone to missing or distortion, affecting the human-computer interaction effect.
By extracting the template lip movement parameters from the real lip movement video, combining the speech feature sequence to generate the initial image sequence of the target object, and performing time domain rendering processing to generate the target image sequence of the target object.
It improves the authenticity and nature of virtual images, enhances the vividness and intelligence of human-computer interaction, and ensures that the lip movements match the voice data consistently.
Smart Images

Figure CN114429767B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, in particular to artificial intelligence, computer vision, virtual / augmented reality, and other technical fields, and specifically to a video generation method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] With the digitalization of everyday life, avatars have gradually gained public acceptance and become a crucial component of intelligent interaction. They have demonstrated potential in numerous fields, including media, customer service, education, social networking, gaming, healthcare, and film and television. In practical applications, avatars have enhanced the interaction between humans and machines. Summary of the Invention
[0003] The present disclosure provides a video generation method, apparatus, electronic device, storage medium, and program product.
[0004] According to one aspect of the present disclosure, a video generation method is provided, which may include: determining a lip shape feature sequence corresponding to a speech feature sequence, wherein the speech feature sequence is extracted from speech data of a video to be generated; determining a target template lip movement parameter sequence from a plurality of template lip movement parameters based on the lip shape feature sequence, wherein the plurality of template lip movement parameters are extracted from a video including real lip movements; generating an initial image sequence of a target object based on the target template lip movement parameter sequence; and performing time-domain rendering processing on the initial image sequence of the target object to obtain a target image sequence of the target object.
[0005] According to another aspect of the present disclosure, a video generation device is provided, which may include: a feature determination module for determining a lip shape feature sequence corresponding to a speech feature sequence, wherein the speech feature sequence is extracted from speech data of a video to be generated; a lip movement determination module for determining a target template lip movement parameter sequence from a plurality of template lip movement parameters based on the lip shape feature sequence, wherein the plurality of template lip movement parameters are extracted from a video including real lip movements; a generation module for generating an initial image sequence of a target object based on the target template lip movement parameter sequence; and a time domain rendering module for performing time domain rendering processing on the initial image sequence of the target object to obtain a target image sequence of the target object.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a method as disclosed herein.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method of the present disclosure when executed by a processor.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 Schematically illustrates an exemplary system architecture to which the video generation method and apparatus according to an embodiment of the present disclosure may be applied;
[0012] Figure 2 The flowchart of the video generation method according to the embodiment of the present disclosure is schematically shown;
[0013] Figure 3 Schematically shows a flow chart of a video generation method according to another embodiment of the present disclosure;
[0014] Figure 4 A block diagram schematically shows a video generating apparatus according to an embodiment of the present disclosure; and
[0015] Figure 5 The block diagram schematically shows an electronic device suitable for implementing the video generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0017] The present disclosure provides a video generation method, apparatus, electronic device, storage medium, and program product.
[0018] According to an embodiment of the present disclosure, a video generation method is provided, which may include: determining a lip shape feature sequence corresponding to a speech feature sequence, wherein the speech feature sequence is extracted from speech data of a video to be generated; determining a target template lip movement parameter sequence from a plurality of template lip movement parameters based on the lip shape feature sequence, wherein the plurality of template lip movement parameters are extracted from a video including real lip movements; generating an initial image sequence of a target object based on the target template lip movement parameter sequence; and performing time domain rendering processing on the initial image sequence of the target object to obtain a target image sequence of the target object.
[0019] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0020] Figure 1 An exemplary system architecture to which the video generation method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.
[0021] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the video generation method and apparatus may be applied may include a terminal device, but the terminal device may implement the video generation method and apparatus provided by the embodiments of the present disclosure without interacting with a server.
[0022] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0023] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0024] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0025] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports the content browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0026] It should be noted that the video generation method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the video generation apparatus provided in the embodiment of the present disclosure can also be provided in the terminal device 101, 102, or 103.
[0027] Alternatively, the video generation method provided in the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the video generation apparatus provided in the embodiment of the present disclosure may generally be provided in the server 105. The video generation method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the video generation apparatus provided in the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0028] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0029] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0030] Figure 2 The flowchart of the video generation method according to the embodiment of the present disclosure is schematically shown.
[0031] like Figure 2 As shown, the method includes operations S210 to S240.
[0032] In operation S210 , a lip shape feature sequence corresponding to a speech feature sequence is determined, wherein the speech feature sequence is extracted from speech data of a video to be generated.
[0033] In operation S220, a target template lip movement parameter sequence is determined from a plurality of template lip movement parameters based on the lip shape feature sequence, wherein the plurality of template lip movement parameters are extracted from a video including real lip movements.
[0034] In operation S230, an initial image sequence of the target object is generated based on the target template lip movement parameter sequence.
[0035] In operation S240 , a temporal rendering process is performed on the initial image sequence of the target object to obtain a target image sequence of the target object.
[0036] According to an embodiment of the present disclosure, the voice data of the video to be generated can be segmented in chronological order into multiple monosyllabic voice data, and then voice features are extracted from each monosyllabic voice data of the multiple monosyllabic voice data to generate multiple voice features arranged in chronological order, that is, a voice feature sequence.
[0037] According to embodiments of the present disclosure, a neural network model can be used to predict a lip shape feature sequence based on a speech feature sequence. For example, an RNN (Recurrent Neural Network) model can be used to predict a lip shape feature sequence based on a speech feature sequence. However, this is not limiting. A neural network model can also be used to predict lip shape features based on speech features, forming a lip shape feature sequence from multiple lip shape features divided in chronological order.
[0038] According to other embodiments of the present disclosure, a matching rule may be pre-established to determine a lip shape feature sequence corresponding to a speech feature sequence. The matching rule may include a mapping rule between speech features and lip shape features.
[0039] According to an embodiment of the present disclosure, at least one speech feature in the speech feature sequence may include a phonetic posterior gram (PPG), but is not limited thereto and may also include a phoneme feature, a syllable feature, and a Mel Frequency Cepstrum Coefficient (MFCC), etc.
[0040] According to embodiments of the present disclosure, the a posteriori probability of speech is used to represent the a posteriori probability of the speech category for each specific time frame of speech data. The speech category corresponds to the factor state. The a posteriori probability of speech is extracted using a speaker-independent automatic speech recognition system. Using the a posteriori probability of speech as a speech feature can avoid the reduction in lip shape feature prediction accuracy caused by speaker-related factors acting as noise when predicting lip shape feature sequences.
[0041] According to an embodiment of the present disclosure, at least one lip shape feature in the lip shape feature sequence may include a lip shape thumbnail, but is not limited thereto. The lip shape feature may also include a regressed lip shape point, a lip shape blend deformation parameter (BlendShape), and the like.
[0042] According to an embodiment of the present disclosure, the lip shape feature sequence may include a lip change feature sequence, which is used to characterize the movement changes of the lips in a time sequence during speaking.
[0043] According to an embodiment of the present disclosure, at least one template lip movement parameter in the template lip movement parameter sequence may include a lip shape thumbnail, which may also be referred to as a lip shape micro-image. However, this is not limited to this. Template lip movement parameters may also include regression lip shape points, lip shape blending parameters, and the like.
[0044] According to an embodiment of the present disclosure, the template lip movement parameter sequence may include a lip change parameter sequence, which is used to characterize the movement changes of the lips in a time sequence during the speaking process.
[0045] According to an embodiment of the present disclosure, a target template lip movement parameter sequence can be determined from multiple template lip movement parameters based on a lip shape feature sequence, and an initial image sequence of a target object can be generated based on the target template lip movement parameter sequence. For example, the target template lip movement parameter sequence can be combined with the target object's image parameter information and rendered to generate the target object's initial image sequence. However, this is not limiting. Alternatively, a lip shape feature sequence can be used to generate an initial image sequence of a target object. For example, the lip shape feature sequence can be combined with the target object's image parameter information and rendered to generate the target object's initial image sequence.
[0046] According to an embodiment of the present disclosure, at least one template lip movement parameter in a template lip movement parameter sequence can be extracted from a video including real lip movements. For example, it can be extracted from a video of a target subject speaking. Therefore, compared to using a lip shape feature sequence to generate an initial image sequence of the target subject, determining a target template lip movement parameter sequence from multiple template lip movement parameters based on a lip shape feature sequence, and generating an initial image sequence of the target subject based on the target template lip movement parameter sequence, can make the lip movements of the target subject's initial image more realistic and accurate, avoiding problems such as missing or distorted details.
[0047] It should be noted that the acquisition of the video including real lip movements involved in the embodiments of the present disclosure and the video including real expressions involved below is authorized by the user corresponding to the video.
[0048] According to embodiments of the present disclosure, a video including the target object's lip movements can be generated based on the target object's initial image sequence. However, this is not limiting. Alternatively, the target object's initial image sequence can be subjected to temporal rendering to obtain a target image sequence of the target object, and a video including the target object's lip movements can be generated based on the target image sequence of the target object.
[0049] According to embodiments of the present disclosure, temporal rendering is performed on the initial image sequence of a target object, and the temporal relevance of the initial image sequence can be exploited to re-render the initial image sequence. For example, when re-rendering the current initial image, the temporal relevance of the initial image of the next frame and the initial image of the previous frame can be combined, thereby improving the naturalness and authenticity of the rendered target image sequence of the target object.
[0050] According to an embodiment of the present disclosure, after operation S240 , the video generation method may further include an operation of generating a target video based on a target image sequence of a target object and voice data of a video to be generated.
[0051] According to an embodiment of the present disclosure, the target video can be played on the device terminal, the microphone plays the voice data, and the display screen shows the target object who is speaking and whose lip movements match the voice data.
[0052] By using the video generation method provided by the embodiment of the present disclosure, the target template lip movement parameter sequence can be used to generate the initial image sequence of the target object, so that the lip movements matching the voice data are real and vivid, avoiding problems such as missing or distortion. At the same time, the initial image sequence is processed by time-series rendering to make the target image of the target object more natural and realistic, thereby making the interaction between people and machines more intelligent and vivid.
[0053] Figure 3 The flowchart of a video generation method according to another embodiment of the present disclosure is schematically shown.
[0054] like Figure 3 As shown, the method includes operations S310 to S360.
[0055] In operation S310, a target template lip movement parameter sequence is determined.
[0056] In operation S320 , a target template expression parameter sequence is determined.
[0057] In operation S330, image parameter information of the target object is determined.
[0058] In operation S340 , the target template expression parameter sequence and the target template lip movement parameter sequence are fused to obtain an expression-fused lip movement parameter sequence.
[0059] In operation S350, an initial image sequence of the target object is generated based on the expression-fused lip movement parameter sequence and the image parameter information.
[0060] In operation S360 , a temporal rendering process is performed on the initial image sequence of the target object to obtain a target image sequence of the target object.
[0061] According to an embodiment of the present disclosure, the target template lip movement parameter sequence may include a lip change parameter sequence that matches the speech data of the video to be generated.
[0062] According to an embodiment of the present disclosure, with respect to operation S310 , determining a target template lip movement parameter sequence may include the following operations.
[0063] For example, a target template lip shape feature that matches the lip shape feature is determined from multiple template lip shape features. Based on the target template lip shape feature, the target template lip movement parameters are determined from multiple template lip movement parameters in the lip shape mapping relationship. The lip shape mapping relationship represents the mapping relationship between multiple template lip movement parameters and multiple template lip shape features.
[0064] According to an embodiment of the present disclosure, a target template lip-shape feature that matches a lip-shape feature is determined from multiple template lip-shape features. The lip-shape feature can be matched one-to-one with each of the multiple template lip-shape features, and the similarity between the lip-shape feature and each of the multiple template lip-shape features is calculated to obtain multiple similarity results corresponding to the multiple template lip-shape features. The multiple similarity results are sorted in descending order, and the template lip-shape feature that ranks first is used as the target template lip-shape feature. The similarity calculation method is not limited, and for example, it can be Euclidean similarity or cosine similarity.
[0065] According to an embodiment of the present disclosure, a lip shape mapping relationship between multiple template lip movement parameters and multiple template lip shape features can be pre-established, so that the target template lip movement parameters can be obtained based on the target template lip shape features by matching the lip shape mapping relationship.
[0066] According to an embodiment of the present disclosure, for each lip shape feature in the lip shape feature sequence, a target template lip movement parameter that matches the lip shape feature can be determined in the above manner. Multiple target template lip movement parameters are arranged in a time sequence to obtain a target template lip movement parameter sequence that matches the lip shape feature sequence.
[0067] According to the embodiments of the present disclosure, the target template lip movement parameters are closer to the actual lip shape dynamic parameters than the lip shape features directly predicted from speech features. The target template lip movement parameters can be determined from multiple pre-generated template lip movement parameters using an indexing approach based on a lip movement mapping relationship. However, this is not a limitation. Alternatively, the lip shape features can be directly optimized and adjusted to obtain target lip movement parameters that are closer to the actual lip shape.
[0068] According to an embodiment of the present disclosure, determining the target template lip movement parameters by indexing is simpler, more convenient and more effective than obtaining the target lip movement parameters by optimizing and adjusting the calculation method.
[0069] According to an embodiment of the present disclosure, the video generation method may further include, for example, an operation of generating a plurality of template lip movement parameters.
[0070] For example, a video frame sequence of the subject's lip movements is obtained, and a plurality of template lip movement parameters are determined based on the video frame sequence of the subject's lip movements.
[0071] According to an embodiment of the present disclosure, the object may be a target object, but is not limited thereto. The object may also be an object of the same type as the target object.
[0072] According to embodiments of the present disclosure, a video of a subject expressing, for example, a segment of speech data can be captured. This video can be deframed to obtain a sequence of video frames of the subject's lip movements. Based on the sequence of video frames of the subject's lip movements, multiple template lip movement parameters can be determined.
[0073] According to an embodiment of the present disclosure, a plurality of template lip movement parameters are determined using a lip movement video frame sequence of a real object, which is more realistic and effective and avoids the problem of distortion.
[0074] According to other embodiments of the present disclosure, lip movement video frames can be preprocessed, such as by orienting objects in the video frames or aligning objects in multiple video frames. Lip movement parameters can be extracted from each lip movement video frame in the preprocessed lip movement video frame sequence to obtain template lip movement parameters. By repeating the above process for different sets of speech data, multiple template lip movement parameters can be determined.
[0075] By preprocessing the lip movement video frames in the embodiment of the present disclosure, the problem of large changes in lip shape in different postures can be avoided, so that the template lip movement parameters are all obtained under the same lip shape posture, thereby improving the standardization of the template lip movement parameters.
[0076] According to embodiments of the present disclosure, the target template expression parameter sequence may include a sequence of facial variation parameters, excluding the lips, that matches the speech data of the video to be generated. The face may include, but is not limited to, the area below the lips, such as the chin. The face may also include areas above the lips, such as the eyebrows, eyes, or cheeks. The target template expression parameter sequence can be used to characterize the temporal changes in facial movements, excluding the lips, during speech.
[0077] According to an embodiment of the present disclosure, operation S320 may be performed in the following manner to determine the target template expression parameter sequence.
[0078] For example, target expression type information is determined. Based on the target expression type information, a target template expression parameter sequence is determined from multiple template expression parameters in an expression mapping relationship, where the expression mapping relationship represents a mapping relationship between multiple template expression parameters and multiple expression type information, and the multiple template expression parameters are extracted from a video including real expressions.
[0079] According to an embodiment of the present disclosure, the target expression type information may include expression tags involving facial movements, such as opening the mouth, laughing, smiling, crying, sadness, raising eyebrows, etc.
[0080] According to an embodiment of the present disclosure, the target expression type information can be determined based on the voice data of the video to be generated. For example, semantic information of the voice data of the video to be generated can be identified and the target expression type information can be determined based on the semantic information. However, this is not limited to this. The target expression type information can also be determined based on expression type information specified by the user.
[0081] According to an embodiment of the present disclosure, the video generation method may further include the following operations.
[0082] For example, the expression video frame sequence of the object is determined from the lip movement video frame sequence of the object. Based on the expression video frame sequence of the object, a plurality of template expression parameters are determined.
[0083] According to embodiments of the present disclosure, an expression video frame sequence of an object can be determined from a lip movement video frame sequence of the object. For example, an expression video frame sequence in which expression changes are detected can be selected from the lip movement video frame sequence. However, this is not limited to this. An expression video frame sequence can also be extracted from a video containing the object's real expression. Any method that can extract an expression video frame sequence is sufficient.
[0084] According to an embodiment of the present disclosure, an expression video frame sequence of an object is determined from a lip movement video frame sequence of the object. The lip movement video frame sequence may be a lip movement video frame sequence that has undergone pre-processing operations such as straightening or aligning the object, thereby improving the accuracy and authenticity of the template expression parameters.
[0085] According to embodiments of the present disclosure, expression-fused lip movement parameter sequences can be used to represent parameter sequences of facial changes, such as lips, cheeks, eyebrows, and eyes, during speech. For example, expression-fused lip movement parameter sequences can be used to represent parameter sequences of facial changes that occur during speech, such as eyebrow raising, blinking, and smiling, which combine lip shape and facial expressions.
[0086] According to an embodiment of the present disclosure, for operation S340, the target template expression parameter sequence and the target template lip movement parameter sequence are fused to obtain an expression-fused lip movement parameter sequence, which may include the following operations.
[0087] For example, based on the duration and initial occurrence time of the target template expression parameter sequence, the target template expression parameter sequence is superimposed on the target template lip movement parameter sequence, for example, the target template lip movement parameter sequence is interpolated, smoothed, etc., and finally an expression-fused lip movement parameter sequence is obtained.
[0088] According to other embodiments of the present disclosure, when the target subject is emotionally stable or has a neutral expression, only the lip shape changes dynamically while the target subject is speaking, and no dynamic changes in the expression occur, the video generation method may include operations S310, S330, S350, and S360. The operations may be determined based on actual circumstances.
[0089] According to the embodiments of the present disclosure, the target image of the target object is generated by combining the target template expression parameter sequence, so that the target image of the target object can be more vivid and natural in the process of, for example, speaking and singing.
[0090] According to an embodiment of the present disclosure, with respect to operation S350 , generating an initial image sequence of the target object based on the expression-fused lip movement parameter sequence and image parameter information may include the following operations.
[0091] For example, the image rendering network model is used to process the lip movement parameter sequence and image parameter information of expression fusion to generate the initial image sequence of the target object.
[0092] According to an embodiment of the present disclosure, the image parameter information may include at least one of the following: head posture information, facial mesh information, facial texture information, and lighting parameter information, etc.
[0093] According to the embodiments of the present disclosure, the structure of the image rendering network model is not limited. For example, it can be a generative adversarial network model, but it is not limited to this. It can also be a deep learning network model. Any network model can generate an initial image sequence of the target object based on the lip movement parameter sequence and image parameter information of the expression fusion.
[0094] According to the embodiments of the present disclosure, the initial image sequence of the target object is based on the lip movement parameter sequence and image parameter information of the expression fusion, so that the target object can control the expression, lip shape, posture, etc. during the speaking process, thereby improving the naturalness and interactivity of the initial image of the target object.
[0095] According to an embodiment of the present disclosure, with respect to operation S360 , performing a temporal rendering process on the initial image sequence of the target object to obtain the target image sequence of the target object may include the following operations.
[0096] For example, a temporal rendering model can be used to perform temporal rendering on the initial image sequence of the target object, resulting in a target image sequence of the target object. This model can combine temporal correlations within the initial image sequence, as well as the continuity of facial movements, to re-render the initial image sequence, taking the initial image sequence (i.e., multiple temporally consecutive video frames) as input and leveraging the temporal correlations between the frames. This results in a more realistic and natural target image sequence.
[0097] According to other embodiments of the present disclosure, a second rendering model that is the same as the image rendering network model can be used to render multiple initial images of the target object, such as multiple initial images that are not temporally related, to obtain multiple target images of the target object. The second rendering model can also be used to compensate for the realism of the initial image, to obtain a target image sequence that is more realistic than the initial image sequence. However, compared to the second rendering model, the time domain rendering model provided by the embodiment of the present disclosure uses multiple temporally related initial images, such as an initial image sequence, as input to perform time domain rendering processing, so that the generated target image sequence is more realistic, vivid, and natural.
[0098] According to embodiments of the present disclosure, the temporal rendering model may include a Unet network, but is not limited thereto. A Unet-HD (High-Resolution) network may also be used. Any network structure that can take multiple temporally consecutive video frames, such as an initial image sequence, as input and a target image sequence as output, and that can exploit the temporal correlation within the initial image sequence, will suffice.
[0099] According to an embodiment of the present disclosure, an initial time-domain rendering model may be trained using training samples to obtain a time-domain rendering model.
[0100] According to an embodiment of the present disclosure, a training sample may include a sample video recorded with a real object, such as a real person, as the target object. Sample voice data and a sample video frame sequence corresponding to the sample voice data may be extracted from the sample video. A sample voice feature sequence is extracted from the sample voice data. A sample lip shape feature sequence corresponding to the sample voice feature sequence is determined. Based on the sample lip shape feature sequence, a target template lip movement parameter sequence is determined from a plurality of template lip movement parameters. Based on the target template lip movement parameter sequence, an initial sample image sequence of the target object is generated. The initial sample image sequence of the target object, i.e., a plurality of, for example, 8 temporally continuous sample images, is input into the initial time domain rendering model to obtain a predicted sample image sequence of the target object. The sample video frame sequence and the predicted sample image sequence are input into the loss function to obtain a loss value. The parameters in the initial time domain rendering model are adjusted until the loss value converges. The model when the loss value converges is used as the time domain rendering model.
[0101] According to the embodiments of the present disclosure, the loss function is not limited. For example, the cross entropy loss function can be used, but it is not limited to this. The loss function can also be adjusted according to the network structure of the initial time domain rendering model, as long as it can achieve the training of the initial time domain rendering model.
[0102] According to the embodiments of the present disclosure, by using training samples with temporal continuity to train the initial time-domain rendering model, the time-domain rendering model can learn the temporal correlation in multiple temporally continuous initial image sequences through training, and then use the time-domain rendering model to render the initial image sequence, so that the target image sequence of the target object obtained is closer to the image of the real person.
[0103] Figure 4 The block diagram schematically shows a video generating device according to an embodiment of the present disclosure.
[0104] like Figure 4 As shown, the video generating device 400 may include a feature determination module 410 , a lip movement determination module 420 , a generation module 430 and a time domain rendering module 440 .
[0105] The feature determination module 410 is configured to determine a lip shape feature sequence corresponding to a speech feature sequence, wherein the speech feature sequence is extracted from speech data of a video to be generated.
[0106] The lip movement determination module 420 is used to determine a target template lip movement parameter sequence from multiple template lip movement parameters based on the lip shape feature sequence, wherein the multiple template lip movement parameters are extracted from a video including real lip movements.
[0107] The generation module 430 is used to generate an initial image sequence of the target object based on the lip movement parameter sequence of the target template.
[0108] The time domain rendering module 440 is configured to perform time domain rendering processing on the initial image sequence of the target object to obtain a target image sequence of the target object.
[0109] According to an embodiment of the present disclosure, the video generating apparatus may further include a type determination module and an expression determination module.
[0110] The type determination module is used to determine the target expression type information based on the voice data of the video to be generated.
[0111] The expression determination module is used to determine a target template expression parameter sequence from multiple template expression parameters in an expression mapping relationship based on target expression type information, wherein the expression mapping relationship represents the mapping relationship between multiple template expression parameters and multiple expression type information, and the multiple template expression parameters are extracted from a video including real expressions.
[0112] According to an embodiment of the present disclosure, the generation module may include an image determination unit, a fusion unit, and a generation unit.
[0113] The image determination unit is used to determine the image parameter information of the target object, wherein the image parameter information includes at least one of the following: head posture information, facial mesh information, facial texture information, and lighting parameter information.
[0114] The fusion unit is used to fuse the target template expression parameter sequence and the target template lip movement parameter sequence to obtain an expression fused lip movement parameter sequence.
[0115] The generating unit is used to generate an initial image sequence of the target object based on the expression-fused lip movement parameter sequence and image parameter information.
[0116] According to an embodiment of the present disclosure, the lip movement determination module may include a matching unit and a lip shape determination unit.
[0117] The matching unit is used to determine, for each lip feature in the lip feature sequence, a target template lip feature that matches the lip feature from multiple template lip features.
[0118] The lip shape determination unit is used to determine a target template lip movement parameter sequence from multiple template lip movement parameters in a lip shape mapping relationship based on the target template lip shape feature, wherein the lip shape mapping relationship represents a mapping relationship between multiple template lip movement parameters and multiple template lip shape features.
[0119] According to an embodiment of the present disclosure, the video generating apparatus may further include an acquisition module and a lip movement extraction module.
[0120] The acquisition module is used to acquire a lip movement video frame sequence of the object.
[0121] The lip movement extraction module is used to determine multiple template lip movement parameters based on the lip movement video frame sequence of the object.
[0122] According to an embodiment of the present disclosure, the video generating apparatus may further include a frame extraction module and an expression extraction module.
[0123] The frame extraction module is used to determine the object's expression video frame sequence from the object's lip movement video frame sequence.
[0124] The expression extraction module is used to determine multiple template expression parameters based on the expression video frame sequence of the object.
[0125] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0126] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a method as in the embodiment of the present disclosure.
[0127] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute a method according to an embodiment of the present disclosure.
[0128] According to an embodiment of the present disclosure, a computer program product includes a computer program. When the computer program is executed by a processor, the method according to the embodiment of the present disclosure is implemented.
[0129] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0130] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0131] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0132] The computing unit 501 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the video generation method by any other appropriate means (e.g., by means of firmware).
[0133] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0135] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0136] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0137] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0138] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0139] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0140] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A video generation method, comprising: Determining a lip shape feature sequence corresponding to a speech feature sequence, wherein the speech feature sequence is extracted from speech data of a video to be generated; Based on the lip shape feature sequence, determining a target template lip movement parameter sequence from a plurality of template lip movement parameters, wherein the plurality of template lip movement parameters are extracted from a video including real lip movements of the target subject; Determining image parameter information of the target object, wherein the image parameter information includes at least one of the following: head posture information, facial mesh information, facial texture information, and lighting parameter information; fusing a target template expression parameter sequence and the target template lip movement parameter sequence to obtain an expression-fused lip movement parameter sequence, wherein the target template expression parameter sequence is determined from a plurality of template expression parameters extracted from a video including real expressions; generating an initial image sequence of the target object based on the expression-fused lip movement parameter sequence and the image parameter information; and The initial image sequence of the target object is subjected to time domain rendering processing to obtain a target image sequence of the target object.
2. The method according to claim 1, further comprising: Determining target expression type information based on the voice data of the video to be generated; as well as Based on the target expression type information, a target template expression parameter sequence is determined from a plurality of template expression parameters in an expression mapping relationship, wherein the expression mapping relationship represents a mapping relationship between the plurality of template expression parameters and a plurality of expression type information.
3. The method according to claim 1, wherein Determining a target template lip movement parameter sequence from a plurality of template lip movement parameters based on the lip shape feature sequence includes: For each lip-shaped feature in the lip-shaped feature sequence, determining a target template lip-shaped feature that matches the lip-shaped feature from a plurality of template lip-shaped features; and Based on the target template lip shape features, the target template lip movement parameters are determined from the multiple template lip movement parameters in the lip shape mapping relationship, wherein the lip shape mapping relationship represents the mapping relationship between the multiple template lip movement parameters and the multiple template lip shape features.
4. The method according to any one of claims 1 to 3, further comprising: Obtain a sequence of lip movement video frames of the subject; as well as The plurality of template lip movement parameters are determined based on a sequence of lip movement video frames of the subject.
5. The method according to claim 4, further comprising: determining an expression video frame sequence of the subject from a lip movement video frame sequence of the subject; as well as The plurality of template expression parameters are determined based on the expression video frame sequence of the object.
6. A video generation device comprising: a feature determination module, configured to determine a lip shape feature sequence corresponding to a speech feature sequence, wherein the speech feature sequence is extracted from speech data of a video to be generated; a lip movement determination module, configured to determine a target template lip movement parameter sequence from a plurality of template lip movement parameters based on the lip shape feature sequence, wherein the plurality of template lip movement parameters are extracted from a video including real lip movements of a target subject; a generating module, configured to generate an initial image sequence of the target object based on the target template lip movement parameter sequence; and A time domain rendering module, configured to perform time domain rendering processing on the initial image sequence of the target object to obtain a target image sequence of the target object; Wherein, the generation module includes: An image determination unit, configured to determine image parameter information of the target object, wherein the image parameter information includes at least one of the following: head posture information, facial mesh information, facial texture information, and lighting parameter information; a fusion unit, configured to fuse a target template expression parameter sequence and the target template lip movement parameter sequence to obtain an expression-fused lip movement parameter sequence, wherein the target template expression parameter sequence is determined from a plurality of template expression parameters; and A generating unit is used to generate an initial image sequence of the target object based on the expression-fused lip movement parameter sequence and the image parameter information.
7. The apparatus according to claim 6, further comprising: A type determination module, configured to determine target expression type information based on the voice data of the video to be generated; as well as An expression determination module is used to determine a target template expression parameter sequence from multiple template expression parameters in an expression mapping relationship based on the target expression type information, wherein the expression mapping relationship represents a mapping relationship between the multiple template expression parameters and multiple expression type information, and the multiple template expression parameters are extracted from a video including real expressions.
8. The device according to claim 6, wherein The lip movement determination module includes: a matching unit, configured to determine, for each lip feature in the lip feature sequence, a target template lip feature that matches the lip feature from a plurality of template lip features; and A lip shape determination unit is used to determine the target template lip movement parameters from the multiple template lip movement parameters in the lip shape mapping relationship based on the target template lip shape features, wherein the lip shape mapping relationship represents the mapping relationship between the multiple template lip movement parameters and the multiple template lip shape features.
9. The apparatus according to any one of claims 6 to 8, further comprising: An acquisition module, used to acquire a lip movement video frame sequence of the subject; as well as The lip movement extraction module is used to determine the multiple template lip movement parameters based on the lip movement video frame sequence of the object.
10. The apparatus according to claim 9, further comprising: A frame extraction module, configured to determine a facial expression video frame sequence of the subject from a lip movement video frame sequence of the subject; as well as The expression extraction module is used to determine the multiple template expression parameters based on the expression video frame sequence of the object.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Voice-driven virtual face video generation method and device
CN110874557A