Video generation method, and deep learning model training method and device
By dividing the input speech into sub-voices and generating a sequence of key points, the problem of multi-person digital human lip-drive automation in the prior art is solved, and efficient and automated multi-object dialogue video generation is achieved.
Patent Information
- Application Number
- CN202510338889.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to realize automated multi-person digital human lip drive, resulting in high cost and inefficiency.
By dividing the input speech into multiple sub-voices, and determining the key point sequence based on the speech characteristics of the sub-voice and the template characteristics of the object, generating the target video, and automatically generating the multi-object dialogue video.
Without manpower segmentation and splicing videos, we automatically drive the lip shape changes of each object in multi-object videos, improve the generation efficiency of multi-object dialogue videos, and save manpower and time costs.
Smart Images

Figure CN120220212A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technical fields such as computer vision and augmented reality, and can be applied to scenarios such as digital humans. More specifically, the present disclosure provides a video generation method, a training method for a deep learning model, a device, an electronic device, a storage medium, and a computer program product. Background Art
[0002] With the rapid development of artificial intelligence, the technology of voice-driven video generation has also received extensive attention, especially the technology of voice-driven lip movement has been more widely used. Summary of the Invention
[0003] The present disclosure provides a video generation method, a training method for a deep learning model, a device, an electronic device, a storage medium, and a computer program product.
[0004] According to a first aspect, a video generation method is provided. The method includes: dividing an input voice into a plurality of sub-voices according to a plurality of pronunciation objects and the pronunciation order of the plurality of pronunciation objects; for each sub-voice, determining a key point sequence of the object to which the sub-voice belongs according to the voice feature of the sub-voice and the template feature of the object to which the sub-voice belongs, where the key point sequence represents the lip movement of the object to which the sub-voice belongs when making the sub-voice; and generating a target video according to the key point sequences of the objects to which the plurality of sub-voices belong.
[0005] According to a second aspect, a training method for a deep learning model is provided. The method includes: extracting the voice feature of a sample voice and the template feature of the object to which the sample voice belongs from a first sample video; inputting the voice feature and the template feature into the deep learning model to obtain an output key point sequence, where the output key point sequence represents the lip movement of the object to which the sample voice belongs; and adjusting the parameters of the deep learning model according to the difference between the output key point sequence and a reference key point sequence, where the reference key point sequence is obtained by extracting the key points of the object to which the sample voice belongs from the first sample video.
[0006] According to a third aspect, a training method for a deep learning model is provided. The method includes: obtaining the facial key points of a plurality of objects in each image frame of a second sample video; associating the facial key points of the same object in different image frames according to the facial features of the plurality of objects in each image frame to obtain a key point sequence of each object, where the key point sequence represents the lip movement of the object; for each object, inputting the key point sequence of the object into the deep learning model to obtain an output facial image sequence of the object, and determining the loss of the object according to the difference between the output facial image sequence of the object and the original facial image sequence of the object in the second sample video; and adjusting the parameters of the deep learning model according to the losses of the plurality of objects respectively.
[0007] According to a fourth aspect, a video generation device is provided. The device includes: a voice division module configured to divide an input voice into a plurality of sub-voices according to a plurality of pronunciation objects and the pronunciation order of the plurality of pronunciation objects; a first key point sequence determination module configured to, for each sub-voice, determine a key point sequence of the object to which the sub-voice belongs according to the voice feature of the sub-voice and the template feature of the object to which the sub-voice belongs, where the key point sequence represents the lip shape change of the object to which the sub-voice belongs when pronouncing the sub-voice; and a target image determination module configured to generate a target video according to the key point sequences of the objects to which the plurality of sub-voices belong respectively.
[0008] According to a fifth aspect, a training device for a deep learning model is provided. The device includes: an extraction module configured to extract the voice feature of a sample voice and the template feature of the object to which the sample voice belongs from a first sample video; a second key point sequence determination module configured to input the voice feature and the template feature into the deep learning model to obtain an output key point sequence, where the output key point sequence represents the lip shape change of the object to which the sample voice belongs; and a first adjustment module configured to adjust the parameters of the deep learning model according to the difference between the output key point sequence and a reference key point sequence, where the reference key point sequence is obtained by extracting the key points of the object to which the sample voice belongs from the first sample video.
[0009] According to a sixth aspect, a training device for a deep learning model is provided. The device includes: an acquisition module configured to acquire the facial key points of a plurality of objects in each image frame of a second sample video; a third key point sequence determination module configured to associate the facial key points of the same object in different image frames according to the facial features of the plurality of objects in each image frame to obtain a key point sequence of each object, where the key point sequence represents the lip shape change of the object; a loss determination module configured to, for each object, input the key point sequence of the object into the deep learning model to obtain an output facial image sequence of the object, and determine the loss of the object according to the difference between the output facial image sequence of the object and the original facial image sequence of the object in the second sample video; and a second adjustment module configured to adjust the parameters of the deep learning model according to the losses of the plurality of objects respectively.
[0010] According to a seventh aspect, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided according to the present disclosure.
[0011] According to an eighth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method provided according to the present disclosure.
[0012] According to a ninth aspect, there is provided a computer program product including a computer program stored on at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the method provided by the present disclosure.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0015] Figure 1 is a schematic diagram of an exemplary system architecture to which a video generation method and a deep learning model training method according to an embodiment of the present disclosure can be applied;
[0016] Figure 2 is a flowchart of a video generation method according to an embodiment of the present disclosure;
[0017] Figure 3 is a schematic diagram of a video generation method according to an embodiment of the present disclosure;
[0018] Figure 4 is a flowchart of a deep learning model training method according to an embodiment of the present disclosure;
[0019] Figure 5 is a schematic diagram of a deep learning model outputting a key point sequence according to an embodiment of the present disclosure;
[0020] Figure 6 is a flowchart of a deep learning model training method according to an embodiment of the present disclosure;
[0021] Figure 7 is a schematic diagram of a deep learning model generating an image sequence according to an embodiment of the present disclosure;
[0022] Figure 8 is a block diagram of a video generation device according to an embodiment of the present disclosure;
[0023] Figure 9 is a block diagram of a deep learning model training device according to an embodiment of the present disclosure;
[0024] Figure 10 is a block diagram of a deep learning model training device according to an embodiment of the present disclosure; and
[0025] Figure 11A block diagram of an electronic device that is at least one of a video generation method and a deep learning model training method according to an embodiment of the present disclosure. Detailed implementation manners
[0026] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0027] Currently, lip driving methods are mainly applied to single-person videos. For multi-person scenarios, to achieve the function of multi-person driving, generally, the video is first cut into multiple single-person videos. For example, each image frame in the video is truncated, and the truncated image only contains a single person. Then, the single-person videos are processed, model trained, and driven to generate new single-person videos. Finally, the single-person videos generated by driving are spliced into a video. This solution requires manual human effort to cut and splice the video, and cannot achieve automated multi-person digital human lip driving, with high costs and low efficiency.
[0028] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0029] In the technical solution of the present disclosure, before obtaining or collecting user personal information, the authorization or consent of the user is obtained.
[0030] Figure 1 It is a schematic diagram of an exemplary system architecture that can apply the video generation method and the deep learning model training method according to an embodiment of the present disclosure. It should be noted that Figure 1 The illustration shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.
[0031] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0032] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices, including but not limited to smartphones, tablets, laptop computers, and so on.
[0033] Server 105 can be a server that provides various services. For example, it can be a back-end management server (only as an example) that supports websites browsed by users using terminal devices 101, 102, and 103. The back-end management server can analyze and process data such as user requests received, and feedback the processing results to the terminal devices.
[0034] At least one of the video generation method and the deep learning model training method provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the video generation device and the deep learning model training device provided by the embodiments of the present disclosure can generally be set in server 105.
[0035] Figure 2 It is a flowchart of a video generation method according to an embodiment of the present disclosure.
[0036] As Figure 2 shown, the video generation method 200 includes operations S210 to S230.
[0037] In operation S210, the input speech is divided into multiple sub-speeches according to multiple pronunciation objects and the pronunciation order of the multiple pronunciation objects.
[0038] The input speech can be the conversation of multiple objects in the video that the user wants to generate. According to the pronunciation order of the multiple objects, the input speech can be divided into multiple sub-speeches, and each sub-speech corresponds to one object. For example, the first sub-speech is the speech of the first object, the second sub-speech is the speech of the second object, the third sub-speech is the speech of the third object, or the third sub-speech is again the speech of the first object, and so on.
[0039] In operation S220, for each sub-speech, according to the speech feature of the sub-speech and the template feature of the object to which the sub-speech belongs, a key point sequence of the object to which the sub-speech belongs is determined.
[0040] For example, the key point sequence characterizes the lip shape change of the object to which the sub-speech belongs when making the sub-speech.
[0041] For each sub - speech, feature extraction can be performed on the sub - speech to obtain the speech features of the sub - speech. For example, the Wav2Vec (Wave to Vector, the conversion from audio waveform to vector) model can be used as a speech feature extractor. By inputting the sub - speech into this speech feature extractor, speech features representing lip shape changes can be obtained.
[0042] For each sub - speech, an image of the object to which the sub - speech belongs can be obtained, and template features can be determined based on the image of the object to which the sub - speech belongs. For example, the template features can be obtained by extracting the facial key points of the object from the image of the object to which the sub - speech belongs and performing feature representation on the facial key points. Thus, the template features can represent the facial contour of the object that emits the sub - speech.
[0043] For each sub - speech, based on the speech features and template features of the sub - speech, a key - point sequence belonging to the sub - speech can be generated, and this key - point sequence represents the facial contour and lip shape changes of the object that emits the sub - speech.
[0044] For example, the key - point sequence includes multi - frame key - point information. Each frame of key - point information includes multiple key points on the object's face. The positions of the lip key points in the key - point information of different frames have relative changes, and this relative change represents the lip shape changes of the object. That is to say, the key - point sequence represents the facial contour and lip shape changes of the object that emits the sub - speech.
[0045] In operation S230, a target video is generated based on the key - point sequences of the objects to which multiple sub - speeches belong respectively.
[0046] For example, the key - point sequence of each sub - speech can generate a target image sequence of the object to which the sub - speech belongs. This target image sequence contains the facial picture information of the object to which the sub - speech belongs and the lip shape change information of the object to which the sub - speech belongs when emitting the sub - speech. The target image sequences of multiple sub - speeches are spliced together in the order of generation of the sub - speeches, and a target video of multiple objects having a conversation can be obtained.
[0047] Embodiments of the present disclosure divide the input speech into multiple sub - speeches according to the pronunciation objects and pronunciation order. For each sub - speech, based on the speech features of the sub - speech and the template features of the object, a key - point sequence of the object to which the sub - speech belongs is generated. A target video of multiple objects having a conversation is obtained based on the key - point sequences of the objects to which multiple sub - speeches belong respectively, which can automatically drive the lip shape changes of each object in the multi - object video using speech and improve the generation efficiency of the multi - object conversation video.
[0048] According to the embodiments of the present disclosure, there is no need for manual video segmentation and splicing, and the lip shape changes of each object in the multi - object video are automatically driven to realize the generation of the multi - object conversation video, saving labor costs and time costs.
[0049] Figure 3 It is a schematic diagram of a video generation method according to an embodiment of the present disclosure.
[0050] As Figure 3 shown, this embodiment includes a key point generation model 320 and an image generation model 330. The key point generation model 320 is used to generate a key point sequence of the object to which each sub-speech belongs. The image generation model 330 is used to generate a target image sequence of the object to which each sub-speech belongs.
[0051] The input speech 310 can be a conversation of multiple objects in the video that the user wants to generate. According to the pronunciation order of the multiple objects, the input speech can be divided into multiple sub-speeches, and each sub-speech corresponds to one object. For example, the sub-speech 311 is the speech uttered by the first object, and the sub-speech 312 is the sub-speech uttered by the second object.
[0052] According to an embodiment of the present disclosure, determining the key point sequence of the object to which the sub-speech belongs includes: extracting the speech feature of the sub-speech; determining the template feature of the object to which the sub-speech belongs according to the template image of the object to which the sub-speech belongs in a predetermined state; and determining the key point sequence of the object to which the sub-speech belongs according to the speech feature of the sub-speech and the template feature of the object to which the sub-speech belongs. Determining the template feature of the object to which the sub-speech belongs includes: detecting the facial key points of the object to which the sub-speech belongs in the template image; extracting the features of the facial key points to obtain the template feature.
[0053] For example, for the sub-speech 311, the Wav2Vec model can be used to extract the speech feature. And the image of the object to which the sub-speech 311 belongs in the closed-mouth state can be determined as the template image. Then, the facial key points of the object to which the sub-speech 311 belongs in the template image can be detected, and the facial key points can be characterized to obtain the template feature. Next, the speech feature and the template feature can be input into the key point generation model 320, and the key point generation model 320 outputs the key point sequence of the object to which the sub-speech 311 belongs. This key point sequence characterizes the lip movement of the object to which the sub-speech 311 belongs when uttering the sub-speech 311.
[0054] According to an embodiment of the present disclosure, the length of the sub-speech corresponds to multiple image frames. The steps for the key point generation model 320 to generate the key point sequence of the object to which the sub-speech belongs include: expanding the length of the template feature to be the same as the length of the speech feature of the sub-speech, where the length of the speech feature is determined according to the number of frames of the image frames corresponding to the sub-speech; fusing the expanded template feature and the speech feature to obtain a fused feature; and generating a key point sequence according to the fused feature.
[0055] In one example, the video that the user wants to generate may have an original video with original speech. The user wants to replace the original speech with input speech 310 and make the lip movements of each object in the original video conform to the input speech 310.
[0056] In this scenario, the number of frames of the image frames corresponding to each sub-speech in the input speech 310 can be determined, and the length of the speech feature of each sub-speech is determined by the number of frames of the image frames corresponding to the sub-speech. For example, sub-speech 311 corresponds to N image frames, and the speech feature of sub-speech 311 can be expressed as (N, K), where N is the number of frames and K is the dimension of the speech feature.
[0057] In the above scenario, the template image of the object to which sub-speech 311 belongs in the closed-mouth state can be detected from the original video. Specifically, the key points of the upper and lower lips of the object to which sub-speech 311 belongs can be detected, and the closed-mouth state can be determined according to the positional relationship of the key points of the upper and lower lips. For example, when the distance between the key points of the upper and lower lips of the object to which sub-speech 311 belongs is less than a certain threshold, it is determined that the object to which sub-speech 311 belongs is in the closed-mouth state.
[0058] After determining the template image of the object to which sub-speech 311 belongs in the closed-mouth state, the facial key points of the object in the template image can be detected, and the facial key points are input into a fully connected layer to obtain a template feature with the same dimension as the above speech feature. The template feature can be expressed as (1, K). Subsequently, the template feature can be copied N times to extend the length of the template feature, so that the size of the extended template feature is the same as the size of the speech feature, for example, both are (N, K). Next, the extended template feature and the speech feature are input into a fully connected layer for feature fusion, and the fused feature is input into the key point generation model 320 to obtain the key point sequence of the object to which sub-speech 311 belongs. The key point generation model 320 can be a neural network model including a Transform Decoder module. The fused feature is input into the Transform Decoder, and the key point sequence is output.
[0059] The key point sequence can include N frames of key point information, and each frame of key point information includes the coordinates of multiple key points on the face of the object to which sub-speech 311 belongs, especially including the coordinates of multiple key points on the mouth of the object to which sub-speech 311 belongs. Multiple frames of key points can represent the facial contour and lip movement changes of the object that emits sub-speech 311.
[0060] According to an embodiment of the present disclosure, the generation of the target video 340 includes: for each sub-voice, generating a target image sequence of the object to which the sub-voice belongs according to the key point sequence of the object to which the sub-voice belongs; and generating a target video according to the target image sequences of the objects to which the multiple sub-voices belong. Among them, the generation of the target image sequence includes: visualizing the key point sequence of the object to which the sub-voice belongs to obtain a key point image sequence of the object to which the sub-voice belongs; and generating a target image sequence of the object to which the sub-voice belongs according to the key point image sequence of the object to which the sub-voice belongs.
[0061] For example, after the key point generation model 320 outputs the key point sequence of the object to which the sub-voice 311 belongs, the key point sequence can be visualized into an image according to the coordinates to obtain a key point image sequence. Then, the key point image sequence is input into the image generation model 330, and the image generation model 330 generates a target image sequence 331 of the object to which the sub-voice 311 belongs based on the key point image sequence. The target image sequence 331 includes the facial picture information of the object to which the sub-voice 311 belongs and the lip movement change information of the object to which the sub-voice 311 belongs when uttering the sub-voice 311.
[0062] For the sub-voice 312, similar steps to those for the sub-voice 311 can be taken to obtain the key point sequence of the object to which the sub-voice 312 belongs and the target image sequence 332. The target image sequence 332 includes the facial picture information of the object to which the sub-voice 312 belongs and the lip movement change information of the object to which the sub-voice 312 belongs when uttering the sub-voice 312.
[0063] Finally, the target image sequence 331 and the target image sequence 332 can be spliced together in the pronunciation order of the sub-voice 311 and the sub-voice 312 to obtain the target video 340. The target video 340 can show a scene where multiple objects are having a conversation.
[0064] In the case of having an original video, the facial region videos of the objects in the target video can be replaced with the target videos of the respective objects, so as to achieve the effect of editing the original video. Compared with the original video, the voice in the edited video is the input voice 310, and the lip movement changes in the facial regions of the objects in the video conform to the input voice 310.
[0065] According to an embodiment of the present disclosure, by generating a key point sequence of the object to which a sub-voice belongs for different pronunciation objects based on voice features and template features, visualizing the key point sequence into a key point image sequence, generating a target image sequence of the object to which the sub-voice belongs according to the key point image sequence, and then generating a target video according to the target image sequences of the respective sub-voices, it is possible to automatically generate a multi-object conversation video and improve the video generation efficiency.
[0066] Figure 4It is a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure.
[0067] As Figure 4 shown, the training method 400 of the deep learning model includes operations S410 to S430. The deep learning model may be the above key point generation model.
[0068] In operation S410, extract the speech features of the sample speech in the first sample video and the template features of the object to which the sample speech belongs.
[0069] The first sample video may be a single-object video or a multi-object video. In the case where the first sample video is a single-object video, the sample speech in the first sample video is the speech of the object in the first sample video. Feature extraction can be performed on the sample speech to obtain speech features. And an image frame in which the object is in a closed-mouth state can be detected from the first sample video as a template image, and then facial key points of the object can be detected from the template image, and feature extraction is performed on the facial key points to obtain template features.
[0070] In the case where the first sample video is a multi-object video, it is necessary to distinguish the speeches of different objects. For example, for each image frame in the first sample video, detect the facial key points of each object in the image frame, intercept the facial image of each object according to the facial key points, and then determine the same object in different image frames according to the facial image. Thus, each object can be distinguished from each image frame of the first sample video, and further the speeches of each object can be distinguished.
[0071] For each object, feature extraction can be performed on the sample speech of the object in the first sample video to obtain speech features. Then an image frame in which the object is in a closed-mouth state can be determined from the first sample video as a template image. Then facial key point detection can be performed on the template image to obtain facial key points, and the features of the facial key points are extracted to obtain template features.
[0072] In operation S420, input the speech features and the template features into the deep learning model to obtain an output key point sequence.
[0073] For example, the output key point sequence characterizes the lip shape changes of the object that emits the sample voice. The sample voice corresponds to N frames of images, and the voice features of the sample voice can be represented as (N, K), where N is the number of frames and K is the dimension of the voice features. The template features can be extended to be consistent with the dimension of the voice features, and then the extended template features and the voice features are input into the deep learning model to obtain the key point sequence output by the deep learning model. The key point sequence can include N frames of key points, and each frame of key points includes the facial key point coordinates and lip key point coordinates of the object to which the sample voice belongs. The N frames of key points characterize the facial contour and lip changes of the object that emits the sample voice.
[0074] In operation S430, according to the difference between the output key point sequence and the reference key point sequence, the parameters of the deep learning model are adjusted.
[0075] The reference key point sequence is obtained by extracting the key points of the object to which the sample voice belongs from the first sample video. For example, in the case where the first sample video is a single-object video, facial key point detection can be performed on the first sample video to obtain the reference key point information of the object to which the sample voice belongs in the first sample video. In the case where the first sample video is a multi-object video, the same object in different image frames can be determined based on the facial features of each object, and then the key points of multiple frames belonging to the same object are associated to obtain the reference key point sequence of each object.
[0076] For each object, the difference between the key point sequence of the object output by the deep learning model and the reference key point sequence of the object can be calculated to obtain the loss of the object. The overall loss is determined according to the sum of the losses of multiple objects, and the parameters of the deep learning model are adjusted according to the overall loss.
[0077] Embodiments of the present disclosure extract the voice features of the sample voice and the template features of the object to which the sample voice belongs, input the voice features and the template features into the deep learning model to obtain the output key point sequence, use the reference key point sequence extracted from the first sample video as supervision, determine the loss according to the difference between the output key point sequence and the reference key point sequence, and adjust the model parameters, so that the trained deep learning model can generate a key point sequence characterizing the lip shape changes of the object based on voice driving.
[0078] Figure 5 It is a schematic diagram of the output key point sequence of the deep learning model according to an embodiment of the present disclosure.
[0079] As Figure 5 shown, this embodiment includes a voice feature extraction model 510 and a key point generation model 520.
[0080] According to an embodiment of the present disclosure, determining the template features of the object to which the sample voice belongs includes: determining a template image in which the object to which the sample voice belongs is in a predetermined state from a first sample video; detecting facial key points of the object to which the sample voice belongs in the template image; and extracting the features of the facial key points to obtain template features.
[0081] The sample voice 501 is the voice in the first sample video. When the first sample video is a single-object video, the object in the first sample video is the object to which the sample voice 501 belongs. An image frame in which the object to which the sample voice belongs is in the closed-mouth state can be determined from the first sample video, facial key points are detected to obtain facial key points, and feature extraction is performed on the facial key points to obtain template features 502.
[0082] When the first sample video is a multi-object video, first, the object-related parts belonging to the same object among the multiple image frames of the first sample video are associated. Then, for the same object, the sample voice of the object and the image frame in which the object is in the closed-mouth state are determined from the first sample video, key point detection is performed on the image frame in which the object is in the closed-mouth state to obtain facial key points, and feature extraction is performed on the facial key points to obtain the template features 502 of the object.
[0083] The sample voice 501 is input into a voice feature extraction model 510 to obtain voice features characterizing lip shape changes. Then, the voice features and the template features 502 are input into a key point generation model 520 to obtain an output key point sequence 503.
[0084] According to an embodiment of the present disclosure, the key point generation model 520 generating the output key point sequence includes: expanding the length of the template features according to the number of image frames corresponding to the sample voice in the first sample video; fusing the expanded template features and the voice features to obtain fused features; and inputting the fused features into a deep learning model to obtain the output key point sequence.
[0085] For example, the sample voice corresponds to N frames of images. The voice features of the sample voice can be expressed as (N, K), where N is the number of frames and K is the dimension of the voice features. The template features 502 can be obtained by inputting the facial key points into a fully connected layer for feature extraction, and the template features can be expressed as (1, K). Subsequently, the template features 502 can be copied N times to obtain expanded template features, expressed as (N, K). The expanded template features and the voice features are input into a fully connected layer for feature fusion to obtain fused features, and the fused features are input into the key point generation model 520 to obtain the output key point sequence 503. The key point sequence can include N frames of key points, and each frame of key points includes the facial key point coordinates and lip key point coordinates of the object to which the sample voice belongs. The N frames of key points characterize the facial contour and lip changes of the object that emits the sample voice.
[0086] The key point generation model 520 can be a neural network model including a Transformer decoder. The fused features can be input into the Transformer decoder, and a key point sequence can be obtained through decoding by the Transformer decoder.
[0087] In the embodiments of the present disclosure, by inputting the speech features of the sample speech and the template features of the object to which the sample speech belongs into the key point generation model, the key point generation model can generate a key point sequence representing the facial contour and lip shape changes of the object to which the speech belongs based on speech driving.
[0088] Figure 6 It is a flowchart of a training method of a deep learning model according to an embodiment of the present disclosure.
[0089] As Figure 6 shown, the training method 600 of the deep learning model includes operation S610 to operation S640. The deep learning model can be the above-mentioned image generation model.
[0090] In operation S610, facial key points of multiple objects in each image frame of the second sample video are obtained.
[0091] The second sample video can be a multi-object video. For each frame image in the second sample video, facial key point detection can be performed on each object in the image to obtain the facial key points of multiple objects in each frame image.
[0092] In operation S620, according to the facial features of multiple objects in each image frame, the facial key points belonging to the same object in different image frames are associated to obtain a key point sequence for each object.
[0093] For example, the key point sequence represents the facial contour and lip shape changes of the object. For each frame image, according to the facial key points of each object in the image, the original facial image of each object can be intercepted. Feature extraction is performed on the original facial image to obtain the facial features of each object.
[0094] Taking the first frame image as a reference, the similarity between the facial features in each subsequent frame image and the facial features in the first frame image can be calculated, and the facial features with a similarity greater than the threshold are determined to be the facial features of the same object. Thus, the same object in different frame images can be associated.
[0095] It is also possible to calculate the similarity between the facial features in each adjacent two frame images, and determine the facial features with a similarity greater than the threshold as the facial features of the same object. Thus, the same object in different frame images can also be associated.
[0096] Based on the same object associated in different image frames, the key points belonging to the same object can be associated to obtain the key point sequence of the object. For example, the second sample video includes N frames of images, and the key point sequence of each object can include N frames of key points, and the N frames of key points characterize the facial contour and lip shape change of the object.
[0097] In operation S630, for each object, input the key point sequence of the object into the deep learning model to obtain the output facial image sequence of the object, and determine the loss of the object according to the difference between the output facial image sequence of the object and the original facial image sequence of the object in the second sample video.
[0098] For example, for each object, the key point sequence of the object can be visualized as a key point image sequence, and then the key point image sequence is input into the deep learning model, and the deep learning model outputs the facial image sequence of the object, that is, the output facial image sequence. The output facial image sequence contains the facial picture information of the object and the lip shape change.
[0099] The original facial image sequence of each object can be extracted from the second sample video. For each object, the loss of the object can be determined according to the difference between the original facial image sequence of the object and the output facial image sequence. For example, the mean square error between the corresponding original facial image and the output facial image can be calculated as the loss between the original facial image and the output facial image. The loss of the object is determined according to the losses between multiple original facial images and output facial images.
[0100] For example, for each object, both the key point image sequence and the original facial image sequence of the object can be normalized. The normalization process can be to normalize each pixel in the image. Specifically, the pixel value of each pixel is divided by 255 and then subtracted by 1, so that the pixel value of each pixel is between [-1, 1]. The normalized key point image sequence can be input into the deep learning model, and the deep learning model outputs the normalized output image sequence. Then, according to the difference between the normalized original facial image sequence and the image sequence output by the model, the loss of the model is determined.
[0101] In operation S640, adjust the parameters of the deep learning model according to the losses of multiple objects respectively.
[0102] For example, the sum or weighted sum of the losses of multiple objects can be used as the overall loss of the deep learning model, and the parameters of the deep learning model are adjusted according to the overall loss to obtain the trained deep learning model.
[0103] Embodiments of the present disclosure distinguish different objects in a second sample video. For each object, the key point sequence of the object is input into a deep learning model to obtain an output image sequence of the object. The model parameters are adjusted according to the difference between the output image sequence and the original image sequence of the object, enabling the model to have the ability to generate a real image sequence driven by the key point sequence.
[0104] Figure 7 It is a schematic diagram of generating an image sequence by a deep learning model according to an embodiment of the present disclosure.
[0105] As Figure 7 shown, this embodiment includes a key point detection model 710, a feature comparison model 720, and an image generation model 730.
[0106] According to an embodiment of the present disclosure, according to the facial key points of multiple objects in each image frame, the original facial images of multiple objects in each image frame are extracted; and for each object, feature extraction is performed on the original facial image of the object in each image frame to obtain the facial features of the object in each image frame.
[0107] The second sample video 701 can be a video in which multiple objects have a conversation. For each frame image in the second sample video 701, the key point detection model 710 can be used to detect the facial key points of each object in the image, and the original facial images of each object in each frame image can be obtained according to the facial key points. For example, the original facial image 703 is the original facial image of the first object, and the original facial image 704 is the original facial image of the second object.
[0108] The original facial images of each object in the first frame image can be input into the feature comparison model 720 to obtain the facial features of each object, such as obtaining facial feature 1 and facial feature 2. For each subsequent frame image, the original facial images of each object in the image can be input into the same feature comparison model 720 to obtain the facial features of each object in the image, such as facial feature 3 and facial feature 4. Taking the first frame image as a reference, the feature comparison model 720 can match the facial features of the objects in each subsequent frame image with the facial features of the objects in the first frame image. For example, calculate the similarity between facial feature 3 and facial feature 1 and facial feature 2 respectively, and calculate the similarity between facial feature 4 and facial feature 1 and facial feature 2 respectively. For two facial features with a similarity greater than the threshold, they can be considered as the facial features of the same object. For example, facial feature 1 and facial feature 3 belong to the same object, and facial feature 2 and facial feature 4 belong to the same object. By analogy, the facial features belonging to the same object in different image frames can be obtained.
[0109] For the case where a new object appears in the middle image frame of the second sample video 701, the facial features of the object do not match those of each object in the first frame image. Then, this object can be regarded as a new object, and for each subsequent frame image, the comparison with the facial features of this new object is added.
[0110] By using the feature comparison model 720 to match the facial features of objects in different image frames, the same object in different image frames can be determined. Thus, the key points belonging to the same object in different image frames can be associated to obtain the key point sequence of each object. The key point sequence of each object can be expressed as (B, N, M, 2), where B represents the number of objects, N represents the number of frames, M represents the number of key points, and 2 represents the x and y coordinates of the key points. The key point sequence of each object can characterize the facial contour and lip shape change of the object.
[0111] According to an embodiment of the present disclosure, the image generation model 730 generates a facial image sequence of an object, including: performing visualization processing on the key point sequence of the object to obtain the key point image sequence of the object; and inputting the key point image sequence of the object into a deep learning model to obtain the output facial image sequence of the object.
[0112] For example, by performing visualization processing on the key point sequences of each object, the key point image sequences of each object can be obtained. For example, the key point image sequence 705 is the key point image sequence of the first object, and the key point image sequence 706 is the key point image sequence of the second object.
[0113] For the key point image sequence of each object, the key point image sequence can be input into the image generation model 730, and the image generation model 730 outputs the facial image sequence of the first object. For example, by inputting the key point image sequence 705 belonging to the first object into the image generation model 730, the image generation model 730 outputs the image sequence 707 of the first object. By inputting the key point image sequence 706 belonging to the second object into the image generation model 730, the image generation model 730 outputs the image sequence 708 of the second object.
[0114] The embodiment of the present disclosure enables the image generation model to drive the generation of the image sequence of the object based on the key point sequence of the object by inputting the key point image sequence into the image generation model.
[0115] According to an embodiment of the present disclosure, the present disclosure also provides a video generation device, a training device for a deep learning model, and another training device for a deep learning model.
[0116] Figure 8 It is a block diagram of a video generation device according to an embodiment of the present disclosure.
[0117] AsFigure 8 As shown, the video generation device 800 includes a voice division module 810, a first key point sequence determination module 820, and a video generation module 830.
[0118] The voice division module 810 is configured to divide the input voice into a plurality of sub - voices according to a plurality of pronunciation objects and the pronunciation order of the plurality of pronunciation objects.
[0119] The first key point sequence determination module 820 is configured to, for each sub - voice, determine the key point sequence of the object to which the sub - voice belongs according to the voice feature of the sub - voice and the template feature of the object to which the sub - voice belongs, where the key point sequence represents the lip - shape change of the object to which the sub - voice belongs when emitting the sub - voice.
[0120] The video generation module 830 is configured to generate a target video according to the key point sequences of the objects to which the plurality of sub - voices belong respectively.
[0121] According to an embodiment of the present disclosure, the first key point sequence determination module 820 includes a first template feature determination unit and a key point sequence determination unit.
[0122] The first template feature determination unit is configured to determine the template feature of the object to which the sub - voice belongs according to the template image of the object to which the sub - voice belongs in a predetermined state.
[0123] The key point sequence determination unit is configured to determine the key point sequence of the object to which the sub - voice belongs according to the voice feature of the sub - voice and the template feature of the object to which the sub - voice belongs.
[0124] The first template feature determination unit includes a detection subunit and a template feature determination subunit.
[0125] The detection subunit is configured to detect the facial key points of the object to which the sub - voice belongs in the template image.
[0126] The template feature determination subunit is configured to extract the features of the facial key points to obtain the template feature.
[0127] According to an embodiment of the present disclosure, the length of the sub - voice corresponds to a plurality of image frames. The key point sequence determination unit includes an extension subunit, a fusion subunit, and a key point sequence determination subunit.
[0128] The extension subunit is configured to extend the length of the template feature to be the same as the length of the voice feature of the sub - voice, where the length of the voice feature is determined according to the number of frames of the image frames corresponding to the sub - voice.
[0129] The fusion subunit is configured to fuse the extended template feature and the voice feature to obtain a fusion feature.
[0130] The key point sequence determination subunit is configured to generate a key point sequence according to the fusion feature.
[0131] The video generation module 830 includes a target image sequence generation unit and a target video generation unit.
[0132] The target image sequence generation unit is configured to generate, for each sub-speech, a target image sequence of the object to which the sub-speech belongs according to the key point sequence of the object to which the sub-speech belongs.
[0133] The target video generation unit is configured to generate a target video according to the target image sequences of the objects to which multiple sub-speeches belong respectively.
[0134] The target image sequence generation unit includes a key point image sequence generation subunit and a target image sequence generation subunit.
[0135] The key point image sequence generation subunit is configured to perform visualization processing on the key point sequence of the object to which the sub-speech belongs to obtain a key point image sequence of the object to which the sub-speech belongs.
[0136] The target image sequence generation subunit is configured to generate a target image sequence of the object to which the sub-speech belongs according to the key point image sequence of the object to which the sub-speech belongs.
[0137] Figure 9 It is a block diagram of a training device for a deep learning model according to an embodiment of the present disclosure.
[0138] As Figure 9 shown, the training device 900 for the deep learning model includes an extraction module 910, a second key point sequence determination module 920, and a first adjustment module 930.
[0139] The extraction module 910 is configured to extract the speech feature of the sample speech and the template feature of the object to which the sample speech belongs in the first sample video.
[0140] The second key point sequence determination module 920 is configured to input the speech feature and the template feature into the deep learning model to obtain an output key point sequence, and the output key point sequence characterizes the lip shape change of the object to which the sample speech belongs.
[0141] The first adjustment module 930 is configured to adjust the parameters of the deep learning model according to the difference between the output key point sequence and the reference key point sequence, where the reference key point sequence is obtained by extracting the key points of the object to which the sample speech belongs from the first sample video.
[0142] The extraction module 910 includes a template image determination unit, a detection unit, and a second template feature determination unit.
[0143] The template image determination unit is configured to determine a template image of the object to which the sample speech belongs in a predetermined state from the first sample video.
[0144] The detection unit is used to detect the facial key points of the object to which the sample speech belongs in the template image.
[0145] The second template feature determination unit is used to extract the features of the facial key points to obtain the template features.
[0146] The second key point sequence determination module 920 includes an expansion unit, a fusion unit, and an output unit.
[0147] The expansion unit is used to expand the length of the template features according to the number of image frames corresponding to the sample speech in the first sample video.
[0148] The fusion unit is used to fuse the expanded template features and the speech features to obtain the fusion features.
[0149] The output unit is used to input the fusion features into the deep learning model to obtain the output key point sequence.
[0150] Figure 10 It is a block diagram of a training device for a deep learning model according to another embodiment of the present disclosure.
[0151] As Figure 10 shown, the training device 1000 of the deep learning model includes an acquisition module 1001, a third key point sequence determination module 1002, a loss determination module 1003, and a second adjustment module 1004.
[0152] The acquisition module 1001 is used to acquire the facial key points of multiple objects in each image frame of the second sample video.
[0153] The third key point sequence determination module 1002 is used to associate the facial key points of the same object in different image frames according to the facial features of multiple objects in each image frame, to obtain the key point sequence of each object, and the key point sequence represents the lip shape change of the object.
[0154] The loss determination module 1003 is used to input the key point sequence of the object into the deep learning model for each object, to obtain the output facial image sequence of the object, and to determine the loss of the object according to the difference between the output facial image sequence of the object and the original facial image sequence of the object in the second sample video.
[0155] The second adjustment module 1004 is used to adjust the parameters of the deep learning model according to the respective losses of multiple objects.
[0156] The third key point sequence determination module 1002 includes a visualization processing unit and a facial image sequence determination unit.
[0157] The visualization processing unit is used to perform visualization processing on the key point sequence of the object to obtain the key point image sequence of the object.
[0158] The facial image sequence determination unit is configured to input the key point image sequence of an object into a deep learning model to obtain the output facial image sequence of the object.
[0159] The training device 1000 of the deep learning model further includes an original facial image determination module and a facial feature determination module.
[0160] The original facial image determination module is configured to extract the original facial images of multiple objects in each image frame according to the facial key points of the multiple objects in each image frame; and
[0161] The facial feature determination module is configured to perform feature extraction on the original facial image of an object in each image frame for each object to obtain the facial features of the object in each image frame.
[0162] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0163] Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0164] As Figure 11 shown, the device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1102 or the computer program loaded from the storage unit 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.
[0165] Multiple components in device 1100 are connected to I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0166] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as at least one of the video generation method and the training method of the deep learning model. For example, in some embodiments, at least one of the video generation method and the training method of the deep learning model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of at least one of the video generation method and the training method of the deep learning model described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute at least one of the video generation method and the training method of the deep learning model in any other suitable manner (e.g., by means of firmware).
[0167] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0168] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0169] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0170] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0171] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0172] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other.
[0173] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0174] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A video generation method, comprising: Dividing the input speech into a plurality of sub-speech according to a plurality of pronunciation objects and the pronunciation order of the plurality of pronunciation objects; For each sub-speech, determining a key point sequence of the object to which the sub-speech belongs according to the speech features of the sub-speech and the template features of the object to which the sub-speech belongs, wherein the key point sequence represents the lip shape changes of the object to which the sub-speech belongs when the sub-speech is uttered; as well as A target video is generated according to key point sequences of the objects to which the multiple sub-voices respectively belong.
2. The method according to claim 1, wherein: The step of determining, for each sub-speech, a key point sequence of the object to which the sub-speech belongs according to the speech feature of the sub-speech and the template feature of the object to which the sub-speech belongs comprises: for each sub-speech, Determining the template features of the object to which the sub-speech belongs according to the template image of the object to which the sub-speech belongs in a predetermined state; and A key point sequence of the object to which the sub-speech belongs is determined according to the speech features of the sub-speech and the template features of the object to which the sub-speech belongs.
3. The method according to claim 2, wherein: Determining the template features of the object to which the sub-speech belongs according to the template image of the object to which the sub-speech belongs is in a predetermined state comprises: Detecting facial key points of the object to which the sub-speech belongs in the template image; and The features of the facial key points are extracted to obtain the template features.
4. The method according to claim 2 or 3, wherein: The length of the sub-speech corresponds to a plurality of image frames; and determining the key point sequence of the object to which the sub-speech belongs according to the speech feature of the sub-speech and the template feature of the object to which the sub-speech belongs comprises: Extending the length of the template feature to be consistent with the length of the voice feature of the sub-speech, wherein the length of the voice feature is determined according to the number of image frames corresponding to the sub-speech; Fusing the expanded template feature with the speech feature to obtain a fused feature; and The key point sequence is generated according to the fusion features.
5. The method according to claim 1, wherein: Generating a target video according to the key point sequences of the objects to which the plurality of sub-voices respectively belong comprises: For each sub-speech, generating a target image sequence of the object to which the sub-speech belongs according to a key point sequence of the object to which the sub-speech belongs; and The target video is generated according to the target image sequence of the object to which the multiple sub-voices respectively belong.
6. The method according to claim 5, wherein: For each sub-speech, generating a target image sequence of the object to which the sub-speech belongs according to a key point sequence of the object to which the sub-speech belongs comprises: for each sub-speech, Visualizing a key point sequence of the object to which the sub-speech belongs to obtain a key point image sequence of the object to which the sub-speech belongs; and A target image sequence of the object to which the sub-speech belongs is generated according to the key point image sequence of the object to which the sub-speech belongs.
7. A method for training a deep learning model, comprising: Extracting speech features of a sample speech in a first sample video and template features of an object to which the sample speech belongs; Inputting the speech feature and the template feature into a deep learning model to obtain an output key point sequence, wherein the output key point sequence represents the lip shape change of the object to which the sample speech belongs; as well as Adjusting the parameters of the deep learning model according to the difference between the output key point sequence and a benchmark key point sequence, wherein the benchmark key point sequence is obtained by extracting the key points of the object to which the sample speech belongs from the first sample video.
8. The method according to claim 7, wherein: The step of extracting speech features of a sample speech in the first sample video and template features of an object to which the sample speech belongs includes: Determine, from the first sample video, a template image of the object to which the sample voice belongs in a predetermined state; Detecting facial key points of the object to which the sample voice belongs in the template image; and The features of the facial key points are extracted to obtain the template features.
9. The method according to claim 7 or 8, wherein: The step of inputting the speech features and the template features into a deep learning model to obtain an output key point sequence comprises: Extending the length of the template feature according to the number of image frames corresponding to the sample speech in the first sample video; Fusing the expanded template feature with the speech feature to obtain a fused feature; and The fused features are input into the deep learning model to obtain an output key point sequence.
10. A method for training a deep learning model, comprising: Obtaining facial key points of multiple objects in each image frame in the second sample video; According to the facial features of the multiple objects in each image frame, the facial key points belonging to the same object in different image frames are associated to obtain a key point sequence of each object, wherein the key point sequence represents the lip shape change of the object; For each object, inputting a key point sequence of the object into a deep learning model to obtain an output facial image sequence of the object, and determining a loss of the object based on a difference between the output facial image sequence of the object and an original facial image sequence of the object in the second sample video; as well as Adjusting parameters of the deep learning model based on the losses of each of the multiple objects.
11. The method according to claim 10, wherein: For each object, inputting the key point sequence of the object into the deep learning model to obtain the output facial image sequence of the object comprises: for each object, Performing visualization processing on the key point sequence of the object to obtain a key point image sequence of the object; and The key point image sequence of the object is input into a deep learning model to obtain an output facial image sequence of the object.
12. The method according to claim 10, further comprising: Extracting original facial images of the multiple objects in each image frame according to facial key points of the multiple objects in each image frame; as well as For each object, feature extraction is performed on the original facial image of the object in each image frame to obtain the facial features of the object in each image frame.
13. A video generating device, comprising: A speech division module, used for dividing the input speech into a plurality of sub-speech according to a plurality of pronunciation objects and the pronunciation order of the plurality of pronunciation objects; A first key point sequence determination module is used to determine, for each sub-speech, a key point sequence of the object to which the sub-speech belongs according to the speech features of the sub-speech and the template features of the object to which the sub-speech belongs, wherein the key point sequence represents the lip shape changes of the object to which the sub-speech belongs when the sub-speech is uttered; as well as The video generation module is used to generate a target video according to the key point sequences of the objects to which the multiple sub-voices respectively belong.
14. A training device for a deep learning model, comprising: An extraction module, used to extract speech features of a sample speech in a first sample video and template features of an object to which the sample speech belongs; A second key point sequence determination module is used to input the speech feature and the template feature into a deep learning model to obtain an output key point sequence, wherein the output key point sequence represents the lip shape change of the object to which the sample speech belongs; as well as A first adjustment module is used to adjust the parameters of the deep learning model according to the difference between the output key point sequence and a benchmark key point sequence, wherein the benchmark key point sequence is obtained by extracting the key points of the object to which the sample speech belongs from the first sample video.
15. A training device for a deep learning model, comprising: An acquisition module, used for acquiring facial key points of multiple objects in each image frame in the second sample video; A third key point sequence determination module is used to associate facial key points belonging to the same object in different image frames according to the facial features of the multiple objects in each image frame to obtain a key point sequence for each object, wherein the key point sequence represents the lip shape change of the object; a loss determination module, configured to input, for each object, a key point sequence of the object into a deep learning model to obtain an output facial image sequence of the object, and determine the loss of the object according to a difference between the output facial image sequence of the object and an original facial image sequence of the object in the second sample video; as well as The second adjustment module is used to adjust the parameters of the deep learning model according to the losses of each of the multiple objects.
16. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 12.
18. A computer program product, comprising a computer program, wherein the computer program is stored on at least one of a readable storage medium and an electronic device, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.