A video translation method and system, electronic equipment and storage medium
By eliminating mouth movements and translating audio from the original video, and combining this with a target rendering model to generate a face-rendered video, the problems of insufficient audio-visual synchronization and inter-frame continuity in existing technologies are solved, thus improving the user's viewing experience.
Patent Information
- Application Number
- CN202410585232.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-11
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-05-11
AI Technical Summary
Existing video translation technologies are inadequate in terms of audio-visual synchronization and inter-frame continuity, resulting in poor viewing experience for users.
By removing mouth movements from the original video, a target closed-mouth video is obtained. The original audio is then translated, and the closed-mouth video is driven by the target rendering model to match the translated audio, generating a face-rendered video that ensures that mouth movements are synchronized with the audio.
It improves the audio-visual synchronization rate and inter-frame continuity of the translated video, enhancing the user viewing experience.
Smart Images

Figure CN118474441B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer technology, and particularly relates to a video translation method and system, an electronic device and a storage medium. BACKGROUND
[0002] Video translation refers to generating a translation audio conforming to a speaking tone of a target person and a target language type through voice tone replication and machine translation, given a target person video and the target language type, and combining a target person image in the input target person video and a driving module to complete mouth movement rendering, so as to finally obtain a translation video in the target language type. The video translation task needs to ensure the accuracy of the translation audio content, the matching degree of the target person expression and posture, and the audio-visual synchronization rate of the mouth movement. In the prior art, the audio-visual synchronization rate and the video frame continuity of the final translation video may be poor, resulting in poor user viewing effect. SUMMARY
[0003] The present disclosure provides a technical solution of a video translation method and system, an electronic device and a storage medium.
[0004] According to an aspect of the present disclosure, a video translation method is provided, including: performing mouth movement elimination on a target person in an original video to obtain a target closed-mouth video; performing audio translation on original audio corresponding to the original video to obtain a translation audio, wherein the original audio and the translation audio correspond to different language types; and driving the target closed-mouth video based on the translation audio by using a target rendering model to obtain a face rendering video, wherein the mouth movement of the target person in the face rendering video matches the translation audio.
[0005] In a possible implementation manner, the performing mouth movement elimination on the target person in the original video to obtain the target closed-mouth video includes: performing mouth movement elimination on the target person in the original video to obtain a first closed-mouth video; performing video reconstruction based on the first closed-mouth video and the original video to obtain a second closed-mouth video, wherein the second closed-mouth video has the same posture change as the original video; and performing fusion based on the second closed-mouth video and the original video to obtain the target closed-mouth video, wherein the target closed-mouth video has the same person identity information as the original video.
[0006] In a possible implementation, the mouth action elimination on the target person in the original video to obtain a first closed-mouth video includes: determining original 3DMM coefficients of the original video based on a 3DMM fitting algorithm, where the original 3DMM coefficients include original pose coefficients and original expression coefficients; replacing the original expression coefficients in the original 3DMM coefficients with preset neutral expression coefficients to obtain target 3DMM coefficients; and performing expression modification on the original video based on the target 3DMM coefficients to obtain the first closed-mouth video.
[0007] In a possible implementation, the video reconstruction based on the first closed-mouth video and the original video to obtain a second closed-mouth video includes: performing feature extraction on the original video to obtain 3D key point motion features of the original video; and performing video reconstruction on the first closed-mouth video based on the 3D key point motion features of the original video to obtain the second closed-mouth video.
[0008] In a possible implementation, the fusion based on the second closed-mouth video and the original video to obtain the target closed-mouth video includes: performing face analysis on the target person in the original video, and performing region extraction on the original video based on a face analysis result to obtain a first to-be-fused video, where the first to-be-fused video includes an upper half face region and a nose region of the target person; performing face analysis on the target person in the second closed-mouth video, and performing region extraction on the second closed-mouth video based on a face analysis result to obtain a second to-be-fused video, where the second to-be-fused video includes a lower half face region of the target person and does not include the nose region; and fusing the first to-be-fused video and the second to-be-fused video to obtain the target closed-mouth video.
[0009] In a possible implementation, the driving of the target closed-mouth video based on the translated audio by using a target rendering model to obtain a face rendered video includes: performing face region extraction on the target person in the target closed-mouth video to obtain a reference video and a to-be-rendered video; inputting the to-be-rendered video, the reference video, and the translated audio into the target rendering model to obtain the face rendered video output by the target rendering model.
[0010] In a possible implementation, the reference video includes a plurality of reference video frames, and the video to be rendered includes a plurality of video frames to be rendered, each of which does not include a lower half face region of the target person; the inputting the video to be rendered, the reference video, and the translated audio into the target rendering model to obtain the face rendering video output by the target rendering model includes: for any one of the video frames to be rendered, using the target rendering model, performing person identity information calibration on the video frame to be rendered and a reference video frame corresponding to the video frame to be rendered based on a cross-attention mechanism to obtain cross-attention visual features of the video frame to be rendered, wherein the video frame to be rendered is obtained by removing the lower half face region from the corresponding reference video frame; using the target rendering model, determining audio features of the video frame to be rendered from the translated audio, and using affine coefficients determined based on the audio features of the video frame to be rendered to deform the cross-attention visual features of the video frame to be rendered to obtain deformed visual features of the video frame to be rendered; using the target rendering model, performing rendering processing on the video frame to be rendered based on the audio features and the deformed visual features of the video frame to be rendered to obtain a rendered video frame corresponding to the video frame to be rendered; and based on the rendered video frame corresponding to each of the video frames to be rendered in the video to be rendered, obtaining the face rendering video.
[0011] In a possible implementation, the training data of the target rendering model includes a sample video to be rendered, a reference sample video, and a sample audio; and the training process of the target rendering model includes: using the target rendering model to determine a face rendering video corresponding to the sample video to be rendered; determining a target training loss based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio; and training the target rendering model based on the target training loss to obtain the trained target rendering model.
[0012] In a possible implementation, the reference sample video includes a plurality of reference sample video frames, and the sample video to be rendered includes a plurality of sample video frames to be rendered, and each sample video frame to be rendered does not include a lower half face region of a person; the determining, by using the target rendering model, of the face rendering video corresponding to the sample video to be rendered includes: for any one of the sample video frames to be rendered, performing, by using the target rendering model, identity information calibration on the sample video frame to be rendered and a first reference sample video frame corresponding to the sample video frame to be rendered based on a cross-attention mechanism to obtain cross-attention visual features of the sample video frame to be rendered, where the first reference sample video frame corresponding to the sample video frame to be rendered is any one of the plurality of reference sample video frames; determining, by using the target rendering model, audio features of the sample video frame to be rendered from the sample audio, and deforming, by using an affine coefficient determined based on the audio features of the sample video frame to be rendered, the cross-attention visual features of the sample video frame to be rendered to obtain deformed visual features of the sample video frame to be rendered; performing, by using the target rendering model, rendering processing on the sample video frame to be rendered based on the audio features and the deformed visual features of the sample video frame to be rendered to obtain a rendered video frame corresponding to the sample video frame to be rendered; and obtaining the face rendering video corresponding to the sample video to be rendered based on the rendered video frame corresponding to each of the sample video frames to be rendered in the sample video to be rendered.
[0013] In a possible implementation, the determining, based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio, of the target training loss includes: determining, based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio, of an initial training loss, where the initial training loss includes at least one of the following: a reconstruction loss, a first synchronization loss, a second synchronization loss, a discriminator loss, and a multi-scale loss; and determining, based on the initial training loss, of the target training loss.
[0014] In a possible implementation, the determining the initial training loss based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio includes: determining the reconstruction loss based on each face rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and each second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video, wherein each sample video frame to be rendered is obtained by removing a lower half face region from the corresponding second reference video frame; and / or determining the first synchronization loss based on the face rendering video corresponding to the sample video to be rendered and the sample audio; and / or determining the second synchronization loss based on each face rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and each second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video; and / or determining the discriminator loss based on the face rendering video corresponding to the sample video to be rendered and the reference sample video; and / or determining the multi-scale loss based on each face rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and each second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video.
[0015] In a possible implementation, the determining the target training loss based on the initial training loss includes: in a case where the initial training loss includes the reconstruction loss, the first synchronization loss, the second synchronization loss, the discriminator loss, and the multi-scale loss, performing weighted summation on the reconstruction loss, the first synchronization loss, the second synchronization loss, the discriminator loss, and the multi-scale loss to obtain the target training loss.
[0016] In a possible implementation, the method further includes: performing re-drawing processing on a fusion video of the target mute video and the face rendering video by using a stable diffusion model to obtain a target video.
[0017] In a possible implementation, the stable diffusion model includes a plurality of isomer blocks; and the performing re-drawing processing on the fusion video of the target mute video and the face rendering video by using the stable diffusion model to obtain the target video includes: inputting a preset noise timestamp and a feature vector corresponding to each video frame in the fusion video into each isomer block; performing noise forward diffusion and reverse diffusion on the fusion video based on the plurality of isomer blocks to obtain the target video.
[0018] According to an aspect of the present disclosure, a video translation system is provided, comprising: a mouth movement elimination model configured to eliminate mouth movement of a target person in an original video to obtain a target mute video; an audio translation model configured to perform audio translation on original audio corresponding to the original video to obtain translated audio, wherein the original audio and the translated audio correspond to different language types; and a target rendering model configured to drive the target mute video based on the translated audio to obtain a face rendering video, wherein mouth movement of the target person in the face rendering video matches the translated audio.
[0019] According to an aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory configured to store processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the above method.
[0020] According to an aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer program instructions, which, when executed by a processor, implement the above method.
[0021] In the embodiments of the present disclosure, mouth movement of a target person in an original video is eliminated to obtain a target mute video, so as to reduce the influence of mouth movement in the original video on a subsequent rendering process and improve audio-visual synchronization rate. Original audio corresponding to the original video is translated to obtain translated audio corresponding to different language types from the original audio. Then, a target rendering model is used to drive the target mute video based on the translated audio, so as to effectively change the language type of the original video and obtain a face rendering video in which mouth movement of the target person matches the translated audio.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure. Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0023] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the technical solutions of the present disclosure.
[0024] Figure 1 A flowchart of a video translation method according to an embodiment of the present disclosure is shown.
[0025] Figure 2 A schematic diagram of a video translation system according to an embodiment of the present disclosure is shown.
[0026] Figure 3 A schematic diagram of an audio translation model according to an embodiment of the present disclosure is shown.
[0027] Figure 4 A schematic diagram illustrating a mouth movement elimination model according to an embodiment of the present disclosure is shown.
[0028] Figure 5 A schematic diagram illustrating a target rendering model according to an embodiment of the present disclosure is shown.
[0029] Figure 6 A block diagram of a video translation system according to an embodiment of the present disclosure is shown.
[0030] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0031] Various exemplary embodiments, features and aspects of the present disclosure will be explained in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote like elements or features. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically noted.
[0032] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0033] The term "and / or" used herein only means an association relation of associated objects, and means that three relations can exist, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term "at least one" herein means any one of multiple or any combination of at least two of multiple, for example, at least one of A, B, and C can mean that any one or more elements selected from a set composed of A, B, and C are included.
[0034] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some examples, methods, means, elements and circuits that are well known to those skilled in the art are not described in detail in order to highlight the main idea of the present disclosure.
[0035] Video translation refers to generating a translated video in a target language type by voice tone replication and machine translation, given a target person video and the target language type, and combining the target person image in the input target person video and the driving module to complete the mouth movement rendering, so as to finally obtain the translated video in the target language type. The video translation task needs to ensure the accuracy of the translated audio content, the matching degree of the target person's expression and posture, and the audio-visual synchronization rate of the mouth movement. Video translation has a wide range of market application scenarios, such as video conversion between different language types, video conference translation conversion, video game multi-language type compatibility, film and television work dissemination, etc. In the security field, the development of counterattacks is limited to data acquisition, while the video translation task can provide a large amount of simulation data. More importantly, in the field of digital humans, the video translation task can fully combine natural language processing, computer vision, speech processing and other multi-modal cross-technology, and the research and development of the video translation task will greatly promote the development of identity digital humans and service digital humans. In the prior art, the audio-visual synchronization rate and the video frame continuity of the translated video obtained by the video translation task may be poor, resulting in poor user viewing effect.
[0036] The embodiments of the present disclosure provide a video translation method, which can improve the audio-visual synchronization rate and the video frame continuity of the target video obtained after video translation, thereby effectively improving the user viewing effect. The video translation method provided by the embodiments of the present disclosure is described in detail below.
[0037] Figure 1 A flowchart of a video translation method according to an embodiment of the present disclosure is shown. The method can be performed by an electronic device such as a terminal device or a server, and the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor invoking computer-readable instructions stored in a memory. Alternatively, the method can be performed by a server. As shown in Figure 1 The method includes:
[0038] In step S11, the mouth movement of the target person in the original video is eliminated to obtain a target closed-mouth video.
[0039] The original video here can be a video in which the target person speaks in a first language type. The specific type of the first language type can be determined according to actual conditions, and the present disclosure does not make specific limitations thereto.
[0040] The video translation is to change the language type of the target person in the original video. Different language types correspond to different mouth movements. The semantic information carried by the mouth movement of the target person in the original video will affect the driving effect of the subsequent translated audio based on different language types on the mouth movement. Therefore, the mouth movement of the target person in the original video is eliminated, so that the target closed-mouth video without carrying semantic information is effectively obtained.
[0041] In the following, the specific process of how to eliminate the mouth movement of the target person in the original video to obtain the target closed-mouth video will be described in detail in combination with possible implementation manners of the present disclosure, which will not be repeated here.
[0042] In step S12, the original audio corresponding to the original video is subjected to audio translation to obtain translated audio, wherein the original audio and the translated audio correspond to different language types.
[0043] The original audio of the first language type is converted into translated audio of the second language type, so as to effectively change the language type of the original video. The second language type can be set to any language type different from the first language type according to actual scene needs, which is not specifically limited in the present disclosure.
[0044] In the following, the specific process of how to translate the original audio corresponding to the original video to obtain the translated audio will be described in detail in combination with possible implementation manners of the present disclosure, which will not be repeated here.
[0045] In step S13, the target rendering model is used to drive the target closed-mouth video based on the translated audio to obtain a face rendering video, wherein the mouth movement of the target person in the face rendering video matches the translated audio.
[0046] The target rendering model is used to drive the mouth movement of the target person in the target closed-mouth video based on the translated audio, so as to effectively change the language type of the original video, and obtain the face rendering video in which the mouth movement of the target person matches the translated audio.
[0047] The face rendering video here can be a video frame sequence including only the face region of the target person and excluding the background region; or can be a video frame sequence in which the background region is restored and has the same size as the original video.
[0048] In the following, the training process of the target rendering model and the specific process of how to use the target rendering model to drive the target closed-mouth video based on the translated audio to obtain the face rendering video will be described in detail in combination with possible implementation manners of the present disclosure, which will not be repeated here.
[0049] In the embodiments of the present disclosure, the mouth movement of the target person in the original video is eliminated to obtain a target closed-mouth video, so as to reduce the influence of the mouth movement of the original video on the subsequent rendering process and improve the audio-visual synchronization rate. The original audio corresponding to the original video is translated to obtain translated audio of a different language type corresponding to the original audio. Then, the target rendering model is used to drive the target closed-mouth video based on the translated audio, so as to effectively change the language type of the original video and obtain a face rendering video in which the mouth movement of the target person matches the translated audio.
[0050] Figure 2 A schematic diagram of a video translation system according to an embodiment of the present disclosure is shown. As shown in Figure 2 The video translation system includes an audio translation model, which is configured to translate the original audio corresponding to the original video input into the video translation system to obtain translated audio.
[0051] In an example, the audio translation model of the embodiments of the present disclosure can adopt a mainstream audio translation processing flow in the related art, which is not specifically limited in the present disclosure.
[0052] Figure 3 A schematic diagram of an audio translation model according to an embodiment of the present disclosure is shown. As shown in Figure 3 The audio translation model performs the following audio translation processing flow: obtaining original audio by performing audio separation on the original video; obtaining original clean audio of a first language type (a language type corresponding to the original video) and original background music (BGM) noise by performing noise separation on the original audio; obtaining original text of the first language type by performing audio recognition on the original clean audio; obtaining target text of a second language type (a language type required for final video translation) by performing text translation on the original text; obtaining target clean audio of the second language type by performing text-to-speech (TTS) processing on the target text; performing timbre replication processing on the target clean audio based on the original clean audio, so that the target clean audio and the original clean audio have the same timbre, and fusing the target clean audio after timbre replication and the original BGM noise to obtain translated audio.
[0053] The above noise separation, audio recognition, TTS, timbre replication, and text translation can adopt processing manners in the related art, which are not specifically limited in the present disclosure.
[0054] As shown in Figure 2 The video translation system includes a mouth movement elimination model, which is configured to eliminate the mouth movement of a target person in the original video to obtain a target closed-mouth video.
[0055] In a possible implementation, the mouth action elimination is performed on the target person in the original video to obtain a target closed-mouth video, including: performing mouth action elimination on the target person in the original video to obtain a first closed-mouth video; performing video reconstruction based on the first closed-mouth video and the original video to obtain a second closed-mouth video, wherein the second closed-mouth video has the same pose change as the original video; and performing fusion based on the second closed-mouth video and the original video to obtain the target closed-mouth video, wherein the target closed-mouth video has the same person identity information as the original video.
[0056] The mouth action elimination is performed on the target person in the original video to obtain a first closed-mouth video in which the mouth action is preliminarily eliminated. To further optimize the closed-mouth effect, video reconstruction is performed on the first closed-mouth video to ensure that a second closed-mouth video has the same pose as the original video, and fusion processing is performed on the second closed-mouth video to ensure that the target closed-mouth video does not lose the person identity information of the original video.
[0057] Figure 4 A schematic diagram of a mouth action elimination model according to an embodiment of the present disclosure is shown. As shown in Figure 4 The mouth action elimination model includes an expression editing module configured to perform mouth action elimination on the target person in the original video to obtain a first closed-mouth video in which the mouth action is preliminarily eliminated.
[0058] The expression editing module can be a neural network model DNet capable of performing an expression modification task, or can be another module capable of performing an expression modification task, which is not limited in the present disclosure.
[0059] In a possible implementation, the mouth action elimination is performed on the target person in the original video to obtain a first closed-mouth video, including: determining original 3DMM coefficients of the original video based on a 3D Morphable Models (3DMM) fitting algorithm, wherein the original 3DMM coefficients include original pose coefficients and original expression coefficients; replacing the original expression coefficients in the original 3DMM coefficients with preset neutral expression coefficients to obtain target 3DMM coefficients; and performing expression modification on the original video based on the target 3DMM coefficients to obtain the first closed-mouth video.
[0060] The original 3DMM coefficients of the original video are extracted, and then only the original expression coefficients in the original 3DMM coefficients are replaced with preset neutral expression coefficients without changing the original pose coefficients to obtain target 3DMM coefficients including the original pose coefficients and the preset neutral expression coefficients; and the expression modification is performed on the original video by using the target 3DMM coefficients as a driving signal to obtain a first closed-mouth video in which the mouth action is preliminarily eliminated.
[0061] As Figure 4As shown, the original 3DMM coefficients of the original video are extracted, and the original expression coefficients are replaced based on the preset neutral expression coefficients to obtain target 3DMM coefficients. The target 3DMM coefficients are input into the expression editing module as driving signals.
[0062] In an example, when the expression editing module is a DNet model, the DNet model converts the target 3DMM coefficients as driving signals into hidden vectors through a mapping network, and then outputs the first closed-mouth video by an encoder-decoder-based feature warping and refinement network.
[0063] The first closed-mouth video after expression editing based on 3DMM coefficients may have problems such as inter-frame discontinuity, incorrect construction of the upper half of the face area, incomplete modeling of the lips, loss of character identity information, and changes in character posture. Therefore, the first closed-mouth video needs to be further optimized.
[0064] As shown in Figure 4 The mouth movement elimination model includes a video reconstruction module and a fusion module. The video reconstruction module is used to reconstruct a video based on the first closed-mouth video and the original video to obtain a second closed-mouth video with the same posture change as the original video. The fusion module is used to fuse the second closed-mouth video and the original video to obtain a target closed-mouth video with the same character identity information as the original video.
[0065] The video reconstruction module can be a neural network model Face-vid2vid capable of performing a video reconstruction task, or other modules capable of performing a video reconstruction task. The fusion module can be a module capable of performing a fusion task based on face analysis. The present disclosure does not make specific limitations thereto.
[0066] In a possible implementation, video reconstruction based on the first closed-mouth video and the original video to obtain the second closed-mouth video includes: performing feature extraction on the original video to obtain 3D key point motion features of the original video; and performing video reconstruction on the first closed-mouth video based on the 3D key point motion features of the original video to obtain the second closed-mouth video.
[0067] The 3D key point motion features of the original video are used to correct the first closed-mouth video, so as to effectively reconstruct the second closed-mouth video with high inter-frame continuity and the same posture change as the original video.
[0068] As shown in Figure 4 The original video and the first closed-mouth video are input into the video reconstruction module, and the video reconstruction module can output the corrected second closed-mouth video.
[0069] For example, the video reconstruction module (e.g., Face-vid2vid model) extracts the 3D key point motion features of the i-th video frame in the original video based on a motion encoder, and then the video reconstruction module takes the 3D key point motion features of the i-th video frame in the original video as a driving signal to perform video frame reconstruction on the i-th video frame in the first closed-mouth video, so as to obtain the i-th video frame in the second closed-mouth video, wherein the i-th video frame in the second closed-mouth video has the same pose information as the i-th video frame in the original video. Similarly, the above video frame reconstruction process is performed on each video frame in the first closed-mouth video to obtain the second closed-mouth video.
[0070] In a possible implementation, the target closed-mouth video is obtained by fusing the second closed-mouth video and the original video, including: performing face analysis on the target person in the original video, and performing region extraction on the original video based on the face analysis result to obtain a first to-be-fused video, wherein the first to-be-fused video includes the upper half face region and the nose region of the target person; performing face analysis on the target person in the second closed-mouth video, and performing region extraction on the second closed-mouth video based on the face analysis result to obtain a second to-be-fused video, wherein the second to-be-fused video includes the lower half face region of the target person and does not include the nose region; and fusing the first to-be-fused video and the second to-be-fused video to obtain the target closed-mouth video.
[0071] Considering that the chin closing state is usually shorter than when speaking, the upper half face region of the face usually contains more personal identity information, and the nose region is unnatural after the mouth movement is eliminated, the upper half face region, the nose region of the original video, and the lower half face region of the second closed-mouth video which does not include the nose region are fused based on face analysis technology, so as to effectively obtain the target closed-mouth video with closed lips, no chin movement, and no loss of personal identity information.
[0072] As shown in FIG. 7, the original video and the second closed-mouth video are input into the fusion module, and the fusion module can output the corrected target closed-mouth video. Figure 4
[0073] In an example, the face analysis is performed on the target person in the original video, and the region extraction is performed on the original video based on the face analysis result to obtain the first to-be-fused video, including: performing face analysis on the target person in the original video to obtain a first mask, wherein the first mask is used to indicate the upper half face region and the nose region of the target person; performing feathering operation on the first mask to obtain a second mask; and fusing the second mask with the original video to obtain the first to-be-fused video.
[0074] In one example, facial analysis is performed on the target person in the second closed-mouth video, and region extraction is performed on the second closed-mouth video based on the facial analysis results to obtain a second video to be fused. This includes: performing facial analysis on the target person in the second closed-mouth video to obtain a third mask, wherein the third mask is used to indicate the lower half of the target person's face and the nose region; performing a feathering operation on the third mask to obtain a fourth mask; and fusing the fourth mask with the second closed-mouth video to obtain the second video to be fused.
[0075] Feathering is a common image processing operation, primarily used to create masks with soft, natural edges. The principle of feathering is to control the blurring of the transition area between the selected region and the outside of the mask, creating a gradual transition and achieving a natural blend.
[0076] After obtaining the translated audio and the target closed-mouth video using the above method, a target rendering model can be used to drive the target closed-mouth video based on the translated audio to obtain a face rendering video.
[0077] In one possible implementation, a target rendering model is used to drive a target silent video based on translated audio to obtain a face-rendered video. This includes: extracting the face region of the target person in the target silent video to obtain a reference video and a video to be rendered; and inputting the video to be rendered, the reference video, and the translated audio into the target rendering model to obtain the face-rendered video output by the target rendering model.
[0078] The facial region of the target person in the target silent video is extracted to obtain a reference video containing only the facial region. Then, the lower half of the face region in the reference video is removed to obtain the video to be rendered. The video to be rendered, the reference video, and the translated audio are input into the target rendering model, and the face rendering video can be output end-to-end.
[0079] The video to be rendered does not include the lower half of the face, so as to remind the target rendering model to focus on rendering the lower half of the face.
[0080] like Figure 2 As shown, the video translation system includes: a target rendering model; the translated audio output from the audio translation model, which serves as the driving audio input to the target rendering model; and the target closed-mouth video output from the mouth motion elimination model, after face region extraction, with the resulting video to be rendered and reference video input to the target rendering model.
[0081] In a possible implementation, the reference video includes a plurality of reference video frames, the video to be rendered includes a plurality of video frames to be rendered, and each video frame to be rendered does not include a lower half face region of a target person; the target rendering model is input with the video to be rendered, the reference video, and the translated audio to obtain a face rendering video output by the target rendering model, including: for any one video frame to be rendered, using the target rendering model, performing person identity information calibration on the video frame to be rendered and a reference video frame corresponding to the video frame to be rendered based on a cross-attention mechanism to obtain cross-attention visual features of the video frame to be rendered, wherein the video frame to be rendered is obtained by removing the lower half face region from the corresponding reference video frame; using the target rendering model to determine audio features of the video frame to be rendered from the translated audio, and using an affine coefficient determined based on the audio features of the video frame to be rendered to deform the cross-attention visual features of the video frame to be rendered to obtain deformed visual features of the video frame to be rendered; using the target rendering model to perform rendering processing on the video frame to be rendered based on the audio features and the deformed visual features of the video frame to be rendered to obtain a rendered video frame corresponding to the video frame to be rendered; and obtaining the face rendering video based on the rendered video frame corresponding to each video frame to be rendered in the video to be rendered.
[0082] Compared with the method of randomly selecting one of the plurality of reference video frames as the reference video frame providing the person identity information in the related art, in the embodiment of the present disclosure, the target mute video is decoupled, that is, after the face region of any one video frame in the target mute video is extracted, the video frame can be used as the reference video frame providing the person identity information, and after the lower half face region of the reference video frame is removed, the video frame is used as the video frame to be rendered, so that the person identity information in the real video frame can be effectively utilized, and the change probability of the person identity information can be effectively reduced.
[0083] For example, after the face region extraction is performed on the target person in the target video, a reference video is obtained, and after the lower half face region is removed from each video frame in the reference video, a video frame to be rendered is obtained. That is, the i th video frame to be rendered in the video to be rendered is obtained by removing the lower half face region from the i th reference video frame in the reference video, that is, the i th reference video frame is a real video frame of the i th video frame to be rendered, and therefore, the i th reference video frame can be used as the reference video frame corresponding to the i th video frame to be rendered to provide effective person identity information.
[0084] For any given video frame to be rendered, using the target rendering model and based on the cross-attention mechanism, the character identity information of the video frame to be rendered and the corresponding reference video frame is calibrated. This allows cross-attention to be achieved between the lower half of the face region of the video frame to be rendered and the upper half of the face region of the reference video frame. This enables the rendering of the lower half of the face region of the video frame to be rendered to be prioritized, while maintaining the character identity information provided by the upper half of the face region of the reference video frame to a certain extent, thus obtaining the cross-attention visual features of the video frame to be rendered.
[0085] Figure 5 A schematic diagram of a target rendering model according to an embodiment of the present disclosure is shown. For example... Figure 5 As shown, the target rendering model includes a visual feature encoder and an information calibration encoder. For any video frame to be rendered (e.g., the i-th video frame in the video to be rendered), the visual feature encoder extracts features from the video frame to be rendered to obtain the first visual feature in the target dimension. The visual feature encoder also extracts features from the reference video frame corresponding to the video frame to be rendered (the real video frame before removing the lower half of the face region from the video frame to be rendered, e.g., the i-th video frame in the reference video) to obtain the second visual feature in the target dimension. The value of the target dimension can be flexibly set according to the actual situation, for example, 512 dimensions; this disclosure does not specifically limit this. The visual feature encoder can be a VGG-19 network structure, or other visual feature extraction structures can be flexibly set according to the actual scene requirements; this disclosure does not specifically limit this.
[0086] The first and second visual features are input into the information calibration encoder. The information calibration encoder can adjust the feature weights of the first and second visual features, and then process them in conjunction with the cross-attention mechanism. For example, the second visual feature is used as the key value in the cross-attention mechanism, and the first visual feature is used as the query in the cross-attention mechanism. Finally, the information calibration encoder can output the cross-attention visual features of the target dimension of the video frame to be rendered.
[0087] For any given video frame to be rendered, the target rendering model can be used to determine the audio segment corresponding to the video frame from the input translated audio. Then, feature extraction is performed on the audio segment corresponding to the video frame to obtain the audio features of the video frame to be rendered, which can be used as the driving signal for the video frame to be rendered.
[0088] like Figure 5As shown, the target rendering model includes an audio feature encoder. For any video frame to be rendered, the audio feature encoder extracts features from the translated audio of the corresponding audio segment of the video frame to obtain the target dimension audio features of the video frame. That is, the cross-interest visual features and audio features corresponding to the video frame to be rendered have the same dimension to facilitate subsequent processing. The audio feature encoder can have the same structure as the video feature encoder. For example, the audio feature encoder and the video feature encoder can both be VGG-19 network structures. This disclosure does not specifically limit this.
[0089] Since the character's mouth movements need to be driven by the audio features of the video frame to be rendered, it is necessary to determine the affine coefficients of the audio features of the video frame to be rendered in different channel dimensions.
[0090] like Figure 5 As shown, the target rendering model includes a fully connected layer. For any video frame to be rendered, the audio features of the video frame are input into the fully connected layer, and after processing by the fully connected layer, affine coefficients are output.
[0091] For example, for any video frame to be rendered, determine the affine coefficients [R, T, S] of the audio features of the video frame in different channel dimensions, where, Indicates the rotation coefficient. , Indicates the translation coefficient. denoted by scale factor, and c represents the number of dimensions of the audio feature corresponding to the video frame to be rendered.
[0092] Based on the affine coefficients of the audio features of the video frame to be rendered in different channel dimensions, the cross-interest visual features of the video frame to be rendered are deformed in spatial dimension to obtain the deformed visual features of the video frame to be rendered.
[0093] like Figure 5 As shown, the target rendering model includes an affine transformation module. For any video frame to be rendered, the affine coefficients output from the fully connected layer of the audio features of the video frame to be rendered, and the cross-interest visual features output from the information calibration encoder, are input into the affine transformation module. The affine transformation module performs deformation processing on the cross-interest visual features based on the affine coefficients to obtain the deformed visual features of the video frame to be rendered. The affine transformation module can be the Adaptive Affine Transformation (AdaAT) module in related technologies, and other affine transformation modules can also be flexibly set according to the actual scenario. This disclosure does not specifically limit this.
[0094] For example, spatial deformation can be performed based on the following formula (1):
[0095] (1).
[0096] wherein, represents the pre-deformation coordinate corresponding to the channel c, represents the post-deformation coordinate corresponding to the channel c.
[0097] For any one to be rendered video frame, based on the audio feature of the to be rendered video frame, the deformed visual feature, the to be rendered video frame is rendered and processed, and the rendered video frame corresponding to the to be rendered video frame is obtained.
[0098] As Figure 5 shown, the target rendering model includes a rendering decoder. For any one to be rendered video frame, the deformed visual feature of the to be rendered video frame output by the affine transformation module and the audio feature of the to be rendered video frame output by the audio feature encoder are input into the rendering decoder, and the to be rendered video frame corresponding to the to be rendered video frame is output after being rendered by the rendering decoder. The rendering decoder can be a decoder structure with UNet as the main body, or the structure can be flexibly set according to the actual scene, and the present disclosure does not make specific limitations thereto.
[0099] After the above processing is performed for each to be rendered video frame in the to be rendered video, the rendered video frame corresponding to each to be rendered video frame is obtained, and the final face rendering video is constituted.
[0100] In the case where only the face region of the target person is included in the face rendering video, the face rendering video needs to be fused with the background part in the original target mute video to obtain a fused video.
[0101] For example, the face rendering video in which the mouth movement of the person has been translated by the translated video is combined with the position of the face detection frame corresponding to the face region extracted at the time, is first converted into the original face size, and is then pasted back to the specified position in the target mute video to obtain the fused video.
[0102] However, the fused video obtained based on the above fusion method may have pseudo-reverse ghosting or tomographic phenomena at the fusion boundary, so that there is a significant inter-frame discontinuity problem in the video frame sequence of the fused video. Therefore, the fused video needs to be further optimized to obtain an inter-frame continuous target video.
[0103] In one possible implementation, the method further includes: using a stable diffusion model to perform redrawing processing on the fused video of the target mute video and the face rendering video to obtain a target video.
[0104] Since the face rendered video only includes the face area of the target person, the face rendered video needs to be fused with the target mute video to obtain a fused video. However, the fused video may have unnatural transition in the boundary area, and may have the phenomenon of pseudo ghosting or dissection. Therefore, the stable diffusion model is used to redraw the fused video to improve the fusion effect of the boundary area and obtain a target video with better video frame continuity.
[0105] The target video herein can be a video in which the target person speaks in a second language type, and the target video maintains the same task identity information and posture as the target person in the original video.
[0106] As shown in Figure 2 The video translation system includes a stable diffusion model. The target mute video output by the mouth movement elimination model and the face rendered video output by the target rendering model are input into the stable diffusion model, and after redraw processing by the stable diffusion model, the final target video is output.
[0107] In one possible implementation, the stable diffusion model includes a plurality of heterogeneous blocks. The stable diffusion model is used to redraw the fused video of the target mute video and the face rendered video to obtain the target video, including: inputting a preset noise timestamp and a feature vector corresponding to each video frame in the fused video into each heterogeneous block; based on the plurality of heterogeneous blocks, performing forward diffusion and reverse diffusion of noise on the fused video to obtain the target video.
[0108] The stable diffusion model constructs a forward diffusion and reverse diffusion process based on the input fused video to eliminate the dissection or pseudo ghosting phenomenon at the boundary of the face area, and generates a target video with natural boundary area transition and better inter-frame continuity.
[0109] The preset noise timestamp is a parameter for controlling the introduction and change of noise in the diffusion process, and its specific value can be flexibly set according to actual conditions, which is not specifically limited in the present disclosure.
[0110] The plurality of heterogeneous blocks in the stable diffusion model refer to modules for performing operations such as feature extraction, transformation, and fusion. The specific structure of the stable diffusion model can refer to related technologies, which is not specifically limited in the present disclosure.
[0111] The preset noise timestamp is mapped to an embedding factor and directly used as input of each heterogeneous block in the diffusion model. Each video frame in the fused video is directly used as the latent feature of the input of each heterogeneous block by adding a feature vector.
[0112] The stable diffusion model constructs a forward diffusion process and a reverse diffusion process based on the input of each isomer block, and finally generates a target video with a natural transition in a boundary region and good interframe continuity.
[0113] In the embodiments of the present disclosure, the mouth movement of the target person in the original video is eliminated to obtain a target closed-mouth video, so as to reduce the influence of the mouth movement in the original video on the subsequent rendering process and improve the audio-visual synchronization rate. The original audio corresponding to the original video is translated to obtain translated audio of different language types corresponding to the original audio. Then, the target rendering model is used to drive the target closed-mouth video based on the translated audio, so as to effectively change the language type of the original video and obtain a face rendering video in which the mouth movement of the target person matches the translated audio. Further, the stable diffusion model is used to redraw a fusion video of the target closed-mouth video and the face rendering video, so as to improve the effect at the fusion boundary and obtain a target video with good continuity between video frames.
[0114] The target rendering model mentioned in the embodiments of the present disclosure is pre-trained before actual application. The training process of the target rendering model is described in detail below.
[0115] In a possible implementation, the training data of the target rendering model includes: a to-be-rendered sample video, a reference sample video, and a sample audio. The training process of the target rendering model includes: determining a face rendering video corresponding to the to-be-rendered sample video by using the target rendering model; determining a target training loss based on the to-be-rendered sample video, the face rendering video corresponding to the to-be-rendered sample video, the reference sample video, and the sample audio; and training the target rendering model based on the target training loss to obtain a trained target rendering model.
[0116] First, the training data of the target rendering model is constructed, and an original sample video is determined. The original sample video can be a video in which a person speaks in a certain language type, and the specific form is not limited in the present disclosure.
[0117] Second, a face region of the original sample video is extracted to obtain a reference sample video including only the face region, and then a to-be-rendered sample video is obtained by removing the lower half of the face region in the reference video.
[0118] Then, after audio translation of the audio of the original sample video, a sample audio of a different language type is obtained. The language type of the sample audio is different from that of the audio of the original sample video, and the specific language type is not limited in the present disclosure.
[0119] Finally, the to-be-rendered sample video, the reference sample video, and the sample audio constitute the training data of the target rendering model.
[0120] The training data of the target rendering model is used to train the target rendering model.
[0121] In a possible implementation, the reference sample video includes a plurality of reference sample video frames, the sample video to be rendered includes a plurality of sample video frames to be rendered, and each sample video frame to be rendered does not include a lower half face region of a person; determining, by using the target rendering model, a face rendering video corresponding to the sample video to be rendered includes: for any one of the sample video frames to be rendered, performing, by using the target rendering model, identity information calibration on the sample video frame to be rendered and a first reference sample video frame corresponding to the sample video frame to be rendered based on the cross-attention mechanism to obtain cross-attention visual features of the sample video frame to be rendered, where the first reference sample video frame corresponding to the sample video frame to be rendered is any one of the plurality of reference sample video frames; determining, by using the target rendering model, audio features of the sample video frame to be rendered from the sample audio, and deforming, by using an affine coefficient determined based on the audio features of the sample video frame to be rendered, the cross-attention visual features of the sample video frame to be rendered to obtain deformed visual features of the sample video frame to be rendered; performing, by using the target rendering model, rendering processing on the sample video frame to be rendered based on the audio features and the deformed visual features of the sample video frame to be rendered to obtain a rendering video frame corresponding to the sample video frame to be rendered; and obtaining the face rendering video corresponding to the sample video to be rendered based on the rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered.
[0122] In the training stage, the process of obtaining the face rendering video corresponding to the sample video to be rendered based on the sample video to be rendered, the reference sample video, and the sample audio is similar to the process of obtaining the face rendering video corresponding to the sample video to be rendered based on the sample video to be rendered, the reference video, and the translated audio, and reference can be made to the related content described above, which is not repeated here.
[0123] It should be noted that, in the training stage, for any one of the sample video frames to be rendered, unlike in the application stage, the first reference sample video frame corresponding to the sample video frame to be rendered can be any one of the reference sample video frames, and it is not necessary to limit that the first reference sample video frame must be a real video frame of the sample video to be rendered in the reference sample video.
[0124] For example, the face region extraction is performed on the original sample video to obtain a reference sample video, and the lower half face region is removed from each reference sample video frame in the reference sample video to obtain a to-be-rendered sample video frame. That is, the i th to-be-rendered sample video frame in the to-be-rendered sample video is obtained by removing the lower half face region from the i th reference sample video frame in the reference sample video, that is, the i th reference sample video frame is the real video frame of the i th to-be-rendered sample video frame. However, in the training stage, a reference sample video frame is randomly selected from the reference sample video as the first reference sample video frame corresponding to the i th to-be-rendered sample video frame. The randomly selected first reference sample video frame can be the i th reference sample video frame (real video frame), or can not be the i th reference sample video frame.
[0125] Based on the to-be-rendered sample video, the to-be-rendered sample video corresponding face rendering video, the reference sample video, and the sample audio, the training loss can be determined from multiple angles to obtain the final comprehensive target training loss.
[0126] In a possible implementation, based on the to-be-rendered sample video, the to-be-rendered sample video corresponding face rendering video, the reference sample video, and the sample audio, the target training loss is determined, comprising: based on the to-be-rendered sample video, the to-be-rendered sample video corresponding face rendering video, the reference sample video, and the sample audio, an initial training loss is determined, wherein the initial training loss includes at least one of the following: reconstruction loss, first synchronization loss, second synchronization loss, discriminator loss, and multi-scale loss; based on the initial training loss, the target training loss is determined.
[0127] In the training process, one or more of the reconstruction loss, the first synchronization loss, the second synchronization loss, the discriminator loss, and the multi-scale loss can be considered comprehensively to determine the target training loss that can more effectively train the target rendering model, and improve the model training accuracy.
[0128] In a possible implementation, the initial training loss is determined based on the sample video to be rendered, the face video rendered corresponding to the sample video to be rendered, the reference sample video, and the sample audio, including: determining a reconstruction loss based on the face video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video, wherein each sample video frame to be rendered is obtained by removing the lower half face region from the corresponding second reference video frame; and / or determining a first synchronization loss based on the face video rendered corresponding to the sample video to be rendered and the sample audio; and / or determining a second synchronization loss based on the face video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video; and / or determining a discriminator loss based on the face video rendered corresponding to the sample video to be rendered and the reference sample video; and / or determining a multi-scale loss based on the face video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video.
[0129] To improve the audio-visual synchronization rate, the reconstruction loss is determined based on the face video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video.
[0130] In an example, the reconstruction loss can be determined based on the following formula (2):
[0131] (2).
[0132] wherein, the reconstruction loss is represented by Lr, k represents the number of video frames of the sample video to be rendered and the reference sample video, the face video frame corresponding to the i th sample video frame to be rendered is represented by F i, the second reference video frame corresponding to the i th sample video frame to be rendered, that is, the real video frame (i th reference sample video frame) corresponding to the i th sample video frame to be rendered, is represented by R i.
[0133] To further improve the audio-visual synchronization rate, the first synchronization loss is determined based on the face video rendered corresponding to the sample video to be rendered and the sample audio.
[0134] In an example, the face rendered video corresponding to the sample video to be rendered and the sample audio can be input into the trained SyncNet network. The SyncNet network is built based on the network structure of the trained image classification model (VGG-19) to build a video encoder and an audio encoder to extract features from the face rendered video corresponding to the sample video to be rendered and the sample audio respectively. The similarity of the encoded features output by the two encoders can be determined by using the following formula (3):
[0135] (3).
[0136] wherein, represents the encoded feature of the face rendered video frame corresponding to the i-th video frame to be rendered, represents the encoded feature of the audio segment corresponding to the i-th video frame to be rendered in the sample audio, represents the similarity between and is a value between 0 and 1, is a parameter set to prevent the denominator in formula (3) from being 0.
[0137] Further, the first synchronization loss is determined by minimizing the similarity loss by using the following formula (4):
[0138] (4).
[0139] In order to further improve the audio-visual synchronization rate, the second synchronization loss is determined based on the face rendered video frame corresponding to each video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each video frame to be rendered in the reference sample video.
[0140] In an example, the sample video to be rendered and the reference sample video can be input into the trained image classification model (VGG-16) to extract features, and then the second synchronization loss is determined by using the following formula (5):
[0141] (5).
[0142] wherein, represents the second synchronization loss, represents the feature extracted from the face rendered video frame corresponding to the i-th sample video frame to be rendered, represents the feature extracted from the second reference video frame (real video frame, i-th reference sample video frame) corresponding to the i-th sample video frame to be rendered.
[0143] To further improve the overall quality of the rendered video, the discriminator loss is determined based on the face rendering video corresponding to the sample video to be rendered and the reference sample video.
[0144] In one example, the discriminator loss can be determined based on the following formula (6):
[0145] (6).
[0146] in, Represents the generator loss function. To help the discriminator determine the face rendering video corresponding to the sample video to be rendered ( This is a real video (reference sample video). The probability of ). This indicates the loss of the discriminator.
[0147] To achieve a coarse-to-fine image quality optimization effect, a multi-scale loss is determined based on the face rendering video frame corresponding to each video sample to be rendered in the video sample to be rendered, and the second reference video frame corresponding to each video sample to be rendered in the reference video.
[0148] In one example, the multiscale loss can be determined based on the following formula (7):
[0149] (7).
[0150] in, This represents a video frame at half the spatial scale of the face rendering video frame corresponding to the i-th sample video frame to be rendered. This represents a 1 / 4 spatial scale video frame corresponding to the face rendering video frame of the i-th sample video frame to be rendered. This represents a half-scale video frame corresponding to the second reference video frame (real video frame, i-th reference sample video frame) of the i-th sample video frame to be rendered. It represents a 1 / 4 spatial scale video frame of the second reference video frame (real video frame, i-th reference sample video frame) corresponding to the i-th sample video frame to be rendered.
[0151] In one possible implementation, the target training loss is determined based on the initial training loss, including: when the initial training loss includes reconstruction loss, first synchronization loss, second synchronization loss, discriminator loss, and multi-scale loss, the target training loss is obtained by weighted summation of the reconstruction loss, first synchronization loss, second synchronization loss, discriminator loss, and multi-scale loss.
[0152] In one example, the target training loss can be determined using the following formula (8):
[0153] (8).
[0154] wherein, to represent the weight coefficients of each loss, the specific values can be flexibly set according to the actual situation, and the present disclosure does not make specific limitations hereon.
[0155] In a round of training process, after adjusting the network parameters of the target rendering model based on the target training loss, the next round of training process is performed using the adjusted target rendering model, and after iterating the training for a preset number of rounds or reaching a preset training condition, the training is ended, and a trained target rendering model is obtained.
[0156] The trained target rendering model can be added to the above-mentioned video translation system to perform a video translation task in an application scenario.
[0157] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to the limited space, the present disclosure will not be repeated. Those skilled in the art can understand that in the above-mentioned method of the specific embodiment, the specific execution order of each step should be determined according to its function and possible internal logic.
[0158] In addition, the present disclosure also provides a video translation system, an electronic device, a computer readable storage medium, and a program, all of which can be used to implement any one of the video translation methods provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding description in the method part, and will not be repeated.
[0159] Figure 6 A block diagram of a video translation system according to an embodiment of the present disclosure is shown. As shown in Figure 6 The system 60 includes:
[0160] A mouth movement elimination model 61 for eliminating the mouth movement of a target person in an original video to obtain a target closed-mouth video.
[0161] An audio translation model 62 for performing audio translation on the original audio corresponding to the original video to obtain translated audio, wherein the original audio and the translated audio correspond to different language types.
[0162] A target rendering model 63 for driving the target closed-mouth video based on the translated audio to obtain a face rendering video, wherein the mouth movement of the target person in the face rendering video matches the translated audio.
[0163] In a possible implementation, the mouth movement elimination model 61 is specifically configured to:
[0164] eliminate the mouth movement of the target person in the original video to obtain a first closed-mouth video;
[0165] perform video reconstruction based on the first mute video and the original video to obtain a second mute video, where the second mute video has the same pose change as the original video;
[0166] perform fusion based on the second mute video and the original video to obtain a target mute video, where the target mute video has the same person identity information as the original video.
[0167] In a possible implementation, the mouth action elimination model 61 is specifically configured to:
[0168] determine original 3DMM coefficients of the original video based on a 3DMM fitting algorithm, where the original 3DMM coefficients include original pose coefficients and original expression coefficients;
[0169] replace the original expression coefficients in the original 3DMM coefficients with preset neutral expression coefficients to obtain target 3DMM coefficients;
[0170] perform expression modification on the original video based on the target 3DMM coefficients to obtain the first mute video.
[0171] In a possible implementation, the mouth action elimination model 61 is specifically configured to:
[0172] perform feature extraction on the original video to obtain 3D key point motion features of the original video;
[0173] perform video reconstruction on the first mute video based on the 3D key point motion features of the original video to obtain the second mute video.
[0174] In a possible implementation, the mouth action elimination model 61 is specifically configured to:
[0175] perform face analysis on a target person in the original video, and perform region extraction on the original video based on a face analysis result to obtain a first to-be-fused video, where the first to-be-fused video includes an upper half face region and a nose region of the target person;
[0176] perform face analysis on the target person in the second mute video, and perform region extraction on the second mute video based on a face analysis result to obtain a second to-be-fused video, where the second to-be-fused video includes a lower half face region of the target person and does not include the nose region;
[0177] perform fusion on the first to-be-fused video and the second to-be-fused video to obtain the target mute video.
[0178] In a possible implementation, the target rendering model 63 is specifically configured to:
[0179] Face region extraction is performed on a target person in a target close-mouth video to obtain a reference video and a video to be rendered;
[0180] Based on the video to be rendered, the reference video, and the translated audio, a face rendered video is obtained.
[0181] In a possible implementation, the reference video includes a plurality of reference video frames, and the video to be rendered includes a plurality of video frames to be rendered, and each video frame to be rendered does not include a lower half face region of the target person;
[0182] The target rendering model 63 is specifically configured to:
[0183] For any one video frame to be rendered, based on the cross-attention mechanism, the video frame to be rendered and the reference video frame corresponding to the video frame to be rendered are calibrated in terms of person identity information, to obtain cross-attention visual features of the video frame to be rendered, wherein the video frame to be rendered is obtained by removing the lower half face region from the corresponding reference video frame;
[0184] The audio features of the video frame to be rendered are determined from the translated audio, and the cross-attention visual features of the video frame to be rendered are deformed by using the affine coefficients determined based on the audio features of the video frame to be rendered, to obtain deformed visual features of the video frame to be rendered;
[0185] Based on the audio features and the deformed visual features of the video frame to be rendered, the video frame to be rendered is rendered to obtain a rendered video frame corresponding to the video frame to be rendered;
[0186] Based on the rendered video frame corresponding to each video frame to be rendered in the video to be rendered, a face rendered video is obtained.
[0187] In a possible implementation, the training data of the target rendering model 63 includes a sample video to be rendered, a reference sample video, and a sample audio.
[0188] The system 60 further includes a training module specifically configured to:
[0189] The target rendering model 63 is used to determine a face rendered video corresponding to the sample video to be rendered;
[0190] Based on the sample video to be rendered, the face rendered video corresponding to the sample video to be rendered, the reference sample video, and the sample audio, a target training loss is determined.
[0191] The target rendering model 63 is trained based on the target training loss, to obtain a trained target rendering model 63.
[0192] In a possible implementation, the reference sample video includes a plurality of reference sample video frames, and the sample video to be rendered includes a plurality of sample video frames to be rendered, and each sample video frame to be rendered does not include a lower half face region of a person;
[0193] The target rendering model 63 is specifically configured to:
[0194] For any one sample video frame to be rendered, the identity information of the person in the sample video frame to be rendered and the first reference sample video frame corresponding to the sample video frame to be rendered are calibrated based on the cross-attention mechanism to obtain cross-attention visual features of the sample video frame to be rendered, wherein the first reference sample video frame corresponding to the sample video frame to be rendered is any one of the plurality of reference sample video frames.
[0195] The audio features of the sample video frame to be rendered are determined from the sample audio, and the cross-attention visual features of the sample video frame to be rendered are deformed using the affine coefficients determined based on the audio features of the sample video frame to be rendered to obtain deformed visual features of the sample video frame to be rendered.
[0196] The sample video frame to be rendered is rendered based on the audio features and the deformed visual features of the sample video frame to be rendered to obtain a rendered video frame corresponding to the sample video frame to be rendered.
[0197] The face rendering video corresponding to the sample video to be rendered is obtained based on the rendered video frame corresponding to each sample video frame to be rendered in the sample video to be rendered.
[0198] In a possible implementation, the training module is specifically configured to:
[0199] Based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio, an initial training loss is determined, wherein the initial training loss includes at least one of the following: a reconstruction loss, a first synchronization loss, a second synchronization loss, a discriminator loss, and a multi-scale loss.
[0200] The target training loss is determined based on the initial training loss.
[0201] In a possible implementation, the training module is specifically configured to:
[0202] Based on the face rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video, a reconstruction loss is determined, wherein each sample video frame to be rendered is obtained by removing the lower half face region from the corresponding second reference video frame; and / or,
[0203] determine the first synchronization loss based on the face rendered video corresponding to the to-be-rendered sample video and the sample audio; and / or,
[0204] determine the second synchronization loss based on the face rendered video frame corresponding to each to-be-rendered video sample frame in the to-be-rendered sample video and the second reference sample video frame corresponding to each to-be-rendered video sample frame in the reference sample video; and / or,
[0205] determine the discriminator loss based on the face rendered video corresponding to the to-be-rendered sample video and the reference sample video; and / or,
[0206] determine the multi-scale loss based on the face rendered video frame corresponding to each to-be-rendered video sample frame in the to-be-rendered sample video and the second reference sample video frame corresponding to each to-be-rendered video sample frame in the reference sample video.
[0207] In a possible implementation, the training module is specifically configured to:
[0208] In the case where the initial training loss includes the reconstruction loss, the first synchronization loss, the second synchronization loss, the discriminator loss, and the multi-scale loss, the reconstruction loss, the first synchronization loss, the second synchronization loss, the discriminator loss, and the multi-scale loss are weighted and summed to obtain the target training loss.
[0209] In a possible implementation, the system 60 further includes:
[0210] The stable diffusion model is configured to perform re-drawing processing on the fusion video of the target closed-mouth video and the face rendered video to obtain the target video.
[0211] In a possible implementation, the stable diffusion model includes a plurality of isomer blocks.
[0212] The stable diffusion model is specifically configured to:
[0213] The preset noise timestamp and the feature vector corresponding to each video frame in the fusion video are input into each isomer block.
[0214] The fusion video is forward diffused and backward diffused based on the plurality of isomer blocks to obtain the target video.
[0215] The method has specific technical correlation with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing data storage, reducing data transmission, improving hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system in accordance with the natural law.
[0216] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and specific implementation can be referred to the description of the above method embodiments. For brevity, details are not described here again.
[0217] The embodiments of the present disclosure also provide a computer readable storage medium having stored computer program instructions, and the computer program instructions are executed by a processor to implement the above method. The computer readable storage medium can be a volatile or non-volatile computer readable storage medium.
[0218] The embodiments of the present disclosure also provide an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the above method.
[0219] The embodiments of the present disclosure also provide a computer program product, comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in the processor of an electronic device, the processor in the electronic device executes the above method.
[0220] The electronic device can be provided as a terminal, a server or other forms of devices.
[0221] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Referring to Figure 7 , the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932, for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above method.
[0222] The electronic device 1900 can also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input output interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Microsoft Windows Server TM , Apple's graphical user interface-based operating system (Mac OS X TM ), multi-user multi-process computer operating system (Unix TM), a free and open-source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ), or the like.
[0223] In an example embodiment, there is also provided a non-transitory computer- readable storage medium, such as the memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to implement the above-described method.
[0224] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0225] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium can also include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a magnetically encoded device such as magnetic strip cards, an optically encoded device such as a compact disc (CD) or DVD, and / or any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0226] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0227] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0228] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0229] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0230] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0231] The flow diagrams and the block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0232] The computer program product can be embodied in particular by hardware, software or a combination thereof. In an alternative embodiment, the computer program product is embodied in particular as a computer storage medium, in another alternative embodiment, the computer program product is embodied in particular as a software product, such as a software development kit (SDK) or the like.
[0233] The above description of the various embodiments is intended to be illustrative in all respects, and not restrictive. The scope of the present disclosure is indicated by the appended claims, rather than the foregoing description, and all changes that come within the meaning and range of equivalents are intended to be embraced therein.
[0234] Those skilled in the art can understand that, in the above-described method of the specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible inherent logic.
[0235] If the technical solutions of the present application involve personal information, the product applying the technical solutions of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solutions of the present application involve sensitive personal information, the product applying the technical solutions of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his personal information, the individual's authorization is obtained under the condition that the device uses obvious mark / information to inform the individual of the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.
[0236] The above has described various embodiments of the present disclosure, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical application or improvement of technology in the market of the embodiments, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method of video translation, characterized by, The method comprises the following steps: performing mouth action elimination on a target person in an original video to obtain a target mute video; performing audio translation on original audio corresponding to the original video to obtain translated audio, wherein the original audio and the translated audio correspond to different language types; driving the target mute video based on the translated audio by using a target rendering model to obtain a face rendering video, wherein the mouth action of the target person in the face rendering video matches the translated audio; The method of performing mouth action elimination on a target person in an original video to obtain a target mute video comprises the following steps: performing mouth action elimination on the target person in the original video to obtain a first mute video; performing video reconstruction based on the first mute video and the original video to obtain a second mute video, wherein the second mute video has the same attitude change as the original video; performing face analysis on the target person in the original video, and extracting a region from the original video based on the face analysis result to obtain a first to-be-fused video, wherein the first to-be-fused video includes an upper half face region and a nose region of the target person; performing face analysis on the target person in the second mute video, and extracting a region from the second mute video based on the face analysis result to obtain a second to-be-fused video, wherein the second to-be-fused video includes a lower half face region of the target person and does not include the nose region; fusing the first to-be-fused video and the second to-be-fused video to obtain the target mute video, wherein the target mute video has the same person identity information as the original video.
2. The method of claim 1, wherein, The method of performing mouth action elimination on a target person in an original video to obtain a first mute video comprises the following steps: determining original 3DMM coefficients of the original video based on a three-dimensional variable model 3DMM fitting algorithm, wherein the original 3DMM coefficients include original attitude coefficients and original expression coefficients; replacing the original expression coefficients in the original 3DMM coefficients with preset neutral expression coefficients to obtain target 3DMM coefficients; performing expression modification on the original video based on the target 3DMM coefficients to obtain the first mute video.
3. The method of claim 1, wherein, The method of performing video reconstruction based on the first mute video and the original video to obtain a second mute video comprises the following steps: extracting features of the original video to obtain 3D key point motion features of the original video; performing video reconstruction on the first mute video based on the 3D key point motion features of the original video to obtain the second mute video.
4. The method of claim 1, wherein, The method of driving the target mute video based on the translated audio by using a target rendering model to obtain a face rendering video comprises the following steps: extracting a face region of the target person in the target mute video to obtain a reference video and a to-be-rendered video; inputting the to-be-rendered video, the reference video, and the translated audio into the target rendering model to obtain the face rendering video output by the target rendering model.
5. The method of claim 4, wherein, The reference video includes a plurality of reference video frames, and the video to be rendered includes a plurality of video frames to be rendered, each of which does not include a lower half face region of the target person; The inputting the video to be rendered, the reference video and the translated audio into the target rendering model to obtain the face rendering video output by the target rendering model comprises: For any one video frame to be rendered, the target rendering model is used to calibrate the identity information of the person based on the cross-attention mechanism, to obtain the cross-attention visual feature of the video frame to be rendered, wherein the video frame to be rendered is obtained by removing the lower half face region from the corresponding reference video frame; The target rendering model is used to determine the audio feature of the video frame to be rendered from the translated audio, and the affine coefficient determined based on the audio feature of the video frame to be rendered is used to deform the cross-attention visual feature of the video frame to be rendered, to obtain the deformed visual feature of the video frame to be rendered; The target rendering model is used to render the video frame to be rendered based on the audio feature and the deformed visual feature of the video frame to be rendered, to obtain the rendered video frame corresponding to the video frame to be rendered; The face rendering video is obtained based on the rendered video frame corresponding to each video frame to be rendered in the video to be rendered.
6. The method of claim 1, wherein, The training data of the target rendering model includes a sample video to be rendered, a reference sample video, and a sample audio; The training process of the target rendering model includes: The target rendering model is used to determine the face rendering video corresponding to the sample video to be rendered; The target training loss is determined based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio; The target rendering model is trained based on the target training loss, to obtain the trained target rendering model.
7. The method of claim 6, wherein, The reference sample video includes a plurality of reference sample video frames, and the sample video to be rendered includes a plurality of sample video frames to be rendered, each of which does not include a lower half face region of a person; The target rendering model is used to determine the face rendering video corresponding to the sample video to be rendered, comprising: For any one sample video frame to be rendered, the target rendering model is used to calibrate the identity information of the person based on the cross-attention mechanism, to obtain the cross-attention visual feature of the sample video frame to be rendered, wherein the first reference sample video frame corresponding to the sample video frame to be rendered is any one of the plurality of reference sample video frames; The target rendering model is used to determine the audio feature of the sample video frame to be rendered from the sample audio, and the affine coefficient determined based on the audio feature of the sample video frame to be rendered is used to deform the cross-attention visual feature of the sample video frame to be rendered, to obtain the deformed visual feature of the sample video frame to be rendered; The target rendering model is used to perform rendering processing on the sample video frame to be rendered based on the audio feature, the deformed visual feature, and the sample video frame to be rendered, so as to obtain a rendered video frame corresponding to the sample video frame to be rendered. The face rendering video corresponding to the sample video to be rendered is obtained based on the rendered video frame corresponding to each sample video frame to be rendered in the sample video to be rendered.
8. The method according to claim 6 or 7, characterized in that, The target training loss is determined based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio. The initial training loss is determined based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio, and the initial training loss includes at least one of a reconstruction loss, a first synchronization loss, a second synchronization loss, a discriminator loss, and a multi-scale loss. The target training loss is determined based on the initial training loss.
9. The method of claim 8, wherein, The initial training loss is determined based on the sample video to be rendered, the face rendering video corresponding to the sample video to be rendered, the reference sample video, and the sample audio. The reconstruction loss is determined based on the face rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video, and each sample video frame to be rendered is obtained by removing the lower half face region from the corresponding second reference video frame; and / or The first synchronization loss is determined based on the face rendering video corresponding to the sample video to be rendered and the sample audio; and / or The second synchronization loss is determined based on the face rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video; and / or The discriminator loss is determined based on the face rendering video corresponding to the sample video to be rendered and the reference sample video; and / or The multi-scale loss is determined based on the face rendering video frame corresponding to each sample video frame to be rendered in the sample video to be rendered and the second reference sample video frame corresponding to each sample video frame to be rendered in the reference sample video.
10. The method of claim 8, wherein, The target training loss is determined based on the initial training loss. In a case where the initial training loss includes the reconstruction loss, the first synchronization loss, the second synchronization loss, the discriminator loss, and the multi-scale loss, the reconstruction loss, the first synchronization loss, the second synchronization loss, the discriminator loss, and the multi-scale loss are weighted and summed to obtain the target training loss.
11. The method of claim 1, wherein, The method further includes: The fusion video of the target closed-mouth video and the face rendering video is redrawn using a stable diffusion model to obtain a target video.
12. The method of claim 11, wherein, The stable diffusion model includes a plurality of isomer blocks. The target video is obtained by performing redraw processing on the fusion video of the target closed-mouth video and the face rendering video based on the stable diffusion model. The preset noise timestamp and the feature vector corresponding to each video frame in the fusion video are input into each isomer block. The fusion video is forward diffused and backward diffused based on the plurality of isomer blocks to obtain the target video.
13. A video translation system characterized by, The method comprises: a mouth movement elimination model configured to eliminate mouth movement of a target person in an original video to obtain a target closed-mouth video; an audio translation model configured to translate original audio corresponding to the original video to obtain translated audio, wherein the original audio and the translated audio correspond to different language types; a target rendering model configured to drive the target closed-mouth video based on the translated audio to obtain a face rendering video, wherein mouth movement of the target person in the face rendering video matches the translated audio; The mouth movement elimination model is specifically configured to: eliminate mouth movement of the target person in the original video to obtain a first closed-mouth video; perform video reconstruction based on the first closed-mouth video and the original video to obtain a second closed-mouth video, wherein the second closed-mouth video has the same posture change as the original video; perform face analysis on the target person in the original video, and perform region extraction on the original video based on the face analysis result to obtain a first fusion video, wherein the first fusion video includes an upper half face region and a nose region of the target person; perform face analysis on the target person in the second closed-mouth video, and perform region extraction on the second closed-mouth video based on the face analysis result to obtain a second fusion video, wherein the second fusion video includes a lower half face region of the target person and does not include the nose region; fuse the first fusion video and the second fusion video to obtain the target closed-mouth video, wherein the target closed-mouth video has the same person identity information as the original video.
14. An electronic device, comprising: The method comprises: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method of any one of claims 1 to 12.
15. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by the processor, implement the method of any one of claims 1 to 12. The computer program instructions, when executed by the processor, implement the method of any one of claims 1 to 12.
Citation Information
Patent Citations
Generation method and device of digital human animation, electronic equipment and storage medium
CN115830193A
Video generation and model training method and device, equipment and storage medium
CN116385604A
Real-time multi-language processing live broadcast method and system based on deep learning
CN117253486A