Figure processing method and system applied to video conversion
By introducing a lip-type conversion model in language conversion technology, the problem of inconsistent character lip-type and line language after language conversion is solved, and a better user experience is achieved.
Patent Information
- Application Number
- CN202411582147.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-11-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-07
AI Technical Summary
After the language changes, the lip movements of the characters in the film do not correspond to the converted lines, which reduces the user's viewing experience.
By obtaining the audio and character images in the video to be converted, performing language conversion, and inputting the converted audio and original character images into the lip-type conversion model to generate synchronized lip-type action images, thereby maintaining the original sound characteristics and lip-type actions of the character.
After the language conversion is realized, the character's lip movements correspond to the translated audio pronunciation, improving the user's viewing experience.
Smart Images

Figure CN119967223A_ABST
Abstract
Description
[0001] This application claims priority to a Chinese patent application filed with the State Intellectual Property Office with application number 202311479579.X on November 8, 2023. The entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the technical field of voice conversion, and in particular to a character processing method and system applied to video conversion. Background Art
[0003] When users watch movies from different countries, they cannot directly understand the meaning of the lines from the audio of the lines spoken by the characters in the movie because the language of the movie is different from the user's language. Therefore, users need to watch subtitles continuously while watching the movie to understand the meaning of the lines of the characters in the movie.
[0004] To this end, voice conversion technology can be used to convert the language of the film into the language desired by the user while retaining the voice characteristics of the characters in the film, thereby achieving the effect of real-time translation. In the above process, since the language of the characters' lines has changed, the original speaking lip shape of the characters in the film cannot correspond to the converted lines, which reduces the user's viewing experience. Summary of the invention
[0005] In order to solve the problem that the mouth shape of the characters in the picture does not correspond to the pronunciation of the language after language conversion, some embodiments of the present application provide a character processing method and system for video conversion.
[0006] In a first aspect, the present application provides a character processing method applied to video conversion, comprising:
[0007] Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image at least includes a first lip movement of a target person speaking in the first language;
[0008] Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language;
[0009] Inputting the target audio and the first character image into a lip-sync conversion model to obtain a second character image output by the lip-sync conversion model, wherein the second character image at least includes a second lip movement of the target character speaking in the second language; the lip-sync conversion model includes a transformation module, wherein the transformation module is configured to calculate a transformation coefficient corresponding to each feature channel in the image feature of the first character image at least according to the target audio, and perform feature transformation on the image feature channel of the first character image according to the transformation coefficient to obtain the image feature of the first character image;
[0010] A target video is generated according to the target audio and the second person image.
[0011] In some embodiments, before the step of inputting the target audio and the first character image into the lip-sync model, the method further includes:
[0012] Acquire training data; the training data includes a first training image, the first training image includes at least one training character with lip movements;
[0013] Inputting the training data into a model to be trained, so as to extract training audio features and a first training image of the training data through the model to be trained, wherein the model to be trained is a lip-sync model that has not been trained to convergence;
[0014] Based on the training audio features, performing feature channel conversion on the first training image through the model to be trained to generate a second training image;
[0015] Calculate the training loss of the second training image and the lip shape image label according to the loss function;
[0016] If the training loss is greater than the loss threshold, iteratively training the model to be trained;
[0017] If the training loss is less than or equal to the loss threshold, the lip conversion model is input according to the model parameters of the model to be trained.
[0018] In some embodiments, the model to be trained includes an encoder, a transformation module and a decoder, and the step of performing feature channel conversion on the first training image through the model to be trained based on the training audio features includes:
[0019] Encoding the training audio features by the encoder to obtain a training audio code, and encoding the first training image by the encoder to obtain a first training image code;
[0020] Determining, by the transformation module, a feature channel calculation relationship at least according to the training audio code and the first training image code;
[0021] According to the feature channel calculation relationship, performing feature channel conversion on the first training image code to obtain a second training image code;
[0022] The decoder decodes the second training image to obtain a second training image.
[0023] In some embodiments, the step of determining the feature channel calculation relationship at least according to the training audio code and the first training image code by the transformation module includes:
[0024] Obtain a reference image according to the first training image encoding, wherein the reference image is an image of a person identical to the training person in the first training image;
[0025] Splicing the reference image with the first training image to obtain a training spliced image;
[0026] The transform coefficients are calculated according to the training audio code and the spliced training image to determine the feature channel calculation relationship, wherein the transform coefficients are used to indicate the transformation relationship of the first training image relative to the reference image.
[0027] In some embodiments, the step of calculating the transformation coefficients according to the spliced training image includes:
[0028] extracting alignment features of the spliced training image by the encoder;
[0029] A transformation coefficient is calculated based on the aligned features and the training audio features.
[0030] In some embodiments, the transform coefficients include a first transform coefficient, a second transform coefficient, and a third transform coefficient, and the step of performing feature channel conversion on the first training image encoding according to the feature channel calculation relationship includes:
[0031] Calculating a rotation angle of a feature channel encoded by the first training image according to the first transform coefficient;
[0032] Calculating a translation distance of a feature channel encoded by the first training image according to the second transform coefficient;
[0033] Calculate a scaling factor of a feature channel of the first training image encoding according to the third transform coefficient;
[0034] The first training image code is converted into a second training image code according to the rotation angle, the translation distance, and the scaling factor.
[0035] In some embodiments, the encoder includes a hierarchical structure formed by a preset number of encoding units, the input of the encoding unit is the output of the previous encoding unit, the output of the encoding unit is the input of the next encoding unit, and the step of performing decoding on the second training image encoding by the decoder includes:
[0036] Dividing the second training image into a preset number of first image segments by encoding;
[0037] Inputting the first image segment into the encoder, so as to independently perform attention calculation on each of the first image segments through a certain encoding unit to obtain a first attention encoding result;
[0038] Adjusting the first image segment in a preset manner to divide the second training image encoding into a preset number of second image segments; wherein there is a partial overlap between the first image segment and the corresponding second image segment;
[0039] Inputting the first attention encoding result into an encoding unit at a next level, so that the encoding unit performs attention calculation on the second image segment to obtain a second attention encoding result;
[0040] Iterate the above operation to obtain a second feature encoding according to the second attention encoding result;
[0041] The second feature code is decoded to obtain a second training image.
[0042] In some embodiments, the decoder includes a preset number of upsampling layers; the step of decoding the second feature code to obtain a second training image includes:
[0043] Inputting the second feature code into the decoder, and performing a decoding operation on the second feature code according to the upsampling layer to obtain the second training image;
[0044] Wherein, a preset number of coding units in the coding module and the upsampling layer are distributed according to a corresponding structure;
[0045] In the decoder, the second feature code of the decoding operation performed by the upsampling layer is output by a non-corresponding encoding unit.
[0046] In some embodiments, the feature dimensions of the upsampling layer and the coding unit increase with the increase of the level, and the number of feature channels of the upsampling layer and the coding unit decreases with the increase of the level;
[0047] The encoding unit and the upsampling layer are connected at corresponding levels.
[0048] In a second aspect, some embodiments of the present application provide a character processing system for video conversion, a controller and a lip conversion model, wherein the lip conversion model includes a transformation module, wherein the transformation module is configured to calculate a transformation coefficient corresponding to each feature channel in an image feature of a first character image at least according to a target audio, and perform a feature transformation on the image feature channel of the first character image according to the transformation coefficient to obtain the image feature of the first character image; the controller is configured to:
[0049] Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image includes a first lip movement of a target person speaking in the first language;
[0050] Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language;
[0051] Inputting the target audio and the first character image into a lip-sync model to obtain a second character image output by the lip-sync model, wherein the second character image includes a second lip movement of the target character speaking in the second language;
[0052] A target video is generated according to the target audio and the second person image.
[0053] As can be seen from the above technical solutions, the present application provides a character processing method and system for video conversion, wherein the method obtains a video to be converted, performs language conversion on the audio to be converted of the video to be converted according to a second language, inputs the target audio and the first character image into a lip conversion model, obtains the second character image output by the lip conversion model, and then generates a target video according to the target audio and the second character image. The present application adjusts the lip movements in the first character image through the translated audio data, so that the lip movements of the target character in the target video correspond to the pronunciation of the translated target audio, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0055] Figure 1 A flowchart of a method for processing characters in video conversion provided in an embodiment of the present application;
[0056] Figure 2 A flowchart of performing training on a lip-sync conversion model according to an embodiment of the present application;
[0057] Figure 3 This is a training flow chart of the model to be trained in the embodiment of the present application;
[0058] Figure 4 This is a flow chart of calculating the characteristic channel calculation relationship in an embodiment of the present application;
[0059] Figure 5 This is a flow chart of calculating transformation coefficients according to training spliced images and training audio features in an embodiment of the present application;
[0060] Figure 6 This is a flow chart of performing feature transformation on feature channels according to transformation coefficients in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In order to make the purpose and implementation method of the present application clearer, the exemplary implementation method of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0062] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.
[0063] The terms "first", "second", "third", etc. in the specification and the above drawings of this application are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise specified. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.
[0064] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0065] Users can play videos on smart devices, such as movies, TV dramas, operas or variety shows, etc. When users watch movies from different countries, they cannot directly understand the semantics of the lines spoken by the characters in the movie because the language of the movie is different from the user's native language. For example, when a Chinese user watches an English movie, if the user cannot understand English, he cannot directly know the semantics of the lines based on the audio of the English lines spoken by the characters in the movie.
[0066] To this end, some movie resources may include subtitle resources, that is, when playing the movie screen, subtitles in the corresponding language are displayed in the movie screen. For example, when playing a movie in English, Chinese subtitles are added to the movie screen so that Chinese users can understand the meaning of the lines in the movie through Chinese subtitles. This requires users to watch the screen and the subtitles while watching the movie so that they can know the semantics of the characters' lines in the movie. Although the subtitles are located in the movie screen, repeatedly watching content located in different positions affects the user's viewing experience.
[0067] In order to improve the viewing experience of the film, the lines in the film can be converted into different languages before post-dubbing, so that users can intuitively understand the semantics of the lines according to the dubbing. However, this post-dubbing method consumes a lot of time and needs to be completed before playing the film, and it is impossible to translate and play the lines in real time.
[0068] To this end, voice conversion technology can be used to convert the language of the film into the language expected by the user while retaining the voice characteristics of the characters in the film, thereby achieving the effect of real-time translation. In the above process, since the language of the lines spoken by the characters has changed, the film maintains the original speaking lip shape of the characters and cannot correspond to the converted lines. For example, when the lines audio is in the first language, the film will play the lines audio in the first language based on the preset sound characteristics, and correspondingly, the lip shape movements of the characters in the film correspond to the pronunciation lip shape of the lines audio in the first language. After real-time translation, the lines audio is converted from the first language to the second language. Due to the different pronunciations between different languages, the lip shape movements of the characters in the film do not correspond to the lines audio in the second language, resulting in the characters' lip shape movements seen by the user not matching the actual audio pronunciation heard, affecting the user's viewing experience.
[0069] In order to solve the problem that the character's lip shape in the picture does not correspond to the language pronunciation after language conversion, some embodiments of the present application provide a character processing method applied to video conversion, which can be applied to a display device for playing audio and video or to a lip shape conversion system based on language conversion. Figure 1 A flowchart of a method for processing characters in video conversion provided by this application. Figure 1 , taking the application object as a character processing system applied to video conversion as an example, the method includes:
[0070] S100: Obtain the video to be converted.
[0071] The lip conversion system can obtain the video to be converted input by the user. The video to be converted includes the audio to be converted and the first character image. The audio to be converted is the original audio of the video to be converted, and is also the audio data for language conversion to be performed. The language information of the audio to be converted is the first language. The first language is the language of the audio to be converted before the translation is performed. For example, if the user is playing an English movie through a display device, the first language of the English movie is English.
[0072] In some embodiments, the video to be converted may also be a dubbed film, that is, the language of the audio to be converted is different from the original language, in which case the first language is the dubbed language. For example, if the video to be converted is an English film dubbed from Chinese, the first language of the audio to be converted is Chinese.
[0073] The first person image is a person who is speaking in the video to be converted, and the lip conversion system can perform frame processing on the video to be converted to obtain the first person image from the framed images. The first person image includes a first lip movement of the target person speaking in the first language, wherein the target person is a person who is speaking in the first person image.
[0074] In some embodiments, the first character image may also include facial images of other characters. For example, the front of the target character and the front images of other characters in the video to be converted may be simultaneously displayed in the video frame of the video to be converted. In the subsequent process of performing video conversion, the lip movements of other characters may be converted synchronously. In addition to facial images or lip movements, the first character image may also include information such as body movements, postures, and expressions of the target character, so that in the process of performing video conversion, the body movements, postures, and expressions are kept unchanged, or the action postures of the target character in the video to be converted are changed according to preset action information.
[0075] In some embodiments, the lip-sync conversion system can detect the number of target characters in the split-frame picture to collect the first character image according to the number of target characters. When the split-frame picture includes only one target character, the character image of the target character in the split-frame picture when in a speaking state can be used as the first character image. When the split-frame picture includes multiple target characters, the first character images of different target characters in a speaking state can be collected in the split-frame picture respectively, and a character image set can be generated according to the first character images of the corresponding target characters. For example, when the split-frame picture includes character A and character B who are in a speaking state at the same time, at this time, the lip-sync conversion system can obtain the first character image of character A and the first character image of character B respectively, and generate character image set A according to the first character image of character A, and generate character image set B according to the first character image of character B, so as to facilitate the subsequent lip-sync conversion processing of different speaking characters.
[0076] It should be understood that a person who is not speaking in the framed image is not counted as a target person. For example, when in the framed image, person A is in a speaking state and person B is in a non-speaking state, at this time, person A is the target person. Person B is not the target person. Therefore, the lip conversion system can also detect the speaking state of the person in the video to be converted, and update the number of target persons according to the speaking state. For example, in the above example, when person B is converted from a non-speaking state to a speaking state, both person A and person B in the framed image are in a speaking state. At this time, person B is the target person in the current framed image, and at this time, the number of target persons changes from 1 to 2.
[0077] S200: performing language conversion on the audio to be converted according to a second language to obtain a target audio.
[0078] After obtaining the video to be converted, the lip conversion system can extract the audio to be converted in the video to be converted, so as to perform subsequent processing on the audio to be converted. After completing the extraction of the audio to be converted, the lip conversion system can perform language conversion according to the second language or the audio to be converted according to the language conversion instruction input by the user to obtain the target audio, wherein the second language is different from the first language. For example, when the audio to be converted is in Chinese, a second language different from English can be used to perform language conversion on the audio to be converted, for example, English, Russian, Spanish or Portuguese. Taking the second language as English as an example, when the audio to be converted is the voice audio of "Hello", the target audio after language conversion is the voice audio of "Hello". After the audio to be converted is converted in language, the target audio obtained retains the voice characteristics of the target person, but the language information is switched from the first language to the second language.
[0079] In some embodiments, the audio to be converted is the speech audio spoken by the target person. After the lip-sync conversion system extracts the audio to be converted, in order to reduce the interference of other audio in the video to be converted on the speech audio, the lip-sync conversion system can perform noise reduction processing on the audio to be converted to reduce the influence of other audio on the audio to be converted.
[0080] S300: Inputting the target audio and the first character image into a lip-sync conversion model to obtain a second character image output by the lip-sync conversion model.
[0081] After the lip-syncing system performs language conversion on the audio to be converted to obtain the target audio, at this time, due to the change in language, the first lip movement of the target person speaking in the first language is different from the pronunciation movement of the target audio. Therefore, when the user watches the video to be converted, the lip movement of the target person in the video to be converted will not match the pronunciation lip movement of the target audio.
[0082] To this end, the lip-sync conversion system can input the target audio and the first character image into the lip-sync conversion model, and the lip-sync conversion model can extract the audio features of the target audio, and according to the audio features, switch the first lip-sync action of the target character in the first character image to the second lip-sync action to output the second character image. The lip-sync conversion model is obtained by executing a specific training process based on training data, and the training data includes at least one training character image with lip-sync actions. The lip-sync conversion model includes a transformation module, and the transformation module calculates the transformation coefficient corresponding to each feature channel in the image features of the first character image based on at least the target audio. The transformation module can perform feature transformation on each feature channel of the first character image according to the calculated transformation coefficient, thereby obtaining the image features of the first character image.
[0083] The second character image includes a second lip movement of the target character speaking in a second language. The second lip movement is the same as the pronunciation lip movement of the target audio, so that the character's lip movement viewed by the user corresponds to the target audio after language conversion, thereby improving the user's viewing experience.
[0084] S400: Generate a target video according to the target audio and the second person image.
[0085] After obtaining the second character image, the lip shape conversion system can replace the audio to be converted in the video to be converted according to the target audio after the language conversion, and replace the first character image in the video to be converted according to the second character image after the lip shape movement conversion, so that after the language conversion of the video to be converted, the lip shape of the target character is synchronously converted, so that after the language conversion of the video to be converted, the lip shape of the character in the picture corresponds to the pronunciation of the language.
[0086] In some embodiments, since only the lip movement area of the target person in the first person image is converted, the lip conversion system only replaces the lip movement area of the target person when replacing the first person image with the second person image, so that the overall facial area of the target person remains unchanged.
[0087] In order to output the second character image according to the first character image and the target audio, the lip-sync conversion system needs to perform a specific training process on the lip-sync conversion model according to the training data containing the target character to obtain the lip-sync conversion model.
[0088] In some embodiments, the lip-to-lip conversion system may obtain training data for training the lip-to-lip conversion model, and the training data is video data. The training data may include a first training image, and the first training image includes at least one training character with lip movements. The training character may be a target character for performing lip-to-lip conversion during the application process. For example, the training data includes a character image of character A. After the lip-to-lip conversion capability of the lip-to-lip conversion model is trained by the training data, the lip-to-lip conversion system may perform lip-to-lip movement conversion on character A during the application process.
[0089] The training character may also be a character other than the target character. For example, the training data includes a character image of character A. After the lip-syncing capability of the lip-syncing model is trained by the training data, the lip-syncing system can perform lip-syncing action conversion on character B different from character A during application.
[0090] like Figure 2 As shown, the lip-to-lip conversion system can input training data into the model to be trained, and the model to be trained is a lip-to-lip conversion model that has not been trained to convergence. The model to be converted has the same model structure as the lip-to-lip conversion model. After receiving the training data, the model to be trained can extract the training audio features of the training data, so as to determine the pronunciation lip shape according to the training audio features. The model to be trained can also extract the first training image of the training data. Since the training data is video data containing the target person, the first training image contains the facial image of the target person.
[0091] The model to be trained can perform feature channel conversion on the first training image based on the training audio features to generate a second training image. In the second training image, the target person makes the same lip movements as the pronunciation of the training audio features. Since the parameters of the model to be trained have not been trained to convergence, the second training image generated by the model to be trained should have a certain training loss. For this reason, the lip shape conversion system can calculate the training loss of the second training image and the lip shape image label through a loss function, wherein the lip shape image label is used to characterize the pronunciation lip shape movement of the training audio features, so that the training loss determines whether the second training image meets the output standard.
[0092] In order to determine whether the second training image meets the output standard, the lip conversion system can set a loss threshold. If the training loss is greater than the loss threshold, it means that the training loss between the second training image and the lip image label is large, and the second training image cannot be output. Therefore, the lip conversion system can iteratively train the training model to update the model parameters of the model to be trained and regenerate the second training image. When the training loss is less than or equal to the loss threshold, it means that the second training image meets the output standard, and the lip conversion system can obtain the lip conversion model based on the model parameters of the model to be trained.
[0093] In some embodiments, considering the diversity of languages in the training data, the model to be trained can perform speech extraction on the audio part of the training data through the wav2vec model to obtain training audio features. The wav2vec model can perform sequence stratification on the audio part of the training data and extract the audio features of the middle layer of the audio sequence to obtain the training audio features, so that the obtained audio features have good generalization for different languages, different receiving devices, and different speakers.
[0094] The model to be trained may include an encoder, a transformation module, and a decoder. Figure 3 This is a training flow chart of the model to be trained in the embodiment of this application. Figure 3 After the training data is input into the model to be trained, the model to be trained extracts the training audio features and the first training image of the training data respectively, and encodes the training audio features through the encoder to obtain the training audio encoding, and encodes the first training image through the encoder to obtain the first training encoded image.
[0095] After the encoder completes encoding, the training audio code and the first training image code can be input into the transformation module, so that the transformation module performs lip conversion processing on the first training image code based on the training audio code. The transformation module is composed of a fully connected layer, and the transformation module can learn the transformation relationship between different images through the fully connected layer, that is, the first training image code is transformed into a desired lip conversion state according to the training audio code.
[0096] After the training audio code and the first training image code are input into the transformation module, the first training image code is the quasi-transformed picture code, and the transformation module can determine the feature channel calculation relationship corresponding to the first training image code according to the training audio code, thereby realizing the lip movement transformation of the training character in the first training image at the feature channel level. The feature channel calculation relationship corresponds to different transformation relationships of the feature channels, for example, the feature channel calculation relationship can include a rotation relationship, a translation relationship, and a scaling relationship.
[0097] In order to determine the characteristic channel calculation relationship, the lip conversion system can extract a reference image from the training data according to the first training image code. The reference image can be a character image that is the same as the training character in the first training image, which is used to show the image of the training character in the training data and serve as a reference image for obtaining the characteristic channel calculation relationship. For example, the reference image can be a frontal image of the training character. By using the reference image as a reference, the conversion relationship of different characteristic channels such as the position, facial angle or facial size of the first training image relative to the reference image can be determined, thereby determining the characteristic channel calculation relationship according to the conversion relationship of different characteristic channels, so as to facilitate the subsequent lip conversion of the first training image code.
[0098] In some embodiments, the reference image may also be a plurality of images of a person at different angles, so as to determine the characteristic channel calculation relationship from different angle dimensions. For example, the reference image may include a first reference image, a second reference image, and a third reference image at different angles.
[0099] After obtaining the reference image, in order to improve the accuracy of the feature channel calculation relationship, such as Figure 4 As shown, the reference image and the first training image can be spliced to obtain a training spliced image, and the transformation coefficients can be calculated based on the spliced training image and the training audio code to determine the characteristic channel calculation relationship based on the training audio features. Wherein, when performing lip conversion on the first training image code, the transformation module can perform corresponding conversion calculations on the characteristic channels of the first training image code according to the transformation coefficients to obtain the second training image code, where the transformation coefficients are used to indicate the transformation relationship of the first training image relative to the reference image.
[0100] Since the reference image and the first training image are different images, after performing image stitching, feature stitching errors may occur in the training stitched image.
[0101] In order to eliminate feature splicing errors, in some embodiments, such as Figure 5 As shown, the encoder may include an alignment encoder configured to perform image registration on the training stitched images. Image registration is the process of stitching multiple images from different sources, finding matching points between different images and aligning the matching points spatially to minimize the desired error.
[0102] Based on this, the alignment encoder can perform feature alignment on the training spliced image according to the matching points between the first training image and the reference image during the splicing of the first training image and the reference image, and extract the alignment features of the training spliced image. Based on the alignment features, the transformation module can calculate the transformation coefficients according to the alignment features and the training audio features.
[0103] In some embodiments, corresponding to the conversion operations to be performed on different feature channels, the transform coefficient may include multiple sub-transform coefficients. Figure 6As shown, corresponding to the calculation of the rotation relationship of the feature channel, the calculation of the translation relationship of the feature channel and the calculation of the scaling relationship of the feature channel, the transformation coefficient may include a first transformation coefficient, a second transformation coefficient and a third transformation coefficient. Among them, the first transformation coefficient is a sub-transformation coefficient for calculating the rotation angle of the feature channel, and the transformation module can calculate the rotation angle of each feature channel of the first training image code according to the first transformation coefficient. The second transformation coefficient is a sub-transformation coefficient for calculating the translation distance of the feature channel, and the transformation module can calculate the translation distance of each feature channel of the first training image code according to the second transformation coefficient. The third transformation coefficient is a sub-transformation coefficient for calculating the scaling factor of the feature channel, and the transformation module can calculate the scaling factor of each feature channel of the first training image code according to the second transformation coefficient. After performing corresponding rotation, translation and scaling calculations on each feature channel of the first training image code, the first training image code can be converted into the second training image code according to the calculated rotation angle, translation distance and scaling factor.
[0104] It should be noted that the embodiment of the present application only exemplifies the three conversion calculation methods for feature channels, namely rotation, translation and scaling. In the actual training process, the conversion calculation methods for feature channels can be increased or decreased according to the lip conversion effect to be achieved by the lip conversion model. When adding conversion calculation methods, the transformation coefficients can also include more than three sub-transformation coefficients, for example, a fourth transformation coefficient, a fifth transformation coefficient, etc. The embodiment of the present application does not specifically limit the number of conversion calculation methods for feature channels.
[0105] The second training image code is an image code after the training character performs lip shape conversion according to the training audio features. The transformation module can output the second training image code to the decoder, so that the decoder can decode the second training image code to obtain the second training image and complete the lip shape conversion in the training process.
[0106] It can be seen from the above technical solution that this embodiment performs transformation on each feature channel during the lip shape conversion process, which reduces the computational load and improves the efficiency of lip shape conversion. In addition, since it does not regenerate the first character image in the application process, it can better retain the details of the character and the corresponding lip shape of the character in the video to be converted with lip shape masking, thereby completing the lip shape conversion and synthesis of the masked video. In the above process, what the transformation module learns is the transformation process with the training character, rather than training for a specific character. Therefore, in the process of lip shape conversion, it has good generalization and does not limit the character corresponding to the lip shape conversion, so it does not need to be targeted for training with a specific character.
[0107] In some embodiments, the encoder is further configured to perform attention calculation on the second training image encoding to perform attention encoding on the second training image. To this end, the encoder also includes a hierarchical structure formed by a preset number of encoding units, wherein the input of the encoding unit is the output of the previous encoding unit, and the output of the encoding unit is the input of the next encoding unit. The input end of the encoder can also be connected to the output end of the transformation module, that is, the input of the encoder is the second training image encoding output by the transformation module.
[0108] In order to improve the model fitting ability and reduce the overall calculation amount of the model, before the encoder performs attention calculation on the second training image encoding, the lip conversion system can divide the second training image encoding into multiple first image segments according to the calculation window, and each first image segment includes a certain number of pixels. The lip conversion system can input the first image segment into the encoder to independently perform attention calculation on each first image segment through a certain encoding unit to obtain a first attention encoding result.
[0109] However, in order to fully learn the features at different positions in the first feature code, after completing an attention calculation, the lip conversion system needs to adjust the first image segment in a preset manner, so as to re-divide the second training image code into a preset number of second image segments. The first image segment may partially overlap with the corresponding second image segment, and the number of the first image segment and the second image segment may also be different depending on the size of the image division.
[0110] After completing the division, the encoding unit can input the first attention encoding result into the unit of the next level corresponding to the current level, so as to perform attention calculation on the second image segment through the encoding unit based on the first attention encoding result to obtain the second feature encoding result. The encoder iterates the above attention calculation operation, and performs iterative attention calculation based on the second attention encoding result and the number of layers of the remaining encoding units, thereby outputting the second feature code. It should be understood that when the attention calculation is performed each time, the image segment needs to be re-divided based on the first feature code. The division method can be divided according to preset rules, such as division angle, image shape, image area, etc., or the image can be divided randomly, and it should be ensured that each pixel in the first feature code participates in the attention calculation. For example, in the first layer encoding unit, the window division is positioned as the initial window, and in the next layer encoding unit, the window division is moved one pixel to the right and downward based on the initial window. In this way, with the deepening of the layers, each pixel can be taken into account in the calculation of different windows, so that the encoder can simultaneously realize the combination of local information and global information encoded in the second training image during the calculation of the attention mechanism, thereby retaining detailed local information and global context in the process of generating the second training image.
[0111] In some embodiments, a hierarchical design is adopted between the encoding units, wherein each level gradually reduces the resolution of the first training image while increasing the dimension of the features.
[0112] In some embodiments, the decoder may include a preset number of upsampling layers, and the upsampling layers gradually increase the image space dimensions corresponding to the arrangement order, while the number of feature channels gradually decreases. After obtaining the second feature code, the second feature code may be input into the decoder, and a decoding operation is performed on the second feature code according to the upsampling layer, thereby obtaining a second training image.
[0113] In some embodiments, the feature dimensions of the upsampling layer and the encoding unit increase with the increase of the level, and the number of feature channels of the upsampling layer and the encoding unit decreases with the increase of the level.
[0114] Since the coding unit and the upsampling layer are hierarchical structures, and the number of feature channels is gradually reduced as the level increases, while the dimension of the feature increases. To this end, the coding unit in the coding module can be distributed with the upsampling layer in the decoder according to the corresponding structure, that is, the upsampling layer and the corresponding level of the coding unit are connected. For example, the coding unit at the first level is connected to the upsampling layer at the first level, the coding unit at the second level is connected to the upsampling layer at the second level, and so on.
[0115] It should be understood that in the above embodiments, the hierarchical structure formed by the coding units is described as a whole. In the actual encoding process, the actual output of each layer of coding units is the attention coding result obtained by performing attention calculation on the image segment. As the coding unit level goes deeper, the encoder will perform coding calculations on each pixel in the first training image, thereby summarizing the attention coding results into the first training image coding. Correspondingly, the input of the upsampling layer is the attention coding result after the feature channel conversion is performed. After completing the decoding operation of all the attention coding results, the decoder outputs the second training image after the feature channel conversion is performed to complete the lip shape conversion during the training process.
[0116] In order to facilitate the execution of the character processing method for video conversion described in the above embodiments, some embodiments of the present application also provide a character processing system for video conversion, the system comprising: a controller and a lip conversion model, the lip conversion model comprising a transformation module, the transformation module being configured to perform feature transformation on an image feature channel of training data input to the lip conversion model during the training process of the lip conversion model, the training data comprising at least one training character image with lip movements; the controller being configured to:
[0117] S100: Acquire a video to be converted, where the video to be converted includes audio to be converted and a first person image.
[0118] The language information of the audio to be converted is a first language, and the first character image includes a first lip movement of a target character speaking in the first language.
[0119] S200: performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language;
[0120] S300: Inputting the target audio and the first character image into a lip-sync conversion model to obtain a second character image output by the lip-sync conversion model.
[0121] The second character image includes a second lip movement of the target character speaking in the second language.
[0122] S400: Generate a target video according to the target audio and the second person image.
[0123] As can be seen from the above technical solutions, the present application provides a character processing method and system for video conversion, wherein the method obtains a video to be converted, performs language conversion on the audio to be converted of the video to be converted according to a second language, inputs the target audio and the first character image into a lip conversion model, obtains the second character image output by the lip conversion model, and then generates a target video according to the target audio and the second character image. The present application adjusts the lip movements in the first character image through the translated audio data, so that the lip movements of the target character in the target video correspond to the pronunciation of the translated target audio, thereby improving the user experience.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0125] For the convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are intended to better explain the present disclosure, so that those skilled in the art can better use the embodiments.
Claims
1. A character processing method for video conversion, characterized in that: include: Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image at least includes a first lip movement of a target person speaking in the first language; Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language; Inputting the target audio and the first character image into a lip-sync conversion model to obtain a second character image output by the lip-sync conversion model, wherein the second character image at least includes a second lip movement of the target character speaking in the second language; the lip-sync conversion model includes a transformation module, wherein the transformation module is configured to calculate a transformation coefficient corresponding to each feature channel in the image feature of the first character image at least according to the target audio, and perform feature transformation on the image feature channel of the first character image according to the transformation coefficient to obtain the image feature of the first character image; A target video is generated according to the target audio and the second person image.
2. The character processing method for video conversion according to claim 1, characterized in that: Before the step of inputting the target audio and the first character image into the lip-sync conversion model, the method further includes: Acquire training data; the training data includes a first training image, the first training image includes at least one training character with lip movements; Inputting the training data into a model to be trained, so as to extract training audio features and a first training image of the training data through the model to be trained, wherein the model to be trained is a lip-sync model that has not been trained to convergence; Based on the training audio features, performing feature channel conversion on the first training image through the model to be trained to generate a second training image; Calculate the training loss of the second training image and the lip shape image label according to the loss function; If the training loss is greater than the loss threshold, iteratively training the model to be trained; If the training loss is less than or equal to the loss threshold, the lip conversion model is input according to the model parameters of the model to be trained.
3. The character processing method for video conversion according to claim 2, characterized in that: The model to be trained includes an encoder, a transformation module and a decoder. Based on the training audio features, the step of performing feature channel conversion on the first training image through the model to be trained includes: Encoding the training audio features by the encoder to obtain a training audio code, and encoding the first training image by the encoder to obtain a first training image code; Determining, by the transformation module, a feature channel calculation relationship at least according to the training audio code and the first training image code; According to the feature channel calculation relationship, performing feature channel conversion on the first training image code to obtain a second training image code; The decoder decodes the second training image to obtain a second training image.
4. The character processing method for video conversion according to claim 3 is characterized in that: The step of determining the characteristic channel calculation relationship at least according to the training audio code and the first training image code by the transformation module comprises: Obtain a reference image according to the first training image encoding, wherein the reference image is an image of a person identical to the training person in the first training image; Splicing the reference image with the first training image to obtain a training spliced image; The transform coefficients are calculated according to the training audio code and the spliced training image to determine the feature channel calculation relationship, wherein the transform coefficients are used to indicate the transformation relationship of the first training image relative to the reference image.
5. The character processing method for video conversion according to claim 4 is characterized in that: The step of calculating the transformation coefficient according to the spliced training image comprises: extracting alignment features of the spliced training image by the encoder; Transform coefficients are calculated based on the aligned features and the training audio features.
6. The character processing method for video conversion according to claim 5, characterized in that: The transform coefficients include a first transform coefficient, a second transform coefficient and a third transform coefficient, and the step of performing feature channel conversion on the first training image encoding according to the feature channel calculation relationship includes: Calculating a rotation angle of a feature channel encoded by the first training image according to the first transform coefficient; Calculating the translation distance of the feature channel of the first training image encoding according to the second transform coefficient; Calculating a scaling factor of a feature channel of the first training image encoding according to the third transform coefficient; The first training image code is converted into a second training image code according to the rotation angle, the translation distance, and the scaling factor.
7. The character processing method for video conversion according to claim 3 is characterized in that: The encoder includes a hierarchical structure formed by a preset number of encoding units, the input of the encoding unit is the output of the previous encoding unit, and the output of the encoding unit is the input of the next encoding unit. The step of encoding and decoding the second training image by the decoder includes: Dividing the second training image into a preset number of first image segments by encoding; Inputting the first image segment into the encoder, so as to independently perform attention calculation on each of the first image segments through a certain encoding unit to obtain a first attention encoding result; Adjusting the first image segment in a preset manner to divide the second training image encoding into a preset number of second image segments; wherein there is a partial overlap between the first image segment and the corresponding second image segment; Inputting the first attention encoding result into an encoding unit at a next level, so that the encoding unit performs attention calculation on the second image segment to obtain a second attention encoding result; Iterate the above operation to obtain a second feature encoding according to the second attention encoding result; The second feature code is decoded to obtain a second training image.
8. The character processing method for video conversion according to claim 7, characterized in that: The decoder comprises a preset number of upsampling layers; The step of decoding the second feature code to obtain a second training image includes: Inputting the second feature code into the decoder, and performing a decoding operation on the second feature code according to the upsampling layer to obtain the second training image; Wherein, a preset number of coding units in the coding module and the upsampling layer are distributed according to a corresponding structure; In the decoder, the second feature code of the decoding operation performed by the upsampling layer is output by a non-corresponding encoding unit.
9. The character processing method for video conversion according to claim 8, characterized in that: The feature dimensions of the upsampling layer and the coding unit increase with the increase of the level, and the number of feature channels of the upsampling layer and the coding unit decreases with the increase of the level; The encoding unit and the upsampling layer are connected at corresponding levels.
10. A character processing system for video conversion, characterized in that: include: A controller and a lip conversion model, wherein the lip conversion model includes a transformation module, wherein the transformation module is configured to calculate a transformation coefficient corresponding to each feature channel in the image feature of the first person image at least according to the target audio, and perform feature transformation on the image feature channel of the first person image according to the transformation coefficient to obtain the image feature of the first person image; the controller is configured to: Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image at least includes a first lip movement of a target person speaking in the first language; Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language; Inputting the target audio and the first character image into a lip-sync model to obtain a second character image output by the lip-sync model, wherein at least the second character image includes a second lip movement of the target character speaking in the second language; A target video is generated according to the target audio and the second person image.
Citation Information
Patent Citations
Image processing device and image processing method
CN110169072A
Video translation method, system and device and storage medium
CN112562721A
Video generation method and device, equipment and storage medium
CN112752118A
Voice lip reading method and system based on multi-scale video feature fusion
CN113450824A
Audio-driven figure mouth shape method, model and training method thereof
CN115035604A