Figure video processing method and system

By introducing a lip-type conversion model into language conversion technology, the problem of the lip-type and language in the film after language conversion is solved, and a better viewing experience is achieved.

CN119967224APending Publication Date: 2025-05-09NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202411589167.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-08
Filing Date
2024-11-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

After the language changes, the lip movements of the characters in the film do not correspond to the converted lines, which reduces the user's viewing experience.

Method used

By obtaining the audio and images in the video to be converted, performing language conversion, and inputting the converted audio and original image into the lip-type conversion model to generate synchronous lip-type actions to ensure that the lip-type corresponds to the language pronunciation.

Benefits of technology

After the language conversion is realized, the lip movements of the characters in the film correspond to the language pronunciation, which improves the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967224A_ABST
    Figure CN119967224A_ABST
Patent Text Reader

Abstract

The invention provides a character video processing method and system, and the method comprises the steps: obtaining a to-be-converted video, carrying out the language conversion of a to-be-converted audio in the to-be-converted video, and obtaining a target audio. Inputting a first figure image extracted from the target audio and the to-be-converted video into a mouth shape conversion model, and converting a first mouth shape action of the first figure image according to the target audio to obtain a second figure image; and generating a target video by using the target audio and the second figure image. The mouth shape action in the first figure image is adjusted through the translated audio data, so that the mouth shape action of the target figure in the target video corresponds to the pronunciation of the translated target audio, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to a Chinese patent application filed with the State Intellectual Property Office with application number 202311479579.X on November 8, 2023. The entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the technical field of voice conversion, and in particular to a method and system for processing character videos. Background Art

[0003] When users watch movies from different countries, they cannot directly understand the meaning of the lines from the audio of the lines spoken by the characters in the movie because the language of the movie is different from the user's language. Therefore, users need to watch subtitles continuously while watching the movie to understand the meaning of the lines of the characters in the movie.

[0004] To this end, voice conversion technology can be used to convert the language of the film into the language desired by the user while retaining the voice characteristics of the characters in the film, thereby achieving the effect of real-time translation. In the above process, since the language of the characters' lines has changed, the original speaking lip shape of the characters in the film cannot correspond to the converted lines, which reduces the user's viewing experience. Summary of the invention

[0005] In order to solve the problem that the mouth shape of the character in the picture does not correspond to the pronunciation of the language after language conversion, some embodiments of the present application provide a method and system for processing character videos.

[0006] In a first aspect, some embodiments of the present application provide a method for processing a character video, comprising:

[0007] Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image at least includes a first lip movement of a target person speaking in the first language;

[0008] Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language;

[0009] Inputting the target audio and the first person image into a lip-sync conversion model to obtain a second person image output by the lip-sync conversion model, wherein the second person image includes a second lip movement of the target person speaking in the second language; the lip-sync conversion model is trained by training data including the target person;

[0010] A target video is generated according to the target audio and the second person image.

[0011] In some embodiments, before the step of inputting the target audio and the first character image into the lip-sync model, the method further includes:

[0012] Acquire training data, where the training data is video data including the target person;

[0013] Inputting the training data into a model to be trained, so as to extract training audio features of the training data and a first training image of a target person through the model to be trained, wherein the model to be trained is a lip-sync conversion model that has not been trained to convergence;

[0014] Obtain a second training image generated by the to-be-trained model according to the training audio features and the first training image;

[0015] Calculating the training loss of the second training image and the lip image label according to the loss function, wherein the lip image label is used to represent the pronunciation lip movement of the training audio feature;

[0016] If the training loss is greater than the loss threshold, iteratively training the model to be trained;

[0017] If the training loss is less than or equal to the loss threshold, a lip-sync model is obtained according to the model parameters of the model to be trained.

[0018] In some embodiments, the model to be trained includes an encoder and a decoder, wherein the encoder includes a first encoding module and a second encoding module, the first encoding module includes a preset number of convolutional layers, and the step of obtaining a second training image generated by the model to be trained according to the training audio features and the first training image includes:

[0019] splicing the training audio features and the first training image to obtain an audio training image;

[0020] Input the audio training image into the first encoding module to traverse the pixel points in the audio training image through a convolution layer, wherein the pixel points include a first pixel point and a second pixel point, wherein the first pixel point is a valid pixel point in the target person's lip shape area, and the second pixel point is an invalid pixel point in the target person's lip shape area, and the invalid pixel point is used to indicate a pixel point that is covered in the target person's lip shape area;

[0021] Performing a convolution operation on a valid area in the audio training image through the convolution layer to obtain a lip shape image feature, wherein the valid area is an image area formed by the first pixel points;

[0022] Performing an iterative convolution operation on the lip-shaped image feature through the convolution layer to obtain a first feature code;

[0023] Performing attention encoding on the first feature code by the second encoding module, and outputting a second feature code;

[0024] The second feature code is decoded by the decoder to obtain a second training image.

[0025] In some embodiments, after the step of performing an iterative convolution operation on the lip-sync image features through the convolution layer, the method further comprises:

[0026] During the iterative convolution operation, a newly added valid area is obtained according to a preset convolution window, wherein the preset convolution window is a convolution window that moves periodically in the audio training image during the convolution iteration process;

[0027] Updating the valid area according to the newly added valid area;

[0028] A convolution operation is performed based on the updated valid area to update the lip shape image features.

[0029] In some embodiments, the step of acquiring a newly added valid area according to a preset convolution window includes:

[0030] Traversing the first pixel point in the preset convolution window;

[0031] If the preset convolution window includes at least one first pixel point, all the pixels in the preset convolution window are marked as first pixels to obtain a newly added effective area.

[0032] In some embodiments, the second encoding unit includes a hierarchical structure formed by a preset number of encoding units, the input of the encoding unit is the output of the previous encoding unit, and the output of the encoding unit is the input of the next encoding unit; the step of performing attention encoding on the first feature encoding by the second encoding module includes:

[0033] Dividing the first feature code into a preset number of first image segments;

[0034] Inputting the first image segment into a coding unit in the second coding module, so that the coding unit independently performs attention calculation on each of the first image segments to obtain a first attention coding result;

[0035] Adjusting the first image segment in a preset manner to divide the first feature code into a preset number of second image segments; wherein there is a partial overlap between the first image segment and the corresponding second image segment;

[0036] Inputting the first attention encoding result into an encoding unit at a next level, so that the encoding unit performs attention calculation on the second image segment to obtain a second attention encoding result;

[0037] Iterate the above operation to obtain the second feature encoding according to the second attention encoding result.

[0038] In some embodiments, the second encoding unit includes a hierarchical structure formed by a preset number of encoding units, the input of the encoding unit is the output of the previous encoding unit, and the output of the encoding unit is the input of the next encoding unit; after the step of obtaining the audio training image, the method further includes:

[0039] Inputting the audio training image into the second encoding module to divide the audio training image into a preset number of first image segments;

[0040] Inputting the first image segment into a coding unit in the second coding module, so that the coding unit independently performs attention calculation on each of the first image segments to obtain a first attention coding result;

[0041] Adjusting the first image segment in a preset manner to divide the first feature code into a preset number of second image segments; wherein there is a partial overlap between the first image segment and the corresponding second image segment;

[0042] Inputting the first attention encoding result into an encoding unit at a next level, so that the encoding unit performs attention calculation on the second image segment to obtain a second attention encoding result;

[0043] Iterate the above operation to obtain the second feature encoding according to the second attention encoding result.

[0044] In some embodiments, the decoder includes a preset number of upsampling layers; and the step of decoding the second feature code by the decoder to obtain a second training image includes:

[0045] Inputting the second feature code into the decoder, and performing a decoding operation on the second feature code according to the upsampling layer to obtain the second training image;

[0046] Wherein, a preset number of convolutional layers in the first encoding module and / or a preset number of encoding units in the second encoding module and the upsampling layer are distributed according to a corresponding structure;

[0047] In the decoder, the second feature code of the decoding operation performed by the upsampling layer is output by the non-corresponding convolutional layer and the encoding unit, or the second feature code of the decoding operation performed by the upsampling layer is directly output by the encoding unit.

[0048] In some embodiments, the convolution layer, the encoding unit, and the upsampling layer each include a preset number of feature channels, wherein the hierarchical structure of the convolution layer, the encoding unit, and the upsampling layer is positively correlated with the number of the feature channels;

[0049] The convolution layer is connected to the upsampling layer at corresponding levels;

[0050] The encoding unit is connected to the upsampling layer at a corresponding level.

[0051] In a second aspect, some embodiments of the present application provide a character video processing system, including: a controller and a lip conversion model, wherein the lip conversion model is trained by training data containing a target character; the controller is configured to:

[0052] Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image at least includes a first lip movement of a target person speaking in the first language;

[0053] Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language;

[0054] Inputting the target audio and the first character image into a lip-sync model to obtain a second character image output by the lip-sync model, wherein the second character image includes a second lip movement of the target character speaking in the second language;

[0055] A target video is generated according to the target audio and the second person image.

[0056] It can be seen from the above technical solutions that the present application provides a method and system for processing character videos. The method obtains a video to be converted, performs language conversion on the audio to be converted in the video to be converted, and obtains a target audio. The first character image extracted from the target audio and the video to be converted is input into a lip conversion model to convert the first lip movement of the first character image according to the target audio to obtain a second character image; the target audio and the second character image are used to generate a target video. The present application adjusts the lip movement in the first character image through the translated audio data, so that the lip movement of the target character in the target video corresponds to the pronunciation of the translated target audio, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0058] Figure 1 A flowchart of a method for processing a character video provided in an embodiment of the present application;

[0059] Figure 2 A flowchart of performing training on a lip-sync conversion model according to an embodiment of the present application;

[0060] Figure 3 A flowchart of obtaining a second training image during the training process of the lip shape conversion model in an embodiment of the present application;

[0061] Figure 4 This is a flow chart of traversing the pixels of an audio training image according to an embodiment of the present application;

[0062] Figure 5 A flowchart of performing a convolution operation on a valid area according to an embodiment of the present application;

[0063] Figure 6 A flowchart of another embodiment of the lip shape conversion model of the present application embodiment obtaining a second training image during the training process;

[0064] Figure 7 This is a schematic diagram of the connection relationship between the first encoding module and the decoder in the embodiment of the present application;

[0065] Figure 8 Schematic diagram of the connection relationship between the second encoding module and the decoder in the embodiment of the present application. DETAILED DESCRIPTION

[0066] In order to make the purpose and implementation method of the present application clearer, the exemplary implementation method of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0067] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0068] The terms "first", "second", "third", etc. in the specification and the above drawings of this application are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise specified. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.

[0069] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0070] Users can play videos on smart devices, such as movies, TV dramas, operas or variety shows, etc. When users watch movies from different countries, they cannot directly understand the semantics of the lines spoken by the characters in the movie because the language of the movie is different from the user's native language. For example, when a Chinese user watches an English movie, if the user cannot understand English, he cannot directly know the semantics of the lines based on the audio of the English lines spoken by the characters in the movie.

[0071] To this end, some movie resources may include subtitle resources, that is, when playing the movie screen, subtitles in the corresponding language are displayed in the movie screen. For example, when playing a movie in English, Chinese subtitles are added to the movie screen so that Chinese users can understand the meaning of the lines in the movie through Chinese subtitles. This requires users to watch the screen and the subtitles while watching the movie so that they can know the semantics of the characters' lines in the movie. Although the subtitles are located in the movie screen, repeatedly watching content located in different positions affects the user's viewing experience.

[0072] In order to improve the viewing experience of the film, the lines in the film can be converted into different languages ​​before post-dubbing, so that users can intuitively understand the semantics of the lines according to the dubbing. However, this post-dubbing method consumes a lot of time and needs to be completed before playing the film, and it is impossible to translate and play the lines in real time.

[0073] To this end, voice conversion technology can be used to convert the language of the film into the language expected by the user while retaining the voice characteristics of the characters in the film, thereby achieving the effect of real-time translation. In the above process, since the language of the lines spoken by the characters has changed, the film maintains the original speaking lip shape of the characters and cannot correspond to the converted lines. For example, when the lines audio is in the first language, the film will play the lines audio in the first language based on the preset sound characteristics, and correspondingly, the lip shape movements of the characters in the film correspond to the pronunciation lip shape of the lines audio in the first language. After real-time translation, the lines audio is converted from the first language to the second language. Due to the different pronunciations between different languages, the lip shape movements of the characters in the film do not correspond to the lines audio in the second language, resulting in the characters' lip shape movements seen by the user not matching the actual audio pronunciation heard, affecting the user's viewing experience.

[0074] In order to solve the problem that the mouth shape of the character in the picture does not correspond to the pronunciation of the language after language conversion, some embodiments of the present application provide a method for processing a character video. Figure 1 This is a flowchart of a method for processing a person video provided in an embodiment of the present application. The method can be applied to a display device for playing audio and video or to a person video processing system. For example, the application object is a person video processing system. Figure 1 , the method comprising:

[0075] S100: Obtain the video to be converted.

[0076] The lip conversion system can obtain the video to be converted input by the user. The video to be converted includes the audio to be converted and the first character image. The audio to be converted is the original audio of the video to be converted, and is also the audio data for language conversion to be performed. The language information of the audio to be converted is the first language. The first language is the language of the audio to be converted before the translation is performed. For example, if the user is playing an English movie through a display device, the first language of the English movie is English.

[0077] In some embodiments, the video to be converted may also be a dubbed film, that is, the language of the audio to be converted is different from the original language, in which case the first language is the dubbed language. For example, if the video to be converted is an English film dubbed from Chinese, the first language of the audio to be converted is Chinese.

[0078] The first person image is a person who is speaking in the video to be converted, and the lip conversion system can perform frame processing on the video to be converted to obtain the first person image from the framed images. The first person image at least includes a first lip movement of the target person speaking in the first language, wherein the target person is a person who is speaking in the first person image.

[0079] In some embodiments, the first character image may also include facial images of other characters. For example, the front of the target character and the front images of other characters in the video to be converted may be simultaneously displayed in the video frame of the video to be converted. In the subsequent process of performing video conversion, the lip movements of other characters may be converted synchronously. In addition to facial images or lip movements, the first character image may also include information such as body movements, postures, and expressions of the target character, so that in the process of performing video conversion, the body movements, postures, and expressions are kept unchanged, or the action postures of the target character in the video to be converted are changed according to preset action information.

[0080] In some embodiments, the lip-sync conversion system can detect the number of target characters in the split-frame picture to collect the first character image according to the number of target characters. When the split-frame picture includes only one target character, the character image of the target character in the split-frame picture when in a speaking state can be used as the first character image. When the split-frame picture includes multiple target characters, the first character images of different target characters in a speaking state can be collected in the split-frame picture respectively, and a character image set can be generated according to the first character images of the corresponding target characters. For example, when the split-frame picture includes character A and character B who are in a speaking state at the same time, at this time, the lip-sync conversion system can obtain the first character image of character A and the first character image of character B respectively, and generate character image set A according to the first character image of character A, and generate character image set B according to the first character image of character B, so as to facilitate the subsequent lip-sync conversion processing of different speaking characters.

[0081] It should be understood that a person who is not speaking in the framed image is not counted as a target person. For example, when in the framed image, person A is in a speaking state and person B is in a non-speaking state, at this time, person A is the target person. Person B is not the target person. Therefore, the lip conversion system can also detect the speaking state of the person in the video to be converted, and update the number of target persons according to the speaking state. For example, in the above example, when person B is converted from a non-speaking state to a speaking state, both person A and person B in the framed image are in a speaking state. At this time, person B is the target person in the current framed image, and at this time, the number of target persons changes from 1 to 2.

[0082] S200: performing language conversion on the audio to be converted according to a second language to obtain a target audio.

[0083] After obtaining the video to be converted, the lip conversion system can extract the audio to be converted in the video to be converted, so as to perform subsequent processing on the audio to be converted. After completing the extraction of the audio to be converted, the lip conversion system can perform language conversion according to the second language or the audio to be converted according to the language conversion instruction input by the user to obtain the target audio, wherein the second language is different from the first language. For example, when the audio to be converted is in Chinese, a second language different from English can be used to perform language conversion on the audio to be converted, for example, English, Russian, Spanish or Portuguese. Taking the second language as English as an example, when the audio to be converted is the voice audio of "Hello", the target audio after language conversion is the voice audio of "Hello". After the audio to be converted is converted in language, the target audio obtained retains the voice characteristics of the target person, but the language information is switched from the first language to the second language.

[0084] In some embodiments, the audio to be converted is the speech audio spoken by the target person. After the lip-sync conversion system extracts the audio to be converted, in order to reduce the interference of other audio in the video to be converted on the speech audio, the lip-sync conversion system can perform noise reduction processing on the audio to be converted to reduce the influence of other audio on the audio to be converted.

[0085] S300: Inputting the target audio and the first character image into a lip-sync conversion model to obtain a second character image output by the lip-sync conversion model.

[0086] After the lip-syncing system performs language conversion on the audio to be converted to obtain the target audio, at this time, due to the change in language, the first lip movement of the target person speaking in the first language is different from the pronunciation movement of the target audio. Therefore, when the user watches the video to be converted, the lip movement of the target person in the video to be converted will not match the pronunciation lip movement of the target audio.

[0087] To this end, the lip conversion system can input the target audio and the first character image into the lip conversion model, and the lip conversion model can extract the audio features of the target audio, and according to the audio features, switch the first lip movement of the target character in the first character image to the second lip movement to output the second character image. The lip conversion model is trained based on the training data containing the target character.

[0088] The second character image includes a second lip movement of the target character speaking in a second language. The second lip movement is the same as the pronunciation lip movement of the target audio, so that the character's lip movement viewed by the user corresponds to the target audio after language conversion, thereby improving the user's viewing experience.

[0089] S400: Generate a target video according to the target audio and the second person image.

[0090] After obtaining the second character image, the lip shape conversion system can replace the audio to be converted in the video to be converted according to the target audio after the language conversion, and replace the first character image in the video to be converted according to the second character image after the lip shape movement conversion, so that after the language conversion of the video to be converted, the lip shape of the target character is synchronously converted, so that after the language conversion of the video to be converted, the lip shape of the character in the picture corresponds to the pronunciation of the language.

[0091] In some embodiments, since only the lip movement area of ​​the target person in the first person image is converted, the lip conversion system only replaces the lip movement area of ​​the target person when replacing the first person image with the second person image, so that the overall facial area of ​​the target person remains unchanged.

[0092] In order to output the second character image according to the first character image and the target audio, the lip-sync conversion system needs to perform a specific training process on the lip-sync conversion model according to the training data containing the target character to obtain the lip-sync conversion model.

[0093] The lip-syncing system can obtain training data, wherein the training data is video data containing the target person. When the target person is a character in a film, the lip-syncing system can perform character recognition on the target person. If the target person is a public figure such as a film actor or singer, the lip-syncing system can search for video data containing the target person as training data based on the character recognition result, such as online course video data, performance video data, speech video data, film and television video data, etc.

[0094] In some embodiments, the video to be converted may also be a video recorded by the user containing the target person. In this case, the user may use other video data of the person being filmed as training data to train the lip-sync conversion model.

[0095] After obtaining the training data, such as Figure 2 As shown, the lip-to-lip conversion system can input training data into the model to be trained, and the model to be trained is a lip-to-lip conversion model that has not been trained to convergence. The model to be converted has the same model structure as the lip-to-lip conversion model. After receiving the training data, the model to be trained can extract the training audio features of the training data, so as to determine the pronunciation lip shape according to the training audio features. The model to be trained can also extract the first training image of the training data. Since the training data is video data containing the target person, the first training image contains the facial image of the target person.

[0096] The model to be trained can generate a second training image based on the training audio features and the first training image. In the second training image, the target person makes the same lip movements as the pronunciation of the training audio features. Since the parameters of the model to be trained have not been trained to convergence, the second training image generated by the model to be trained should have a certain training loss. For this reason, the lip shape conversion system can calculate the training loss of the second training image and the lip shape image label through a loss function, wherein the lip shape image label is used to characterize the pronunciation lip shape movement of the training audio features, so that the training loss determines whether the second training image meets the output standard.

[0097] In order to determine whether the second training image meets the output standard, the lip conversion system can set a loss threshold. If the training loss is greater than the loss threshold, it means that the training loss between the second training image and the lip image label is large, and the second training image cannot be output. Therefore, the lip conversion system can iteratively train the training model to update the model parameters of the model to be trained and regenerate the second training image. When the training loss is less than or equal to the loss threshold, it means that the second training image meets the output standard, and the lip conversion system can obtain the lip conversion model based on the model parameters of the model to be trained.

[0098] In some embodiments, considering the diversity of languages ​​in the training data, the model to be trained can perform speech extraction on the audio part of the training data through the wav2vec model to obtain training audio features. The wav2vec model can perform sequence stratification on the audio part of the training data and extract the audio features of the middle layer of the audio sequence to obtain the training audio features, so that the obtained audio features have good generalization for different languages, different receiving devices, and different speakers.

[0099] The model to be trained may include an encoder and a decoder. The encoder may encode the training audio features and the first training image through a convolution operation to extract the training audio features and the features in the first training image. The decoder decodes the encoding result obtained by the encoding to output the second training image to complete the lip shape conversion.

[0100] In order to facilitate the operation of lip conversion of the training audio features and the first training image, the lip conversion system can splice the training audio features with the first training image to obtain the audio training image, and input the audio training image into the encoder. The encoder can encode the audio training image, so as to convert the lip movements in the first training image according to the training audio features. After the encoding is completed, the encoded audio training image is input into the decoder to perform a deconvolution operation on the encoded audio training image to obtain the second training image.

[0101] In order to improve the coding accuracy, Figure 3 As shown, the encoder may include a first encoding module and a second encoding module, that is, the audio training image is encoded by two encoding methods in sequence. Among them, the audio training image is first input into the first encoding module, so that the audio training image is convolutionally encoded by the first encoding module, and the first feature code is output. The input end of the second encoding module is connected to the output end of the first encoding module, that is, the output of the first encoding module is the input of the second encoding module. Therefore, the first encoding module can input the first feature code into the second encoding module, so that the first feature code is performed by the second encoding module. Attention encoding is performed on the first feature code and the second feature code is output.

[0102] In some embodiments, the encoding process of the first encoding module is first described:

[0103] The first encoding module can be based on an improved implementation of the Convolution encoding module and is configured to perform partial convolution on the audio training image. The first encoding module includes a preset number of interconnected convolution layers, and the convolution layers are arranged in an order corresponding to the image feature dimensions gradually decreasing, while the number of feature channels gradually increasing. The encoding process of the first encoding module is described below:

[0104] The model to be trained can input the audio training image into the first encoding module to perform a convolution operation on the audio training image through the first convolution layer, thereby extracting the lip shape image features in the audio training image. Since the audio training image is obtained by splicing the training audio feature and the first training image, the lip shape image features are image features of the target person making the lip shape movements corresponding to the pronunciation of the training audio feature, and the lip shape image features correspond to the pronunciation lip shape images of the training audio feature.

[0105] After the first convolutional layer outputs the lip-type image features, the first encoding module can use the lip-type image features output by the convolutional layer as input and input them into the next convolutional layer to further extract deep-level features of the lip-type image features through the convolutional layer, that is, perform an iterative convolution operation on the lip-type image features to obtain the first feature encoding.

[0106] In some embodiments, if the lip area of ​​the target person is blocked in the training data, the first encoding module cannot perform convolution based on the blocked lip area to obtain effective lip image features, resulting in a significant deviation between the output first feature code and the actual lip image. To this end, the first encoding module can traverse the pixels in the audio training image through a convolution layer, the pixels include a first pixel and a second pixel, the first pixel is a valid pixel in the lip area of ​​the target person, the second pixel is an invalid pixel in the lip area of ​​the target person, and the invalid pixel is used to indicate the obscured pixel in the lip area of ​​the target person.

[0107] In order to distinguish the first pixel from the second pixel, the first encoding module may create an initial mask according to the audio training image to assign a first mask value to the first pixel in the audio training image, and assign a second mask value to the second pixel. For example, the first mask value is 1 and the second mask value is 2. Therefore, Figure 4 As shown, in the process of the convolution layer performing the convolution operation on the audio training image, the convolution output can be calculated only for the first pixel point, and the second pixel point does not participate in the convolution calculation, that is, the convolution operation is performed on the valid area, and the convolution operation is not performed on the area formed by the second pixel point. The valid area is the image area formed by the first pixel point.

[0108] In this embodiment, the calculation of the convolution output depends not only on the current input pixel value, but also on the mask to ensure that the convolution output only reflects the information of valid pixels.

[0109] During the convolution operation, the convolution layer creates different convolution windows to perform convolution calculations on different valid areas of the audio training image, thereby extracting lip image features at different locations.

[0110] Based on the above scenario, in this embodiment, the first encoding module can obtain a newly added valid area according to a preset convolution window during the iterative convolution operation, wherein the preset convolution window is a convolution window that periodically moves in the audio training image during the convolution iteration. During the iterative convolution, the position of the preset convolution window will change. Therefore, in the process of the first encoding module performing an iterative convolution operation on the valid area through the convolution layer, more second pixels will be replaced by first pixels according to the movement of the preset convolution window, and the area of ​​the first pixel will continue to expand, thereby generating the obscured part of the target person's lip shape, so as to solve the problem that the first encoding module cannot perform convolution according to the obscured lip shape area to obtain effective lip shape image features.

[0111] Since the first encoding module updates the number of first pixels each time it performs convolution on the effective area, after the first encoding module performs a convolution operation on the effective area, the newly added effective area can be updated according to the preset convolution window, and the effective area can be updated according to the newly added effective area. When the first encoding module performs the next convolution operation, the convolution operation can be performed on the updated effective area. Since the effective area increases, the first encoding module can obtain more lip-shaped image features when performing a convolution operation on the newly added effective area, thereby realizing the update of the lip-shaped image features.

[0112] Figure 5 Flowchart of performing convolution operation according to the convolution window in the embodiment of the present application. Figure 5 , this embodiment can traverse the first pixel point in the preset convolution window during a single convolution operation. If the mask value of at least one pixel point in a convolution window of the audio training image is 1, that is, at least one first pixel point is included in the preset convolution window, then all the pixels in the preset convolution window are marked as first pixels, that is, the mask values ​​of all the pixels in the convolution window are updated to 1 in the convolution output result at this position, and marked as a valid area. At this time, the second pixels of the convolution window are all replaced by the first pixels, and the area formed by these replaced pixels is the newly added valid area. On the contrary, if the mask values ​​of all pixels in a convolution window of the audio training image are 0, that is, all are second pixels in the convolution window, then the mask value corresponding to the convolution output result of the audio training image at this position is still 0.

[0113] It can be seen from the above technical solution that the masking process of the first encoding module can enable the first encoding module to dynamically adjust the weight according to the known part in the occluded area to adapt to different degrees of occlusion. In addition, when the convolution layer performs a convolution operation on the effective area, it can adapt to the size of the receptive field, improve the convolution layer's ability to understand the context of the audio training image, and better handle the occlusion problem in the target person's mouth shape area in the training data, thereby improving the performance of the model in complex environments.

[0114] The second encoding module can be implemented based on an improvement of the Transformer model, and the second encoding unit includes a hierarchical structure formed by a preset number of encoding units, wherein the input of the encoding unit is the output of the previous encoding unit, and the output of the encoding unit is the input of the next encoding unit.

[0115] The input of the second encoding module is the first feature code output by the first encoding module. The second encoding module can divide the first feature code into multiple first image segments, each of which includes a certain number of pixels. After obtaining the first image segment, the first image segment can be input into a coding unit in the second encoding module. The coding unit will independently perform attention calculation on each first image segment to obtain a first attention coding result. The coding unit input for the first time can be a coding unit at the first level of the hierarchical structure, or a coding unit at any level except the last level in the hierarchical structure.

[0116] The coding unit includes an attention mechanism layer and a feedforward network layer. The attention mechanism layer performs attention calculation on the pixels in the image segment, and the obtained attention coding result is output through the feedforward network layer. Since the coding unit is a hierarchical structure, after completing the attention calculation of the coding unit of the current level, it is necessary to input the obtained first attention coding result into the coding unit of the next level for attention calculation again. However, in order to fully learn the features at different positions in the first feature code, the second coding module needs to adjust the first image segment in a preset manner, so as to re-divide the first feature code into a preset number of second image segments. Among them, the first image segment may partially overlap with the corresponding second image segment, and the number of the first image segment and the second image segment may also be different depending on the size of the image division.

[0117] After completing the division, the second encoding module can input the first attention encoding result into the unit of the next level corresponding to the current level, so as to perform attention calculation on the second image segment through the encoding unit based on the first attention encoding result to obtain the second feature encoding result. The second encoding module iterates the above operation, and performs iterative attention calculation based on the second attention encoding result and the number of layers of the remaining encoding units, thereby outputting the second feature encoding. It should be understood that when the attention calculation is performed each time iteratively, the image segment needs to be re-divided based on the first feature coding. The division method can be divided according to preset rules, such as division angle, image shape, image area, etc., or the image can be divided randomly, and it should be ensured that each pixel in the first feature coding participates in the attention calculation. For example, in the first layer encoding unit, the window division is positioned as the initial window, and in the next layer encoding unit, the window division is moved one pixel to the right and downward based on the initial window. In this way, with the deepening of the hierarchy, each pixel can be taken into account in the calculation of different windows, so that the second encoding module can simultaneously realize the combination of local information and global information of the first feature encoding during the calculation process of the attention mechanism, thereby retaining detailed local information and global context in the process of generating the second training image.

[0118] In some embodiments, Figure 6 As shown, the source image, i.e., the audio training image, can also be directly input into the second encoding unit to output the second feature code. To this end, after the training audio feature and the first training image are spliced ​​to obtain the audio training image, the audio training image can be input into the second encoding module to divide the audio training image into a preset number of first image segments. The first image segment is then input into a certain encoding unit in the second encoding module. The encoding unit performs attention calculation on the first image segment obtained by dividing the audio training image to obtain a first attention encoding result.

[0119] After completing one attention calculation, the second encoding module re-divides the audio training image to obtain a preset number of second image segments, and then inputs the second image segments into the encoding unit of the next level, so that the encoding unit performs attention calculation on the second image segments to obtain a second attention encoding result. The above operation is iterated, and the second feature code is output according to the second attention encoding result.

[0120] In some embodiments, a hierarchical design is adopted between the encoding units, wherein each level gradually reduces the feature channels of the first feature encoding or the audio training image while increasing the dimension of the feature.

[0121] In some embodiments, the decoder may include a preset number of upsampling layers, and the upsampling layers gradually increase the picture feature dimensions corresponding to the arrangement order, while the number of feature channels gradually decreases. In order to reduce the amount of calculation, the embodiment of the present application does not adopt the traditional encoding and decoding connection structure, but adopts a non-corresponding connection relationship.

[0122] Among them, the feature dimensions of the upsampling layer and the coding unit increase with the increase of the level, and the number of feature channels of the upsampling layer and the coding unit decreases with the increase of the level. The feature dimensions of the convolution layer decrease with the increase of the level, and the number of feature channels increases with the increase of the level. In this embodiment, Figure 7 : is a structural diagram of the connection between the convolution layer and the upsampling layer in this embodiment. Figure 7 As shown, the corresponding relationship of the same level between the first encoding module and the decoder is different, so the first encoding module and the decoder adopt the following Figure 7 The connection relationship shown, for example, the convolution layer at the first level is connected to the upsampling layer at the first level, the convolution layer at the second level is connected to the upsampling layer at the second level, and so on. This can avoid the loss of relevant details as the network level deepens, and the training problem that the gradient may gradually disappear or explode as the number of layers increases.

[0123] Taking the convolution layer and upsampling layer at the first level as an example, the convolution layer at the first level has the largest feature dimension and the least number of feature channels, and can extract high-resolution features of the audio training image. The upsampling layer at the first level has the smallest feature dimension and the most feature channels, and can extract low-resolution features in the first encoding feature, so that the high-resolution features output by the convolution layer are directly connected to the low-resolution features in the upsampling layer, avoiding the loss of relevant details as the network level deepens, and the training problem that the gradient may gradually disappear or explode as the number of layers increases. At the same time, more spatial information can be included in the second training image during the decoding process, improving the network training effect and stability, and significantly reducing the amount of calculation.

[0124] like Figure 8 As shown in the figure, the feature dimension of the coding unit corresponds to the feature channel and the upsampling layer. Therefore, the coding unit and the upsampling layer are connected at the corresponding levels, and the attention encoding result output by the coding unit can be directly input into the upsampling layer to perform decoding operations.

[0125] It can be seen from the above technical solutions that the present application provides a method and system for processing character videos. The method obtains a video to be converted, performs language conversion on the audio to be converted in the video to be converted, and obtains a target audio. The first character image extracted from the target audio and the video to be converted is input into a lip conversion model to convert the first lip movement of the first character image according to the target audio to obtain a second character image; the target audio and the second character image are used to generate a target video. The present application adjusts the lip movement in the first character image through the translated audio data, so that the lip movement of the target character in the target video corresponds to the pronunciation of the translated target audio, thereby improving the user experience.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0127] For the convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are intended to better explain the present disclosure, so that those skilled in the art can better use the embodiments.

Claims

1. A method for processing a character video, characterized in that: include: Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image at least includes a first lip movement of a target person speaking in the first language; Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language; Inputting the target audio and the first person image into a lip-sync conversion model to obtain a second person image output by the lip-sync conversion model, wherein the second person image includes a second lip movement of the target person speaking in the second language; the lip-sync conversion model is trained by training data including the target person; A target video is generated according to the target audio and the second person image.

2. The character video processing method according to claim 1, characterized in that: Before the step of inputting the target audio and the first character image into the lip-sync conversion model, the method further includes: Acquire training data, wherein the training data is video data including the target person; Inputting the training data into a model to be trained, so as to extract training audio features of the training data and a first training image of a target person through the model to be trained, wherein the model to be trained is a lip-sync conversion model that has not been trained to convergence; Obtain a second training image generated by the to-be-trained model according to the training audio features and the first training image; Calculating the training loss of the second training image and the lip image label according to the loss function, wherein the lip image label is used to represent the pronunciation lip movement of the training audio feature; If the training loss is greater than the loss threshold, iteratively training the model to be trained; If the training loss is less than or equal to the loss threshold, a lip-sync model is obtained according to the model parameters of the model to be trained.

3. The character video processing method according to claim 2, characterized in that: The model to be trained includes an encoder and a decoder, wherein the encoder includes a first encoding module and a second encoding module, the first encoding module includes a preset number of convolutional layers, and the step of obtaining a second training image generated by the model to be trained according to the training audio features and the first training image, the method includes: splicing the training audio features and the first training image to obtain an audio training image; Input the audio training image into the first encoding module to traverse the pixel points in the audio training image through a convolution layer, wherein the pixel points include a first pixel point and a second pixel point, wherein the first pixel point is a valid pixel point in the target person's lip shape area, and the second pixel point is an invalid pixel point in the target person's lip shape area, and the invalid pixel point is used to indicate a pixel point that is covered in the target person's lip shape area; Performing a convolution operation on a valid area in the audio training image through the convolution layer to obtain a lip shape image feature, wherein the valid area is an image area formed by the first pixel points; Performing an iterative convolution operation on the lip-shaped image feature through the convolution layer to obtain a first feature code; Performing attention encoding on the first feature code by the second encoding module, and outputting a second feature code; The second feature code is decoded by the decoder to obtain a second training image.

4. The character video processing method according to claim 3, characterized in that: After the step of performing an iterative convolution operation on the lip-shaped image features through the convolution layer, the method further comprises: During the iterative convolution operation, a newly added valid area is obtained according to a preset convolution window, wherein the preset convolution window is a convolution window that moves periodically in the audio training image during the convolution iteration process; Updating the valid area according to the newly added valid area; A convolution operation is performed based on the updated valid area to update the lip shape image features.

5. The character video processing method according to claim 4, characterized in that: The steps of obtaining a newly added valid area according to a preset convolution window include: Traversing the first pixel point in the preset convolution window; If the preset convolution window includes at least one first pixel point, all the pixels in the preset convolution window are marked as first pixels to obtain a newly added effective area.

6. The character video processing method according to claim 4, characterized in that: The second encoding unit includes a hierarchical structure formed by a preset number of encoding units, the input of the encoding unit is the output of the previous encoding unit, and the output of the encoding unit is the input of the next encoding unit; The step of performing attention encoding on the first feature encoding by the second encoding module comprises: Dividing the first feature code into a preset number of first image segments; Inputting the first image segment into a coding unit in the second coding module, so that the coding unit independently performs attention calculation on each of the first image segments to obtain a first attention coding result; Adjusting the first image segment in a preset manner to divide the first feature code into a preset number of second image segments; wherein there is a partial overlap between the first image segment and the corresponding second image segment; Inputting the first attention encoding result into an encoding unit at a next level, so that the encoding unit performs attention calculation on the second image segment to obtain a second attention encoding result; Iterate the above operation to obtain the second feature encoding according to the second attention encoding result.

7. The character video processing method according to claim 4, characterized in that: The second encoding unit includes a hierarchical structure formed by a preset number of encoding units, the input of the encoding unit is the output of the previous encoding unit, and the output of the encoding unit is the input of the next encoding unit; After the step of obtaining the audio training image, the method further comprises: Inputting the audio training image into the second encoding module to divide the audio training image into a preset number of first image segments; Inputting the first image segment into a coding unit in the second coding module, so that the coding unit independently performs attention calculation on each of the first image segments to obtain a first attention coding result; Adjusting the first image segment in a preset manner to divide the first feature code into a preset number of second image segments; wherein there is a partial overlap between the first image segment and the corresponding second image segment; Inputting the first attention encoding result into an encoding unit at a next level, so that the encoding unit performs attention calculation on the second image segment to obtain a second attention encoding result; Iterate the above operation to obtain the second feature encoding according to the second attention encoding result.

8. The character video processing method according to claim 6 or 7, characterized in that: The decoder includes a preset number of upsampling layers; and the step of decoding the second feature code by the decoder to obtain a second training graph includes: Inputting the second feature code into the decoder, and performing a decoding operation on the second feature code according to the upsampling layer to obtain the second training image; Wherein, a preset number of convolutional layers in the first encoding module and / or a preset number of encoding units in the second encoding module and the upsampling layer are distributed according to a corresponding structure; In the decoder, the second feature code of the decoding operation performed by the upsampling layer is output by the non-corresponding convolutional layer and the encoding unit, or the second feature code of the decoding operation performed by the upsampling layer is directly output by the encoding unit.

9. The method for processing a person video according to claim 8, characterized in that: The feature dimensions of the upsampling layer and the coding unit increase with the increase of the level, and the number of feature channels of the upsampling layer and the coding unit decreases with the increase of the level; The feature dimension of the convolution layer decreases as the level increases, and the number of feature channels of the convolution layer increases as the level increases; The convolution layer and the upsampling layer are connected at the corresponding level, and the encoding unit and the upsampling layer are connected at the corresponding level.

10. A character video processing system, characterized in that: include: A controller and a lip-sync model, wherein the lip-sync model is trained using training data containing a target person; The controller is configured to: Acquire a video to be converted, the video to be converted including audio to be converted and a first person image, the language information of the audio to be converted is a first language, and the first person image at least includes a first lip movement of a target person speaking in the first language; Performing language conversion on the audio to be converted according to a second language to obtain a target audio, wherein the second language is different from the first language; Inputting the target audio and the first character image into a lip-sync model to obtain a second character image output by the lip-sync model, wherein the second character image includes a second lip movement of the target character speaking in the second language; A target video is generated according to the target audio and the second person image.

Citation Information

Patent Citations

  • Audio-driven figure mouth shape method, model and training method thereof

    CN115035604A

  • Digital human synthesis method based on multi-task learning

    CN115601230A

  • Remote sensing image semantic segmentation method based on double-branch feature fusion

    CN115797931A

  • Three-dimensional intelligent reconstruction method of liver tumor medical image

    CN115937423A

  • General instant 3D mouth shape animation generation method and device and storage medium

    CN115984430A