A video conversion method and apparatus

CN119967226BActive Publication Date: 2026-08-07NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING SILICON INTELLIGENCE TECH CO LTD
Filing Date
2024-11-07
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]为了解决用户在观看视频时,无法实现语种转换后视频中的人物口型不对应的问题,本申请部分实施例提供一种视频转换方法及装置

Benefits of technology

[0057]由以上技术方案可知,本申请提供一种视频转换方法及装置,所述方法通过获取用户输入的待转换视频,提取待转换音频的音频特征,以及,提取待转换音频中的第一文本。根据用户输入的第二语种信息对第一文本执行语种转换,得到第二文本。根据音频特征和第二文本对人物视频执行口型转换生成目标视频。本申请通过对待转换视频的音频执行特征提取和文本提取,并根据用户指定的语种对第一文本执行语种转换,再通过音频特征和语种转换后的第二文本对待转换视频中的人物口型进行转换,使人物口型与翻译后的文本内容对应,提高用户的观影体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967226B_ABST
    Figure CN119967226B_ABST
Patent Text Reader

Abstract

The application provides a video conversion method and device, the method obtains a to-be-converted video input by a user, extracts an audio feature of to-be-converted audio and the first text according to the to-be-converted video. The second text is obtained by performing language conversion on the first text according to the second language information input by the user. The target video is generated by performing conversion processing on the character video according to the audio feature and the second text. The application extracts the feature and the text of the audio of the to-be-converted video, performs language conversion on the first text according to the language specified by the user, and then converts the character mouth shape in the to-be-converted video according to the audio feature and the second text after language conversion, so that the character mouth shape corresponds to the translated text content, thereby improving the viewing experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese patent application No. 202311479579.X, filed on November 8, 2023, with the State Intellectual Property Office of China. The entire contents of the latter are incorporated herein by reference. Technical Field

[0002] This application relates to the field of speech conversion technology, and in particular to a video conversion method and apparatus. Background Technology

[0003] When watching films from different countries, users cannot directly understand the meaning of the dialogue from the audio alone because the film's language differs from their own. Therefore, users need to continuously read subtitles while watching the film to comprehend the meaning of the characters' lines.

[0004] To address this, speech-to-text technology can be used to convert the film's language to the user's desired language while preserving the characters' vocal characteristics, thus achieving real-time translation. However, during this process, because the language of the dialogue changes, the original lip movements of the characters in the film cannot correspond to the converted dialogue, reducing the user's viewing experience. Summary of the Invention

[0005] To address the issue of mismatched lip movements in videos after language conversion, some embodiments of this application provide a video conversion method and apparatus.

[0006] In a first aspect, some embodiments of this application provide a video conversion method, including:

[0007] The system acquires a video to be converted from user input, which includes a video of the target person speaking the first text and an audio recording of the target person speaking the first text.

[0008] The audio features of the audio to be converted and the first text are extracted from the video to be converted, wherein the first text is information about the first language.

[0009] The first text is converted to a second text based on the second language information input by the user, and the second language information is different from the first language information.

[0010] The video of the person is converted based on the audio features and the second text to generate a target video, which is a video of the target person speaking the second text.

[0011] In some embodiments, after obtaining the user-inputted video to be converted, the method further includes:

[0012] Detect the number of people in the video to be converted;

[0013] When there are multiple people in the video to be converted, obtain the user's input command to select the people;

[0014] In response to the character selection instruction, the target character is determined from among multiple characters in the video to be converted;

[0015] Obtain the video and audio to be converted of the target person.

[0016] In some embodiments, when no people are detected in the video to be converted, the method further includes:

[0017] Extract the narration audio from the video to be converted, and extract the narration text from the narration audio; wherein the narration audio is the audio data of the video to be converted when no characters are speaking;

[0018] The narration text is converted to a third text based on the second language information.

[0019] Generate a target narration voice based on the narration voice and the third text;

[0020] The target video is obtained by replacing the narration with the target narration.

[0021] In some embodiments, the step of extracting audio features from the audio to be converted includes:

[0022] Perform person recognition on the target person in the video to be converted to obtain the recognition result;

[0023] Based on the recognition results, the speech habit features and language style features of the target person are extracted, including speech tone features and speech rate features;

[0024] Audio features are generated based on the language style features, the speaking tone features, and the speaking speed features.

[0025] In some embodiments, the step of performing language conversion on the first text based on second language information input by the user includes:

[0026] Detect the first text segment in the first text, where the first text segment is a text segment whose translation differs from the homophonic interpretation.

[0027] The other text fragments in the first text, excluding the first text fragment, are converted to a different language to obtain the second text fragment.

[0028] The second text is synthesized from the second text fragment and the first text fragment.

[0029] In some embodiments, the step of detecting a first text fragment in the first text includes:

[0030] Detect the feature elements of the audio to be converted;

[0031] Obtain the progress position of the feature element in the audio to be converted;

[0032] In the first text, semantic detection is performed on the text segment at the progress position to obtain the semantic detection result;

[0033] If the semantic difference of the semantic detection result is greater than or equal to the determination threshold, the text segment at the progress position is marked as the first text segment.

[0034] In some embodiments, after performing language conversion on other text segments besides the first text segment in the first text, the method further includes:

[0035] A prompt interface is generated based on the first text fragment, and the prompt interface includes an explanation of the first text fragment;

[0036] The method further includes the step of performing lip-syncing on the video of the person based on the audio features and the second text.

[0037] The first video frame corresponding to the first text segment is obtained based on the progress position;

[0038] During the process of performing lip-syncing on the first video frame based on the audio features and the second text, a prompt interface is added to the first video frame.

[0039] In some embodiments, when the video to be converted is a live stream video, the method further includes:

[0040] Extract the audio from the live stream video;

[0041] Extracting live stream features from the live stream audio, and extracting live stream text from the live stream audio;

[0042] Based on the live stream features and live stream text, lip-syncing is performed on the person's video to obtain the target video.

[0043] In some embodiments, when the video to be converted includes multiple people, the method further includes:

[0044] The video to be converted is subjected to frame-by-frame processing to obtain the video frames to be converted;

[0045] Set all the characters in the video frame to be converted as the target characters.

[0046] In some embodiments, after the step of extracting the narration text of the narration speech, the method further includes:

[0047] Detect whether the narration contains a preset feedback voice, wherein the feedback voice is used to represent the feedback of a third-party viewer to the video content in the video to be converted;

[0048] In the case where the narration includes the feedback voice, the first narration text and the second narration text in the narration text are determined based on the feedback voice;

[0049] The first narration text is converted to a different language based on the second language information to obtain the third text; an auxiliary text is obtained based on the second narration text, wherein the auxiliary text is used to represent the meaning of the second narration text;

[0050] Generate the target narration voice based on the narration voice and the third text;

[0051] The narration voice is replaced according to the target narration voice, and the target video is obtained according to the target narration voice, the second narration text, and the auxiliary text.

[0052] Secondly, some embodiments of this application provide a video conversion device, including: a processor, an audio processing module, a translation module, and a lip-syncing module, wherein the processor is communicatively connected to the audio processing module, the translation module, and the lip-syncing module, respectively; the processor is configured to:

[0053] The system acquires a video to be converted from user input, which includes a video of the target person speaking the first text and an audio recording of the target person speaking the first text.

[0054] The audio processing module extracts the audio to be converted and the first text from the video to be converted, wherein the first text is information about the first language.

[0055] The translation module performs language conversion on the first text based on the second language information input by the user to obtain a second text, wherein the second language information is different from the first language information;

[0056] The lip-syncing module performs lip-syncing on the person video based on the audio features and the second text to generate a target video, which is a video of the target person speaking the second text.

[0057] As can be seen from the above technical solutions, this application provides a video conversion method and apparatus. The method acquires a video to be converted input by a user, extracts audio features of the audio to be converted, and extracts first text from the audio to be converted. It then performs language conversion on the first text based on second language information input by the user to obtain second text. Finally, it performs lip-syncing conversion on the video to generate a target video based on the audio features and the second text. This application extracts features and text from the audio of the video to be converted, performs language conversion on the first text according to the language specified by the user, and then uses the audio features and the language-converted second text to convert the lip-syncing of the characters in the video to be converted, making the lip-syncing correspond to the translated text content, thereby improving the user's viewing experience. Attached Figure Description

[0058] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 A flowchart illustrating a video conversion method provided in this application embodiment;

[0060] Figure 2 This is a video frame image used to identify the target person in an embodiment of this application;

[0061] Figure 3 This is a flowchart illustrating the process of extracting video of all people in a video frame and the audio to be converted, as described in this application embodiment.

[0062] Figure 4 This is a flowchart illustrating the process of updating the number of characters based on the video frames to be converted, as described in this application embodiment.

[0063] Figure 5 This is a flowchart illustrating the process of identifying audio features for a target person in an embodiment of this application;

[0064] Figure 6 This is a flowchart illustrating the conversion of narration audio into target narration speech in an embodiment of this application;

[0065] Figure 7 This is a flowchart illustrating the process of determining the first text fragment based on feature elements in an embodiment of this application;

[0066] Figure 8 This is a flowchart illustrating the training of an audio encoder in some embodiments of this application;

[0067] Figure 9 This document provides a flowchart illustrating the introduction of category coding during the training of the audio encoder in some embodiments of this application.

[0068] Figure 10 This is a flowchart of the training clustering module in some embodiments of this application;

[0069] Figure 11 This is a flowchart illustrating the audio feature replacement performed by the replacement unit in this embodiment of the application.

[0070] Figure 12 This is a flowchart illustrating the training of the audio encoder in an embodiment of this application;

[0071] Figure 13 This is a flowchart illustrating the mask prediction process performed on the first training data in an embodiment of this application.

[0072] Figure 14 This is a flowchart illustrating the mask prediction process performed by the second encoding unit in an embodiment of this application.

[0073] Figure 15 This is a flowchart illustrating the training of the lip-sync model in an embodiment of this application. Detailed Implementation

[0074] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0075] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0076] The terms "first," "second," "third," etc., used in the specification and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms can be used interchangeably where appropriate.

[0077] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0078] Users can play videos through smart devices, such as movies, TV series, operas, or variety shows. When watching films from different countries, because the film's language is different from the user's, the user cannot directly understand the meaning of the dialogue based on the characters' words. For example, when a Chinese user watches an English film, if the user cannot understand English, they cannot directly know the meaning of the dialogue based on the audio of the characters' words.

[0079] Therefore, some video resources can include subtitle resources, that is, displaying subtitles in the corresponding language on the screen while the video is playing. For example, when playing a video in English, Chinese subtitles can be added on the screen so that Chinese users can understand the meaning of the dialogue. This requires users to watch the video and subtitles simultaneously to understand the meaning of the characters' dialogue. Although the subtitles are on the screen, repeatedly watching content in different positions affects the user's viewing experience.

[0080] To improve the viewing experience, the dialogue in a film can be translated into a different language before being dubbed in post-production, allowing users to intuitively understand the meaning of the lines based on the dubbing. However, this post-production dubbing method is time-consuming and needs to be completed before the film is played, making it impossible to translate and play the dialogue in real time.

[0081] To address this, speech-to-text technology can be used to convert the film's language to the user's desired language while preserving the characters' vocal characteristics, thus achieving real-time translation. However, during this process, because the language of the dialogue changes, the original lip movements of the characters in the film cannot correspond to the converted dialogue, reducing the user's viewing experience.

[0082] To address the issue of mismatched lip movements between characters and translated text in videos, some embodiments of this application provide a video conversion device configured to execute a video conversion method to resolve the mismatch between the lip movements of characters speaking in the video and the translated text. Figure 1 A flowchart illustrating a video conversion method provided in an embodiment of this application. See also... Figure 1 The method includes:

[0083] S100: Obtain the video to be converted by the user input.

[0084] The video to be converted can include films, short videos, live streams, or recorded videos. It includes both a video of the target person speaking the initial text and the audio of the target person speaking the initial text. The target person is the individual appearing in the video. During lip-syncing, the lip movements of the target person need to be translated based on the translated text. Therefore, a video of the target person speaking the initial text is required. This video should include lip movements not yet translated, the target person's facial expressions, and body language. During subsequent video conversion, the target person's facial expressions and body language (excluding lip movements) can be kept consistent with those in the video. Alternatively, the video can be converted based on specific facial expressions or body language; for example, changing the target person's facial expressions or instructing them on specific body language movements within the video.

[0085] In some embodiments, the person video may be a video including only the head or face of a person, or it may be a video including other personal characteristics, such as a video of the target person speaking with specific facial expressions, body movements, clothing, appearance, age, etc. It should be understood that during the video conversion process, this application requires corresponding lip-syncing based on the converted text spoken by the target person. Therefore, the person video should at least include lip-syncing actions that are not subject to lip-syncing.

[0086] The video to be converted may include one or more people. When the video contains only one person, that person is the target person. When the video contains multiple people, the user can input a person selection command into the video conversion device to identify the target person from among the multiple people. Therefore, after acquiring the user-inputted video, the number of people in the video can be detected. When there are multiple people in the video, the user-input person selection command can be acquired; this command is used to select one person from among the multiple people to identify the target person. For example, as... Figure 2 As shown, the video conversion device can respond to a person selection command, and between person A and person B, select person A as the target person, that is... Figure 2 The dotted line figure shown represents the target person A. After identifying the target person A, the video and audio to be converted can be obtained based on the target person A.

[0087] It should be noted that the dotted and solid lines are only used in this embodiment to distinguish between target people and non-target people. In the actual video display, there is no difference in the display between target people and non-target people, thereby improving the effect of people in the video conversion process.

[0088] The above embodiment describes a lip-syncing process for one of multiple individuals. In some embodiments, lip-syncing can be performed on all individuals in the video to be converted. To do this, corresponding video clips and audio clips to be converted can be obtained based on the specific target individual. For example, such as... Figure 3 As shown, when the video to be converted includes both person A and person B, and both person A and person B are the target individuals, their voices can be extracted separately. In this example, the voices are the audio to be converted. After extracting the voices, the video clips of person A and person B are then extracted separately.

[0089] In some embodiments, the number of people in the video to be converted may change during playback. For example, the video to be converted may include two people in the current frame, but only one person 10 frames later. To determine the number of people in the video to be converted, such as... Figure 4 As shown, the video conversion device can perform frame-by-frame processing on the video to be converted, obtaining video frames to be converted. Each video frame to be converted contains a specific number of people. For example, if the current video frame to be converted contains one person, that person is marked as the target person. When the next video frame to be converted contains multiple people, the number of people for the multiple people is updated, and then, based on the number of people, all people are set as the target people, so that lip-syncing can be performed on the target people subsequently.

[0090] S200: Extract the audio features of the audio to be converted and the first text based on the video to be converted.

[0091] Video conversion devices can extract audio features from the audio to be converted using speech recognition technology. These audio features can include the target person's speaking speed, timbre, or emotions, such as anger, happiness, sadness, or fear, so that the subsequent audio is more in line with the target person's speaking context.

[0092] In some embodiments, due to the target person's personal style, the target person will speak with specific speaking habits, such as accent, tone of voice, or specific filler words. To better match the target person's speaking habits, the target person's speaking habit characteristics can be obtained.

[0093] Different target individuals have different speaking habits, therefore, such as Figure 5As shown, the video conversion device can perform person recognition on the target person in the video to be converted, and obtain the recognition result. Based on the recognition result, the video conversion device can extract the target person's speaking habits and language style features. The language style features are used to ensure the target person's speaking style, such as frequently used words or colloquialisms. Speaking habits features can include intonation and speaking speed features, so that when performing lip-syncing on the target person's video, the features are more closely matched to the target person's intonation and speaking speed. The video conversion device can generate audio features based on the language style features, intonation features, and speaking speed features, so as to improve the accuracy of lip-syncing by matching the lip movements in the video with the target person's speaking habits.

[0094] The video conversion device can extract the first text from the audio to be converted through speech recognition. In this embodiment, speech-to-text processing can be performed on the audio to be converted to extract the first text, which is information in the first language.

[0095] In some embodiments, besides speech recognition, the first text can also be obtained through other means. For example, when playing a video to be converted, the subtitle data of the video to be converted can be obtained, and text recognition can be performed on the subtitles to obtain the first text corresponding to the audio to be converted. Alternatively, the text resource track of the video to be converted can be obtained, and the first text can be obtained directly through the text resource track. The video to be converted may include multiple tracks, such as audio resource tracks and video resource tracks. The text resource track is the text content output by each character or narrator in the video to be converted. Therefore, the video conversion device can obtain the first text based on the text resource track.

[0096] After extracting the first text, the number of speakers in the audio to be converted can be identified based on audio features. If there are multiple speakers in the audio, the text segment of the speaker to be converted can be selected as the first text for subsequent translation. It should be noted that when there are multiple speech segments in the audio to be converted, the user selects the speech segments of one or more speakers for subsequent translation based on audio features.

[0097] In some embodiments, in addition to speech recognition, the video conversion device can also obtain the first text in other ways. For example, when playing the video to be converted, the video conversion device can obtain the subtitle data of the video to be converted and perform text recognition on the subtitles to obtain the first text. The video conversion device can also obtain the text resource track of the video to be converted and directly obtain the first text through the text resource track. Among them, the video to be converted can include multiple tracks, such as audio resource tracks, video resource tracks, etc. The text resource track is the text content output by each person or the narrator in the video to be converted. Therefore, the video conversion device can obtain the first text according to the text resource track.

[0098] S300: Perform language conversion on the first text according to the second language information input by the user to obtain the second text.

[0099] The video conversion device can perform language conversion on the first text according to the second language information input by the user. The second language information is different from the first language information. For example, if the first language information is Chinese, the first text can be "Hello" in Chinese. When the second language information is English, the second text can be "hello".

[0100] In some embodiments, in a scenario where the video to be converted includes multiple target persons, the video conversion device can translate the first text of each target person into the second text according to the second language information. However, when in the video to be converted, the languages spoken by different target persons are different, that is, the first language information of different first texts is different.

[0101] At this time, different conversion modes can be performed on different first texts according to the target language. For example, the language information of the first text A corresponding to the target person A is Chinese, and the language information of the first text B corresponding to the target person B is Spanish. When the second language information is English, the video conversion device needs to perform a Chinese-to-English translation operation on the first text A according to the second language information; the video conversion device needs to perform a Spanish-to-English translation operation on the first text B according to the second language information.

[0102] S400: Perform conversion processing on the person video according to the audio feature and the second text to generate a target video, where the target video is the person video in which the target person speaks the second text.

[0103] After obtaining the second text, the video conversion device can perform lip-syncing on the target person's video based on audio features and the second text to generate the target video. The target video is a video of the target person speaking the second text. The video conversion device can preserve the target person's voice characteristics after language conversion based on audio features, so that the audio features in the target video are the same as the target person's audio features. In the target video, the target person's lip movements correspond to the second text, improving the user's viewing experience.

[0104] In some embodiments, the video conversion device can also store the detected target person. In the video to be converted, if the target person is not located within a frame of the video to be converted, audio translation (i.e., lip-syncing) can be performed on other target persons within the frame of the video to be converted. When the target person reappears in a subsequent frame of the video to be converted, the video conversion device can continue to acquire the target person's video to be converted based on the target person's recognition result, and perform speech recognition on the video to be converted to obtain the first text.

[0105] In some embodiments, the video conversion device can also mark the target person and perform real-time detection based on the marked target person. When the marked target person appears in the video to be converted, the audio spoken by the target person can be acquired in real time, and the audio features of the audio and the video of the target person can be extracted, thereby performing the translation process described above, and finally displaying the video of the target person speaking the second text on the screen, thus achieving the purpose of real-time translation.

[0106] In some embodiments, when the target person does not appear in the video to be converted, or when the target person does not speak, audio data will also be played to introduce some conditions, people, scenes and other information in the video to be converted. This audio data can be called narration audio.

[0107] like Figure 6 As shown, when there is no target person in the video to be converted, the video conversion device extracts the narration from the video. The narration is the audio data of the video when no one is speaking. After obtaining the narration, the narration features and narration text can be extracted, and language conversion is performed on the narration text according to the second language information to obtain the third text. It should be understood that the conversion process of the third text is the same as that of the second text obtained by performing language conversion on the first text.

[0108] After obtaining the third text, the video conversion device can directly synthesize the target narration speech based on the narration speech features and the third text. Since there is no target person in the video to be converted, or the target person does not speak, there is no need to convert the target person's lip movements. Therefore, in the process of synthesizing the target narration speech, there is no need to obtain and convert the target person's video; only the target narration speech needs to be synthesized. After obtaining the target narration speech, the narration speech can be replaced to obtain the target video.

[0109] In some embodiments, videos with strong regional characteristics, such as entertainment variety shows, talk shows, and crosstalk, may be included. Therefore, for such videos, directly translating the original text may result in the translation of the relevant content failing to reflect its original meaning well due to puns, jokes, and other elements, leading to abrupt content and affecting the user experience.

[0110] In this embodiment, text fragments such as puns and jokes that are unsuitable for language conversion are defined as first text fragments. A first text fragment is a text fragment whose meaning differs from the original meaning after translation. During the translation of the first text based on the second language information, the video conversion device detects the first text fragments within the first text. After obtaining the first text fragment, since it is unsuitable for language conversion, it can be played directly using the original language information. Therefore, the video conversion device can perform language conversion on other text fragments besides the first text fragment to obtain second text fragments, and then combine the second text fragment and the first text fragment to form the second text.

[0111] In some embodiments, taking a comedy video as an example, when a first text fragment such as a pun appears in a comedy video, specific sound effects or expressive comments such as "hahaha" will appear. For example, the appearance of laughter or applause sound effects, or the display of comments like "hahaha," indicates the presence of a first text fragment in a comedy video. Therefore, in order to accurately obtain the first text fragment, such as... Figure 7 As shown, the video conversion device can detect feature elements of the audio to be converted and obtain the progress position of the feature elements in the audio. The first text is extracted from the audio to be converted; therefore, the corresponding positions of the first text and the audio to be converted are the same. To determine that the text segment at the progress position is the first text segment, the video conversion device can perform semantic detection on the text segment corresponding to the progress position in the first text, obtain the semantic detection result, and calculate the semantic difference between the semantic detection result and the preset semantic label. When the semantic difference is greater than or equal to a judgment threshold, the text segment corresponding to the progress position is marked as the first text segment. When the semantic difference is less than the judgment threshold, no marking is performed on the text segment.

[0112] It should be noted that the above example uses comedy videos as the video to be converted. In practical applications, the video to be converted can be other types of videos, such as speech videos, instructional videos, or press conference videos. This embodiment does not specifically limit the type of video to be converted.

[0113] In some embodiments, since the video conversion device does not perform language conversion on the first text segment, the user may not be able to directly understand the text content corresponding to the first text segment. Therefore, the video conversion device can generate a prompt interface based on the first text segment, the prompt interface including an explanation of the first text segment. When performing lip-syncing on a person's video based on audio features and the second text, the first video frame corresponding to the first text segment can be obtained based on the progress position of the first text segment, and a prompt interface can be added to the first video frame during the lip-syncing process based on the audio features and the second text. The prompt interface can be added in various forms, such as generating a prompt interface based on the first text segment using background text; or generating a prompt interface using a bullet screen (or similar format).

[0114] It should be understood that, in order to facilitate users' understanding of the interpretation content in a second language, the video conversion device can convert the interpretation content into a second language based on the second language information after it is generated, and generate a prompt pop-up based on the translated interpretation content.

[0115] The following example illustrates why text fragments unsuitable for language conversion. Taking a narration video featuring a third-party audience as an example, the third-party audience can be the video content within the narration video. For instance, the narration video might show a third-party audience listening to a talk show played through an audio system. In this case, the video conversion device can extract the narration audio from the video, i.e., the media asset audio of the talk show, as well as the feedback audio from the third-party audience. The feedback audio from the third-party audience can be used as feature elements to identify content in the narration audio that is unsuitable for language conversion. Therefore, the video conversion device can detect whether the narration audio contains preset feedback audio, such as laughter, hissing, or cheers from the third-party audience, which represent their feedback to the video content.

[0116] When the narration includes feedback audio, the video conversion device can divide the narration text into a first narration text and a second narration text based on the feedback audio. The second narration text is the narration text unsuitable for translation, i.e., the first text segment. Therefore, the video conversion device can perform language conversion on the first narration text (excluding the second narration text) based on the second language information to obtain a third text. After generating the third text, a target narration audio can be generated based on the narration audio and the third text, and the original narration audio can be replaced with the target narration audio.

[0117] It should be noted that, since the second narration text is not suitable for translation, the video conversion device will retain the narration audio corresponding to the second narration text. When replacing the narration audio, the video conversion device only replaces the narration audio corresponding to the first narration text; that is, it replaces the narration audio corresponding to the first narration text with the target narration audio. The narration audio corresponding to the second narration text is retained in the original narration audio and forms the target audio with the replaced target narration audio.

[0118] For second narration texts unsuitable for translation, the video conversion device can display the second narration text as a prompt interface in the target video, without performing speech conversion processing on the second narration text. To this end, the video conversion device can search for corresponding annotation text for the second narration text via LLM or the internet. The annotation text serves as auxiliary text, representing the meaning of the second narration text, so that viewers can understand its meaning through the auxiliary text. In this way, the original language second narration text is retained in the subtitles of the generated target video, and the auxiliary text facilitates viewers' understanding of the meaning of the second narration text in the target video. Furthermore, the language information of the narration speech in the target video is the second language information, achieving real-time audio-video conversion.

[0119] It should be understood that this embodiment is only used as an example to illustrate the concept. In actual application, if the video to be converted is a video of a person, the conversion process of the above embodiment can also be used. The only difference is that the lip movements of the target person in the video of the person are converted.

[0120] In some embodiments, since the video conversion device does not perform language conversion on the first text segment, during the video conversion of a person's video, when the target person speaks the first text segment, video conversion may not be performed on the target person, that is, the original lip movements, facial expressions, body movements, and other features are preserved. For the second text segment, video conversion can be performed according to the language-converted second text.

[0121] In some embodiments, the video to be converted can be a live stream video. Since live streams have high real-time requirements, the video conversion device can extract live stream audio from the live stream video. The live stream can be video data obtained by the video conversion device according to a server request. To improve real-time performance, the video conversion device can extract live stream audio from the live stream video while loading the video stream data sent by the server, so as to facilitate subsequent audio processing.

[0122] After obtaining the live stream audio, live stream features and live stream text can be extracted from the live stream audio. The process of extracting the live stream text can be referred to the extraction process of the first text in the previous embodiment, and will not be repeated in this embodiment. The video conversion device can translate the live stream text according to the second language information, and perform lip-syncing on the target person's video in the live stream video according to the live stream features and live stream text to obtain the target video.

[0123] This application also provides a video conversion apparatus in some embodiments, comprising a processor, an audio processing module, a translation module, and a lip-syncing module. The processor is communicatively connected to the audio processing module, the translation module, and the lip-syncing module, respectively. The processor is configured to execute the video conversion method described in the above embodiments, including:

[0124] S100: Obtain the video to be converted by the user input.

[0125] The video to be converted includes a video of the target person speaking the first text and an audio recording of the target person speaking the first text.

[0126] S200: Extract audio features from the audio to be converted using the audio processing module, and extract the first text from the audio to be converted using the audio processing module.

[0127] The first text is information in the first language.

[0128] S300: The translation module performs language conversion on the first text based on the second language information input by the user to obtain the second text.

[0129] The second language information is different from the first language information.

[0130] S400: The lip-syncing module performs lip-syncing on the person video based on the audio features and the second text to generate the target video.

[0131] The target video is a video of the target person speaking the second text.

[0132] In some embodiments, the lip-sync module may further include a speech conversion model, which generates target audio based on the translated second text and the audio to be converted, and extracts audio features from the target audio. To improve the accuracy of speech conversion, reference audio from the same sound source as the audio to be converted can also be obtained as audio data input to the speech conversion model.

[0133] The speech conversion model may also include an audio encoder and an audio decoder module. The audio encoder is used to extract reference audio features from the reference audio.

[0134] To this end, the category mapping unit can perform category mapping on other features of the training data according to a preset category code, thereby classifying other features according to different speakers and reducing the influence of other speakers' audio on the extraction of reference audio features. The category mapping unit consists of a mapping layer used to perform category mapping according to the category code, thereby extracting reference audio features from other features and improving the extraction accuracy of reference audio features.

[0135] In some embodiments, to achieve the aforementioned speech clustering extraction effect, the audio encoder needs to undergo a specific training process. During training, the feature encoding unit and the category mapping unit are first initialized, meaning some parameters of the feature encoding unit and the category mapping unit are randomly initialized. After initialization, a first training dataset is obtained, including a preset number of first training data points, used to train the audio encoder to be trained. After inputting the first training data into the audio encoder to be trained, the feature encoding unit can extract the first training features from the first training data. The first training features correspond to the clustering features of the speech conversion model in the application stage. Therefore, the speech conversion model performs category mapping on the first training features through the category mapping unit to obtain second training features, which correspond to the reference audio features of the speech conversion model in the application stage.

[0136] After obtaining the second training feature, the feature loss between the second training feature and the audio feature label can be calculated using a loss function. The audio feature label is used to represent the standard audio feature, and the standard audio feature is used as a reference to determine the output of the second training feature. The loss function can be expressed as follows:

[0137]

[0138] Where i is the id of the cluster category, e i Let be the embedding vector of the trainable audio encoder, sim be the cosine similarity, and τ be the hyperparameter.

[0139] like Figure 8As shown, when the feature loss is less than or equal to the first loss threshold, it indicates that the second training features extracted by the audio encoder meet the output standard, and the audio encoder completes the training process. At this point, the audio encoder can be output based on the current parameters of the feature encoding unit and the current parameters of the category mapping unit to complete the training. If the feature loss is greater than the first loss threshold, it indicates that the accuracy of the second training features extracted by the audio encoder has not reached the accuracy of the audio feature labels. Therefore, iterative training of the audio encoder is required, that is, iterative training of the feature encoding unit and the category mapping unit using the first training data.

[0140] In some embodiments, such as Figure 9 As shown, the audio encoder also includes a clustering module and a category encoding unit. The clustering module provides category codes to the audio encoder during training. During training, the category encoding unit obtains the category codes obtained by the clustering module after performing clustering on the speech sample data, and then provides these category codes to the category mapping unit. After obtaining the category codes, the category mapping unit performs category mapping on the first training features according to the category codes, clustering the first training features according to the category codes to obtain the second training features.

[0141] It should be noted that the category coding unit does not participate in the feature extraction process of the audio encoder during application. It is only introduced into the audio encoder during the training process to train the clustering extraction function of the audio encoder.

[0142] In the actual training process, in addition to training and updating the parameters of the feature encoding unit, this embodiment also trains the audio encoder based on the category encoding corresponding to the training data and the real category encoding ID obtained by the clustering module to improve the accuracy of audio feature extraction. Therefore, the speech conversion model also needs to train the clustering module. The clustering module can include a feature extractor configured to extract audio features from the training data. During the training of the clustering module, a second training dataset needs to be obtained. The second training dataset includes a preset number of second training data points, which can be audio data that is the same as or different from the first training data. The second training data can be based on LibriSpeech-960 and AISHELL-3 data; specifically, it obtains speech sample data from 200 speakers, with 200 clusters.

[0143] The second training data is input into the clustering module to be trained. A feature extractor extracts the third training features from the second training data and performs clustering classification on these third training features to generate category codes. During the clustering classification process, the feature extractor can calculate feature similarity on the third training features and cluster the third training data whose feature similarity is less than or equal to a similarity threshold, i.e., clustering similar audio features. After clustering, audio features of the same category can represent the audio features of an independent person, while audio features of different categories differ significantly. This facilitates the extraction of audio features for each person based on category features, thereby improving the feature extraction effect.

[0144] To improve the clustering training effect, the clustering module may further include a first clustering module and a second clustering module. The first and second clustering modules can respectively perform clustering training on the audio data of different individuals in the second training data. The feature extractors of the first and second clustering modules can be any two of three types of feature extractors: SoftHubert, Hubert, and WAV2VEC2.0. For example, the first clustering model can use Hubert as its feature extractor, and the second clustering model can use WAV2VEC2.0 as its feature extractor.

[0145] Therefore, the first and second clustering modules can be trained simultaneously using the second training data, enabling them to classify audio data from different speakers within the second training data. Since the first and second clustering modules employ different feature extractors, they can cluster from different dimensions, improving the extraction of various audio features, such as timbre and prosody. Thus, the combination of different feature extraction methods from the first and second clustering modules can enrich the training results when performing language transformations subsequently.

[0146] In some embodiments, such as Figure 10 As shown, during the training process of the first and second clustering modules, it is necessary to encode the categories output by the first and second clustering modules. For example, after the first clustering module clusters the second training data, it obtains different feature categories, which can be assigned ID1.1, ID1.2, ..., ID1.9, etc. Similarly, after the second clustering model clusters the second training data, it obtains different feature categories, which can be assigned ID2.1, ID2.2, ..., ID2.9, etc. The purpose of category encoding is to ensure that each category after clustering by the clustering module has a unique identifier for differentiation, so as to facilitate category mapping and encoding during subsequent language transformation training.

[0147] In some embodiments, the speech conversion model can also update the loss function of the audio encoder based on the first clustering module and the second clustering module. For example, the average cross-entropy between the true class code of the first clustering model and the predicted class code of the audio encoder is minimized, while the average cross-entropy between the true class code of the second clustering model and the predicted class code of the audio encoder is also minimized. Based on this, the loss function is updated, and the parameters of the audio encoder are updated. Through the above training method, the audio encoder's ability to classify timbre categories is further enhanced.

[0148] In some embodiments, the speech conversion model can also be trained based on feature substitution. For this purpose, the audio encoder further includes a permutation unit and a feature encoding unit. In this case, the audio encoder is configured to extract first audio features from the audio to be converted, and to extract second audio features from a reference audio source with the same sound source as the audio to be converted. The permutation unit is configured to perform speech feature substitution on the first and second audio features according to a permutation network to obtain target audio features. The audio decoding module is configured to decode the target audio features to obtain the target audio.

[0149] During the process of speech feature replacement performed by the permutation unit, such as Figure 11 As shown, the speech conversion model can calculate feature similarity values ​​through a permutation unit. These similarity values ​​are the similarity between the second audio feature and the first audio feature. Furthermore, the permutation unit, based on the KNN algorithm, queries a third audio feature within the second audio feature based on the feature similarity value. This third audio feature is then used to replace the first audio feature according to the permutation network, resulting in the target audio feature. The third audio feature is the second audio feature whose feature similarity value is greater than or equal to a similarity threshold. By replacing the original first audio feature with a third audio feature that is similar to it, the acoustic features of the obtained target audio feature better match the acoustic features of the target person, thereby improving the acoustic similarity of the target audio.

[0150] In some embodiments, the feature encoding unit extracts audio features from the audio through feature encoding. For example, the feature encoding unit performs feature encoding on the first audio to obtain the first audio features as follows: Then, feature encoding is performed on the second audio to obtain the second audio features. Where, x∈R d For the feature representation of a frame, N s N is the length of the first audio signal. t The length of the second audio signal.

[0151] During the calculation of feature similarity values, the permutation unit can generate a fourth audio feature based on the first audio feature using a non-nearest neighbor substitution method. The fourth audio feature is... After generating the fourth audio feature, the permutation unit can use a proximity algorithm to calculate the proximity values ​​between the second and fourth audio features, and generate feature similarity values ​​based on these proximity values. Then, based on these feature similarity values, a third audio feature is determined from the second audio features. The third audio feature is X. t,e =X t ∪X e During the speech feature replacement process, the replacement unit can perform feature sequence X of the third audio feature. t,e Averaging is performed to obtain the speech feature sequence to be replaced, and the feature sequence X in the first audio feature is replaced by the speech sequence to be replaced. s This leads to the final target audio encoding, i.e.

[0152] In some embodiments, for the audio encoder to perform the audio synthesis effect of speech feature replacement described above, a specific training process needs to be performed on the speech conversion model. To this end, before inputting the first and second audio files into the speech conversion model, a first training dataset needs to be obtained. The first training dataset includes a predetermined amount of first training data, which is audio data used to train the audio encoder. The first training data includes at least one speaker's audio to facilitate training the audio encoder's ability to extract audio features.

[0153] Before training the audio encoder, its parameters need to be initialized, i.e., some parameters of the feature encoding unit and the permutation unit are randomly initialized. After initialization, the first training data can be input into the audio encoder to be trained, so that the first training features of the first training data can be extracted through the feature encoding unit. Since the audio encoder needs to perform speech feature replacement between different audio features, during the training process, reference training data different from the first training data can be input into the audio encoder as sample audio data for speech feature replacement. Thus, speech feature replacement is performed using the first training features and the reference training features extracted from the reference training data to complete the training process of the permutation unit.

[0154] In some embodiments, such as Figure 12As shown, two first training features from different first training data can also be input into the permutation unit to perform speech feature replacement on the two first training features, resulting in a second training feature. After obtaining the second training feature, the speech conversion model calculates the feature loss between the second training feature and the audio feature label using a loss function. The audio feature label represents the standard audio features after speech feature replacement and is used to determine the convergence degree of the audio encoder. If the feature loss is less than or equal to the first loss threshold, it indicates that the feature accuracy of the second training feature output by the audio encoder after speech feature replacement reaches the accuracy of the audio feature label. The speech conversion model can then output the audio encoder based on the current parameters of the feature encoding unit and the permutation unit. If the feature loss is greater than the first loss threshold, it indicates that the feature accuracy of the second training feature output by the audio encoder after speech feature replacement has not reached the accuracy of the audio feature label. In this case, it is necessary to iteratively train the encoding unit and the permutation unit using the first training data, continuously updating the parameters of the training encoding unit and the permutation unit during the training process until the calculated feature loss is less than or equal to the first loss threshold, thus completing the training process of the audio encoder.

[0155] In some embodiments, to improve the accuracy of audio feature extraction by the feature coding unit, the feature coding unit may also employ mask prediction to extract audio features. Therefore, as... Figure 13 As shown, the feature encoding unit includes a masking subunit, a first encoding unit, and a second encoding unit. The masking subunit is configured to mask a portion of the audio features of the first training data to facilitate subsequent feature prediction based on the masked audio features. To this end, after the first training data is input into the feature encoding unit, the masking subunit can perform masking processing on the first training data to obtain masked training data. The masking subunit can perform masking on the first training data according to a preset masking ratio, which represents the proportion of data to be masked. The masking ratio can be between 30% and 50%. For example, when the masking ratio is 50%, half of the first training data is masked. The masked first training data can be a single data portion; for example, the first 50% is masked, and the last 50% is not, so that prediction can be performed on the first 50% of the audio features based on the last 50% of the first training data. The first training data that is masked can also be divided into multiple data portions. For example, the first 25% of the first training data is masked, the second 25% is not masked, the third 25% is masked, and the fourth 25% is not masked. This is to facilitate performing predictions on the masked data portions based on the first training data that is not masked.

[0156] After completing the masking process, the masking sub-unit can output the masking training data and input it into the first encoding unit. The first encoding unit is a CNN encoding unit, which can perform feature extraction on the masking training data using a convolutional neural network with a pre-defined kernel size, and output the masked audio features.

[0157] Since some training data in the masked training data is masked, the first encoding unit cannot extract features from the masked data. Instead, the first encoding unit can extract audio features from the unmasked portion of the masked training data to obtain the masked audio features.

[0158] The first masking unit can input the masked audio features into the second coding unit, which can be a Transformer coding unit. The second coding unit can further extract deeper features from the masked audio features and perform feature prediction on the masked part of the masked audio features based on the extracted deeper features to obtain the first training features. Thus, by training with the mask, the accuracy of the feature coding unit in extracting audio features is improved.

[0159] During feature prediction by the second coding unit, prediction needs to be performed based on the attention weights of different feature sequences to improve the accuracy of feature prediction. Therefore, the second coding unit also needs to calculate the attention weights of the feature sequences of the masked audio features before performing feature prediction.

[0160] Figure 14 This is a flowchart illustrating the calculation of attention weights based on mask training features in an embodiment of this application. See also... Figure 14 The mask training features include multiple mask feature sequences. The second encoding unit can obtain the sequence length of the mask audio features and generate a relative position index based on the sequence length. The relative position index is the distance index between any two sequences in the mask audio features. For example, for mask feature sequence i, the second encoding unit can use the relative distance index r(i, j) between mask feature sequence i and mask feature sequence j. After generating the relative distance index, the second encoding unit can generate a relative position vector between mask feature sequence i and mask feature sequence j based on the distance index, and calculate the attention weight of mask feature sequence i relative to mask feature sequence j based on the relative position vector. It should be understood that the above embodiment is the process of calculating the attention weight of mask feature sequence i relative to mask feature sequence j. In the actual training process, attention weights can also be calculated based on any different mask feature sequences. The second encoding unit can perform feature prediction on the masked part of the mask audio features based on the attention weights of multiple sets of mask feature sequences to obtain the first training features.

[0161] After obtaining the target audio, the target audio and a first image of the person in the video can be input into the lip-sync model of the lip-sync module to generate a second image of the person. The lip movements of the second image are the same as those in the target audio. The lip-sync module can be applied to display devices used for playing audio and video or to lip-sync systems based on language conversion. This embodiment uses a lip-sync system as an example.

[0162] In some embodiments, a pre-training process for the lip-sync model is also required. For this purpose, user-input character training data can be obtained, where the character training data consists of video data containing the target character. This video data can include lip movements, facial expressions, or body movements of the target character. When the target character is a person in a video, the lip-sync model can perform character recognition. If the target character is a public figure such as a film or television actor or singer, video data containing the target character can be searched as character training data based on the character recognition results. Examples of such video data include online course videos, performance videos, speech videos, and film / television videos.

[0163] In some embodiments, the video to be converted can also be a video recorded by the user that contains the target person. In this case, the user can use other video data of the person being filmed as training data to train the lip-sync model.

[0164] After obtaining the character training data, such as Figure 15 As shown, training data on a person can be input into the model to be trained. The model to be trained is an untrained lip-sync model, and it has the same model structure as the lip-sync model. After receiving the training data, the model to be trained can extract the training audio features from the training data to determine the pronunciation lip shape based on these features. The model to be trained can also extract the first training image from the training data. Since the training data is video data containing the target person, the first training image contains the target person's facial image.

[0165] The model to be trained can generate a second training image based on the training audio features and the first training image. In the second training image, the target person makes the same mouth movements as those in the training audio features. Since the parameters of the model to be trained have not yet converged, the second training image generated by the model should have a certain training loss. Therefore, the training loss between the second training image and the lip-shape image labels can be calculated using a loss function. The lip-shape image labels are used to represent the mouth movements of the training audio features, and the training loss is used to determine whether the second training image meets the output standard.

[0166] To determine whether the second training image meets the output standard, in some embodiments, the lip-sync model can set a loss threshold. If the training loss is greater than the loss threshold, it indicates that the training loss between the second training image and the lip-sync image label is too large, and the second training image cannot be output. Therefore, the lip-sync system can iteratively train the training model to update the model parameters of the model to be trained and regenerate the second training image. When the training loss is less than or equal to the loss threshold, it indicates that the second training image meets the output standard, and the lip-sync system can obtain the lip-sync model based on the model parameters of the model to be trained.

[0167] In the above embodiments, the training and application processes are only for the target person. If the video data of the target person cannot be obtained well, the lip-syncing process cannot be performed on the lip-syncing model. Therefore, the person in the training data can also be someone other than the target person. For example, the training data can be images of people including person A. After training the lip-syncing ability of the lip-syncing model using the training data, the lip-syncing system can perform lip-syncing on person B, who is different from person A, during the application process.

[0168] In some embodiments, the lip-sync model may include a transformation module, wherein the transformation module is configured to calculate, at least based on the target audio, a transformation coefficient corresponding to each feature channel in the image features of the first person image, and perform feature transformation on the image feature channels of the first person image based on the transformation coefficients to obtain the image features of the first person image.

[0169] Based on the transformation module, the model to be transformed can also include another training method:

[0170] The lip-syncing system can input training data into the model to be trained. After the training data is input into the model to be trained, the model to be trained extracts the training audio features and the first training image from the training data, respectively, and performs encoding on the training audio features through an encoder to obtain the training audio code. In addition, it performs encoding on the first training image through an encoder to obtain the first training encoded image.

[0171] After the encoder completes encoding, the training audio code and the first training image code can be input into the transformation module. The transformation module then performs lip-syncing processing on the first training image code based on the training audio code. The transformation module consists of fully connected layers. Through these fully connected layers, the transformation module learns the transformation relationships between different images, transforming the first training image code to the desired lip-syncing state based on the training audio code.

[0172] After inputting the training audio encoding and the first training image encoding into the transformation module, the first training image encoding becomes the image encoding to be transformed. The transformation module can determine the feature channel calculation relationship corresponding to the first training image encoding based on the training audio encoding, thereby realizing the transformation of the training character's lip movements in the first training image at the feature channel level. The feature channel calculation relationship corresponds to different transformation relationships of the feature channels. For example, the feature channel calculation relationship can include rotation, translation, and scaling relationships.

[0173] To determine the feature channel calculation relationships, the lip-syncing system can extract a reference image from the training data based on the encoding of the first training image. The reference image can be an image of the same person as the training subject in the first training image, used to represent the image of the training subject in the training data and serving as a reference image for obtaining the feature channel calculation relationships. For example, the reference image can be a frontal image of the training subject. By using the reference image as a reference, the transformation relationships of different feature channels, such as the position, facial angle, or facial size of the first training image relative to the reference image, can be determined. Based on these transformation relationships, the feature channel calculation relationships can be determined to facilitate subsequent lip-syncing of the encoding of the first training image.

[0174] In some embodiments, the reference image may also be multiple images of a person from different angles to determine the feature channel calculation relationship from different angular dimensions. For example, the reference image may include a first reference image, a second reference image, and a third reference image from different angles.

[0175] After obtaining the reference image, to improve the accuracy of feature channel calculation relationships, the reference image can be concatenated with the first training image to obtain a training concatenated image. Transform coefficients are then calculated based on the concatenated training image and the training audio encoding to determine the feature channel calculation relationships according to the training audio features. Specifically, when performing lip-syncing on the first training image encoding, the transformation module can perform corresponding transformation calculations on the feature channels of the first training image encoding based on the transform coefficients to obtain the second training image encoding. The transform coefficients indicate the transformation relationship of the first training image relative to the reference image. After obtaining the second training image encoding, a decoder decodes the second training image to obtain the second training image.

[0176] As can be seen from the above technical solutions, this application provides a video conversion method and apparatus. The method acquires a video to be converted input by a user, extracts audio features of the audio to be converted and the first text based on the video to be converted. It then performs language conversion on the first text based on the second language information input by the user to obtain the second text. Finally, it performs conversion on the video of a person based on the audio features and the second text to generate the target video. This application extracts features and text from the audio of the video to be converted, performs language conversion on the first text according to the language specified by the user, and then uses the audio features and the language-converted second text to convert the lip movements of the person in the video to be converted, so that the lip movements correspond to the translated text content, thereby improving the user's viewing experience.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0178] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the foregoing exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be made based on the foregoing teachings. The selection and description of the above embodiments are for the purpose of better explaining the contents of this disclosure, thereby enabling those skilled in the art to better utilize the described embodiments.

Claims

1. A video conversion method, characterized in that, include: The system acquires a video to be converted, which is input by the user and includes a video of the target person speaking the first text. The audio features of the audio to be converted and the first text are extracted from the video to be converted, wherein the first text is information about the first language. The first text is converted to a second text based on the second language information input by the user, and the second language information is different from the first language information. The video of the person is converted based on the audio features and the second text to generate a target video, wherein the target video is a video of the target person speaking the second text. The step of performing language conversion on the first text based on the second language information input by the user includes: Detect the first text segment in the first text, where the first text segment is a text segment whose translation differs from the homophonic interpretation. The other text fragments in the first text, excluding the first text fragment, are converted to a different language to obtain the second text fragment. The second text is synthesized from the second text fragment and the first text fragment.

2. The video conversion method according to claim 1, characterized in that, After obtaining the user-inputted video to be converted, the method further includes: Detect the number of people in the video to be converted; When there are multiple people in the video to be converted, obtain the user's input command to select the people; In response to the character selection instruction, the target character is determined from among multiple characters in the video to be converted; Obtain the video and audio to be converted of the target person.

3. The video conversion method according to claim 1, characterized in that, When no person is detected in the video to be converted, the method further includes: Extract the narration audio from the video to be converted, and extract the narration text from the narration audio; wherein the narration audio is the audio data of the video to be converted when no characters are speaking; The narration text is converted to a third text based on the second language information. Generate a target narration voice based on the narration voice and the third text; The target video is obtained by replacing the narration with the target narration.

4. The video conversion method according to claim 1, characterized in that, The step of extracting audio features from the audio to be converted includes: Perform person recognition on the target person in the video to be converted to obtain the recognition result; Based on the recognition results, the speech habit features and language style features of the target person are extracted, including speech tone features and speech rate features; Audio features are generated based on the language style features, the speaking tone features, and the speaking speed features.

5. The video conversion method according to claim 1, characterized in that, The step of detecting a first text segment in the first text includes: Detect the feature elements of the audio to be converted; Obtain the progress position of the feature element in the audio to be converted; In the first text, semantic detection is performed on the text segment at the progress position to obtain the semantic detection result; If the semantic difference of the semantic detection result is greater than or equal to the judgment threshold, the text segment at the progress position is marked as the first text segment.

6. The video conversion method according to claim 5, characterized in that, After performing language conversion on other text segments in the first text besides the first text segment, the method further includes: A prompt interface is generated based on the first text fragment, and the prompt interface includes an explanation of the first text fragment; The method further includes the step of performing lip-syncing on the video of the person based on the audio features and the second text. Obtain the first video frame corresponding to the first text segment based on the progress position; During the process of performing lip-syncing on the first video frame based on the audio features and the second text, a prompt interface is added to the first video frame.

7. The video conversion method according to claim 1, characterized in that, When the video to be converted is a live stream video, the method further includes: Extract the audio from the live stream video; Extracting live stream features from the live stream audio, and extracting live stream text from the live stream audio; Based on the live stream features and live stream text, lip-syncing is performed on the person's video to obtain the target video.

8. The video conversion method according to claim 2, characterized in that, When the video to be converted includes multiple people, the method further includes: The video to be converted is subjected to frame-by-frame processing to obtain the video frames to be converted; Set all the characters in the video frame to be converted as the target characters.

9. The video conversion method according to claim 3, characterized in that, After the step of extracting the narration text of the narration speech, the method further includes: Detect whether the narration contains a preset feedback voice, wherein the feedback voice is used to represent the feedback of a third-party viewer to the video content in the video to be converted; In the case where the narration includes the feedback voice, the first narration text and the second narration text in the narration text are determined based on the feedback voice; The first narration text is converted to a different language based on the second language information to obtain the third text; an auxiliary text is obtained based on the second narration text, wherein the auxiliary text is used to represent the meaning of the second narration text; Generate a target narration voice based on the narration voice and the third text; The narration voice is replaced according to the target narration voice, and the target video is obtained according to the target narration voice, the second narration text, and the auxiliary text.

10. A video conversion device, characterized in that, include: The processor comprises an audio processing module, a translation module, and a lip-syncing module, wherein the processor is communicatively connected to the audio processing module, the translation module, and the lip-syncing module, respectively; the processor is configured to: The system acquires a video to be converted from user input, which includes a video of the target person speaking the first text and an audio recording of the target person speaking the first text. The audio processing module extracts the audio to be converted and the first text from the video to be converted, wherein the first text is information about the first language. The translation module performs language conversion on the first text based on the second language information input by the user to obtain a second text, wherein the second language information is different from the first language information; The lip-syncing module performs conversion processing on the person video based on the audio features of the audio to be converted and the second text to generate a target video, which is a video of the target person speaking the second text. The translation module performs language conversion on the first text based on the second language information input by the user, including: Detect the first text segment in the first text, where the first text segment is a text segment whose translation differs from the homophonic interpretation. The other text fragments in the first text, excluding the first text fragment, are converted to a different language to obtain the second text fragment. The second text is synthesized from the second text fragment and the first text fragment.

Citation Information

Patent Citations

  • Video translation method, system and device and storage medium

    CN112562721A

  • Model-based dubbing to translate spoken audio in a video

    US11514948B1