Video processing method and apparatus, and electronic device, storage medium, program and product
By automatically extracting and translating text and speech from videos using machine learning models to generate target videos, this technology solves the problem of low efficiency in existing video translation and dubbing techniques, achieving efficient and character-appropriate dubbing effects and improving the user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
In existing technologies, video translation and dubbing processes rely on human translation and dubbing, which is inefficient and costly. It is impossible to translate and dub a large number of videos in a short period of time, and the dubbing effect does not closely match the character's image.
The original speech and text in the video are automatically extracted using a machine learning model. The text is translated using a first machine learning model, and the target speech is generated using a second machine learning model. The image and the target speech are then fused to generate the target video, which imitates the voice characteristics of the character to improve the dubbing effect.
It enables automated translation and dubbing of videos, improving translation and dubbing efficiency, and making the dubbing effects closer to the characters, thus enhancing the viewing experience.
Smart Images

Figure CN2024129039_07052026_PF_FP_ABST
Abstract
Description
Video processing methods, apparatuses, electronic devices, storage media, programs and products Technical Field
[0001] This disclosure relates to the fields of artificial intelligence and computer technology, and in particular to a video processing method, apparatus, electronic device, storage medium, program and product. Background Technology
[0002] With the development of internet technology, video and other content can be rapidly disseminated worldwide. Since users in different countries or regions may use different languages, videos, especially films and television shows, are translated into different languages to provide a better viewing experience. Currently, translators and voice actors typically handle the translation and dubbing of videos.
[0003] Summary of the Invention
[0004] According to some embodiments of this disclosure, a video processing method is provided, comprising: acquiring a first text corresponding to the original speech of one or more characters in a video, wherein the original speech and the first text are expressed in a first language; translating the first text into a second text expressed in a second language using a first machine learning model; generating target speech for one or more characters using a second machine learning model based on the voice features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text; and fusing the images of the video with the target speech to generate a target video.
[0005] According to some other embodiments of this disclosure, a video processing apparatus is provided, comprising: an acquisition module configured to acquire first text corresponding to the original speech of one or more characters in a video, wherein the original speech and the first text are expressed in a first language; a translation module configured to translate the first text into second text expressed in a second language; a generation module configured to generate target speech for one or more characters based on the voice features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text; and a fusion module configured to fuse the images of the video with the target speech to generate a target video.
[0006] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to perform a video processing method of any embodiment of the present disclosure based on instructions stored in the memory.
[0007] According to some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the video processing method of any embodiment described in the present disclosure.
[0008] According to some other embodiments of the present disclosure, a computer program is provided, comprising: instructions that, when executed by a processor, cause the processor to perform a video processing method according to any embodiment of the present disclosure.
[0009] According to further embodiments of the present disclosure, a computer program product is provided, including instructions that, when executed by a processor, implement the video processing method of any embodiment of the present disclosure.
[0010] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0011] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:
[0012] Figure 1 shows a flowchart illustrating a video processing method according to some embodiments of the present disclosure;
[0013] Figure 2 shows a schematic flowchart of a video processing method according to some other embodiments of the present disclosure;
[0014] Figure 3 shows a schematic diagram of the structure of a video processing apparatus according to some embodiments of the present disclosure;
[0015] Figure 4 shows a schematic diagram of the structure of an electronic device according to some embodiments of the present disclosure;
[0016] Figure 5 shows a schematic diagram of the structure of an electronic device according to other embodiments of the present disclosure. Detailed Implementation
[0017] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0018] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of modules and steps set forth in these embodiments should be interpreted as merely exemplary and does not limit the scope of this disclosure.
[0019] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".
[0020] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0023] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0024] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Furthermore, based on one or more embodiments, those skilled in the art can clearly derive any suitable combination from this disclosure based on specific features, structures, or characteristics.
[0025] To provide a better viewing experience and make videos more easily understood and accepted by users in different countries and regions, some videos are currently translated and dubbed manually before being released to other countries or regions. In recent years, online dramas, short dramas, and various other types of videos have developed and grown rapidly. A large number of videos cannot be translated and dubbed manually in a short time. With the development of artificial intelligence (AI) and large-scale modeling technology, subtitles in videos can be automatically translated into other languages. However, dubbing videos according to the translated languages generally still requires manual work. Voice actors can imitate the original voice in the video, more closely resembling the character's image, thus bringing a better audiovisual effect. This method still cannot dub a large number of videos in a short time and is inefficient and costly in terms of labor.
[0026] This disclosure proposes a video processing method that can automatically translate and dub videos. It can generate dubbing by mimicking the voice characteristics of characters in the video, improving the efficiency of translation and dubbing while making the dubbing more closely resemble the characters, thus enhancing the audiovisual effect of the dubbed video and improving the viewing experience. Some embodiments of the video processing method of this disclosure are described below with reference to Figures 1 and 2.
[0027] The video processing method disclosed herein can be executed by a video processing device or an electronic device, and the video processing device can be implemented in software and / or hardware.
[0028] Figure 1 is a flowchart of some embodiments of the video processing method of this disclosure. As shown in Figure 1, the interaction method of this embodiment includes steps S102 to S108.
[0029] In step S102, the first text corresponding to the original voice of one or more characters in the video is obtained.
[0030] The video includes the original voice of one or more characters (or figures), multiple (or frames) of images, and may also include background audio, such as background music, narration, etc. Audio can be extracted from the video, and the original voice and background audio can be separated to obtain the voices of one or more characters (i.e., the original voice) and background audio.
[0031] The video can be a web series, short drama, TV series, movie, or other form of video, not limited to the examples given. The first text can be subtitles. If the video includes subtitles, the subtitles can be extracted from the video as the first text. If the video does not include subtitles, speech recognition technology can be used to convert the original speech into the first text. Both the original speech and the first text are expressed in the first language. For example, the first language is Chinese, not limited to the examples given.
[0032] In step S104, a first machine learning model is used to translate the first text into a second text expressed in a second language.
[0033] For example, the first machine learning model is an LLM (Large Language Model). When translating the first text, the first machine learning model needs to consider the context of each sentence to make the resulting second text more accurate. The second language is different from the first language; for example, the second language is English, but this is not limited to the examples given.
[0034] In step S106, a second machine learning model is used to generate the target speech of one or more characters based on the voice features of each character in one or more characters, the correspondence between one or more characters and multiple sentences in the second text, and the second text.
[0035] For example, the second machine learning model is a text-to-speech (TTS) model, but is not limited to the examples given. The target speech is expressed in a second language, and the speech corresponding to each character in the target speech is matched with the voice features of each character. The voice features of each character can be extracted from the original speech. It is necessary to determine the corresponding sentence in the second text for each character, and then convert the corresponding sentence of each character in the second text into speech that matches the voice features of each character, thus obtaining the target speech.
[0036] Since the generated target speech has the same or similar sound features as the original speech of the character in the video, the target speech sounds just like the voice spoken by the character in the video. Furthermore, since the target speech is expressed in a second language, it sounds as if the character in the video is actually speaking a second language, which can achieve a better dubbing effect.
[0037] In some embodiments, sound features include at least one of acoustic features and prosodic features. Acoustic features include at least one of the following: frequency, timbre, duration, and intensity. Prosodic features include at least one of the following: tone, stress, intonation, and rhythm.
[0038] In step S108, the video images are fused with the target speech to generate the target video.
[0039] The original audio in the video is removed. If the video also includes background noise, the background noise can be synthesized with the target audio to obtain the audio in the video. Then, the images in the video (or the video without audio) are edited together with the audio to generate the target video. Subtitles can also be configured for the target video; a second text can be added as a subtitle to the video's images.
[0040] The timing information of the original speech and the correspondence between the original speech and multiple images can be obtained through video. The correspondence between each sentence in the original speech and each sentence in the target speech is also determined. Therefore, the timing of the target speech and multiple images can be correlated, and then the target video can be obtained by editing and fusing according to the correspondence between the target speech and multiple images.
[0041] In the video processing method described above, firstly, the first text corresponding to the original audio in the video is obtained. Then, the first text is translated into a second text expressed in a different language. Target audio is generated based on the second text. Finally, the images and target audio in the video are fused to generate the target video. During the generation of the target audio from the second text, the voice characteristics of each character are considered, making the generated target audio closely resemble the timbre, pitch, and speech rate of the original audio. This results in the target audio sounding closer to the voice of the character themselves, making the dubbing more fitting to the character's image in the video, improving the dubbing effect, and providing users with a better audiovisual experience when watching the target video. The process of generating the target video is fully automated, requiring no translators or voice actors, thus improving the efficiency of target video generation.
[0042] Before generating the target video, the video can be preprocessed to obtain one or more characters, the correspondence between one or more characters and the original speech, and the speech features of each character. The following describes how to preprocess the video using some examples.
[0043] In some embodiments, multiple keyframes are extracted from the video; multiple facial images are identified in the multiple keyframes; duplicate facial images are removed from the multiple facial images; and one or more characters are determined based on the remaining one or more facial images.
[0044] Frame extraction can be performed on a video to obtain multiple keyframes (images). Face recognition can be performed on these keyframes to identify multiple facial images. By comparing these facial images, duplicate images are removed, and the remaining one or more images are used as the facial images for one or more characters. Each character can be assigned a unique identifier. The identifier for each character and its corresponding facial image are stored in a database.
[0045] Image recognition in a video can automatically identify all characters and their facial features, which can then be used to generate the target video. It can also recognize the original speech in the video to determine the speech features of each character. In some embodiments, based on multiple reference images (or video clips) showing one or more characters speaking in the video; and the correspondence between these reference images (or video clips) and multiple speech clips in the original speech, a reference speech clip for each character is determined; and based on the reference speech clip for each character, the voice features of each character are determined.
[0046] Based on the facial images of each character, images in the video can be identified to determine the video segment or multiple reference images in which each character is speaking. For example, lip-reading technology can be used to determine the video segment or multiple reference images in which each character is speaking. Since there is a correspondence between the images in the video and the original speech, the reference speech segment corresponding to each character can be determined based on the video segment or multiple reference images in which each character is speaking. The reference speech segment corresponding to each character is then parsed to extract the voice features of each character. The voice features of each character can be stored in a database along with the character's identifier.
[0047] For example, the reference speech segments for each character can include reference speech segments corresponding to different emotional information, which can determine the vocal characteristics of each character under different emotions. Emotion recognition can be performed on the original speech to determine the speech under different emotions.
[0048] By preprocessing the video, all characters in the video and the voice features of each character are identified. This information is then used to generate more accurate target speech and target video, improving the dubbing effect. Through video preprocessing, the database can store the correspondence between each character's identifier, facial image, and voice features.
[0049] The following describes how to obtain the first text, using some examples.
[0050] In some embodiments, in response to the video including subtitles, subtitles are obtained as first text; in response to the video not including subtitles, raw speech is obtained from the video, and the raw speech is converted into first text using a third machine learning model.
[0051] For example, OCR (Optical Character Recognition) methods can be used to recognize subtitles in each frame of a video, remove duplicate content, and obtain the first text, not limited to the examples given.
[0052] For example, the third machine learning model is the ASR (Automatic Speech Recognition) model, which can convert raw speech into first text, and is not limited to the examples given.
[0053] When the video includes subtitles, the first text can be determined separately from the subtitles. The original speech is then converted into the first text, and the two first texts are automatically proofread to obtain the proofread first text. For example, LLM (Local Level Management) can be used for automatic proofreading to determine a more accurate first text. By combining subtitle recognition methods with automatic speech recognition methods, the accuracy of the determined first text can be improved, thereby improving the accuracy of subsequent generation of target speech and target video.
[0054] After obtaining the first text, we can also obtain the correspondence between multiple sentences in the first text and multiple images in the video, as well as the time information of each sentence in the first text, which can be used to generate the target video later.
[0055] The following describes, with reference to some embodiments, how to translate a first text into a second text expressed in a second language.
[0056] In some embodiments, a first machine learning model is used to translate the sentences of the first text corresponding to each speech segment into sentences of the second text based on the duration of each speech segment in the original speech and the translation target, wherein the difference between the estimated duration of the translated sentences of the second text after conversion to speech and the duration of the corresponding speech in the original speech is within a preset range.
[0057] Translating a first text into a second text may involve multiple expressions, each using different vocabulary, grammar, and sentence length. During this process, the duration of the first text must be considered. The goal is to ensure that the difference between the duration of each sentence in the second text after conversion to speech and its corresponding duration in the original speech is not too large. This prevents overlap or excessive overlap between the playback times of different sentences after conversion.
[0058] For example, the original speech can be divided into multiple speech segments, each corresponding to one or more sentences in a first text. Based on the text segment (one or more sentences) in the first text corresponding to each speech segment, the duration of each speech segment, and the translation target, a prompt is generated. This prompt is then input into a first machine learning model to obtain the sentence in the second text. For instance, the text in the first text corresponding to each speech segment and the duration of each speech segment are used as input fields in the prompt, and the translation target can be used as the task description information in the prompt. Alternatively, the translation target can be set to ensure that the duration of the sentence in the second text after conversion to speech is the same as the duration of the corresponding speech in the original speech, or it can be set to minimize the difference between the duration of the sentence in the second text after conversion to speech and the duration of the corresponding speech in the original speech.
[0059] In the process of fusing video images with target speech, the duration of the target speech can be adjusted. However, excessive adjustment can lead to poor auditory quality of the target speech. Therefore, during the translation process, the length, words, and grammatical structure of the second text are controlled to ensure that the target speech generated from the second text matches and synchronizes with the timing of the images in the video, making the target speech playback smoother and more natural.
[0060] In some embodiments, a first machine learning model is used to translate the sentences of the first text corresponding to each speech segment into sentences of the second text based on the duration of each speech segment in the original speech, the number of syllables of words in the second language, and the translation target. The estimated duration of the sentences of the second text after being converted into speech is determined based on the number of words in the sentences of the second text and the number of syllables of each word.
[0061] The first machine learning model can be used to judge the number of syllables in words or sentences of different expressions obtained from translation, and then estimate the duration of words or sentences of different expressions. An expression method can be selected so that the difference between the duration of the second text after the sentence is converted into speech and the duration of the corresponding original speech is within a preset range or minimized.
[0062] The method described in the above embodiments sets a translation target for the translation process, enabling the first machine learning model to control the expression of the second text during the translation process. This ensures that the duration of the generated second text, after being converted into target speech, is basically consistent with that of the original speech, allowing the target speech and images in the video to match better during the fusion process, thereby improving the audiovisual effect of the generated target video.
[0063] The following describes how to generate target speech using some examples.
[0064] In some embodiments, the speech segment corresponding to each statement in the second text in the original speech is determined; the emotional information of the speech segment corresponding to each statement in the second text is identified as the emotional information of each statement in the second text; and a second machine learning model is used to generate the target speech of one or more characters based on the voice features of each character in one or more characters, the correspondence between one or more characters and multiple statements in the second text, the emotional information of each statement in the second text, and the second text.
[0065] The correspondence between each statement in the second text and each statement in the first text is known, as is the correspondence between each statement in the first text and each speech segment in the original speech. Therefore, the speech segment corresponding to each statement in the second text in the original speech can be determined. For example, each statement in the second text, its corresponding emotional information, and the voice features of the corresponding character can be input into a second machine learning model to obtain the speech corresponding to each statement in the output.
[0066] The same sentence can be expressed with different emotions, resulting in variations in speech rate, intonation, stress, and other vocal characteristics. By identifying the emotional information of each segment of the original speech and referencing this information during the target speech generation process, the generated target speech can more closely resemble the emotional tone and vocal characteristics of the original speech, thus improving the accuracy of the generated target speech.
[0067] In some embodiments, the second text is divided into texts corresponding to each character based on the correspondence between one or more characters and multiple statements in the second text; a second machine learning model is used to generate the speech of each character in the target speech based on the corresponding text and sound features of each character.
[0068] By referencing the vocal features of each character, the generated target speech can maintain consistency with the original speech in terms of timbre, pitch, and speech rate, effectively replicating (cloning) each character's voice. For example, the vocal features of each character can be retrieved from a database, and these features, along with the corresponding text, can be input into a second machine learning model to obtain the output target speech for each character. The generation of each character's speech can be performed in parallel, improving generation efficiency.
[0069] The following describes, with reference to some examples, how to determine the text or statement corresponding to each role in the second text.
[0070] In some embodiments, the correspondence between one or more characters and multiple statements in the first text is determined based on at least one of the correspondences between multiple images (or multiple video segments) in the video and multiple statements in the first text, and the correspondence between multiple statements in the first text and multiple speech segments in the original speech; the correspondence between one or more characters and multiple statements in the second text is determined based on the correspondences between multiple statements in the first text and multiple statements in the second text, and the correspondences between one or more characters and multiple statements in the first text.
[0071] It is possible to determine the correspondence between one or more characters and multiple statements in the first text, that is, the statement corresponding to each character in the first text. Given the statement corresponding to each character in the first text, and the correspondence between each statement in the first text and each statement in the second text, it is possible to determine the statement corresponding to each character in the second text.
[0072] The corresponding statement for each character in the first text can be determined based on the image (or video clip) corresponding to each statement in the first text and / or the audio clip in the original audio of each statement in the first text.
[0073] In some embodiments, the presence of a speaking face image is identified in each of a plurality of images (or multiple video segments in a video); the speaking character is determined based on the speaking face image and the face image of each character; and the correspondence between the speaking character and the multiple statements in the first text is determined based on the correspondence between the plurality of images (or multiple video segments) and multiple statements in the first text, and the image (or video segment) where the speaking character is located.
[0074] Image recognition can be performed on videos. For example, lip-reading technology can be used to identify the frame (or video segment) containing the image of a speaking face. By matching the image of the speaking face with the facial images of each character in a database, the speaking character can be identified. Based on the image (or video segment) containing the speaking character and the corresponding statement in the first text for each image (or video segment), the correspondence between the speaking character and the statement in the first text can be determined.
[0075] The method described in the above embodiments identifies the speaking character by recognizing images in a video, and then determines what the speaking character is saying. This method can be used to associate one or more characters with statements in a second text, thereby enabling dubbing of each statement using the character's voice characteristics, thus improving the accuracy and effect of dubbing.
[0076] In some embodiments, for each image (or video segment) of a face that does not have a speaking face, the corresponding speech segment in the original speech is determined; based on the sound features of the speech segments corresponding to each image (or video segment) of a face that does not have a speaking face and the sound features of each character, the character corresponding to each image (or video segment) of a face that does not have a speaking face is determined; based on the correspondence between each image (or video segment) of a face that does not have a speaking face and the statements in the first text, and the character corresponding to each image (or video segment) of a face that does not have a speaking face, the correspondence between the character corresponding to each image (or video segment) of a face that does not have a speaking face and one or more statements in the first text is determined.
[0077] In some cases, the speaker may not be in the frame, or the speech may be narrated. For each image (or video clip) where there is no image of a speaking face, the speaker can be identified by comparing the corresponding speech's sound features with the sound features of each character. Based on the image corresponding to the speaking character and the statement in the first text corresponding to each image, the correspondence between the speaking character and the statement in the first text can be determined.
[0078] The method described in the above embodiments identifies voice features and associates one or more characters with statements in a second text, thereby enabling dubbing of each statement using the character's voice features, thus improving the accuracy and effectiveness of dubbing.
[0079] If multiple images correspond to the same speech segment, or speech segments with the same sound features, then these multiple images (or video segments) can correspond to the same character. For multiple images (or video segments) containing images of a speaking face, while identifying the speaking character based on the facial images, the speaking character can also be identified through the corresponding speech sound features. This verifies the accuracy of the correspondence between images (or video segments) and characters, thereby more accurately determining the character corresponding to each image (or video segment) and the corresponding statement in the first text for each character.
[0080] In some embodiments, the speech segment in the original speech corresponding to each image (or video segment) in a plurality of images in a video is determined; the character corresponding to each image (or video segment) is determined based on the sound features of the speech segment in the original speech corresponding to each image and the sound features of each character; and the correspondence between one or more characters and the multiple statements in the first text is determined based on the correspondence between the plurality of images (or video segments) and the multiple statements in the first text and the character corresponding to each image (or video segment).
[0081] Since the database stores the voice features of each character, the voice features of each character can be compared with the voice features of each speech segment in the original language to determine the speech segment corresponding to each character, and then the statement in the first text corresponding to each character can be determined.
[0082] In the process of fully automated translation and dubbing of videos, enabling the video processing device to identify each character, the words spoken by each character, and the voice characteristics of each character is a key step in the subsequent generation of the target video. The method in the above embodiment can accurately determine the correspondence of each character, each sentence, and each audio segment through image recognition and sound feature recognition, thereby generating target audio and target video more accurately.
[0083] To further improve the audiovisual effect of the generated target video, lip movements of corresponding characters can be generated based on the target speech, replacing the lip movements of characters in the video to match the target speech. In some embodiments, the sound features of the target speech are determined; based on the sound features of the target speech and the mapping relationship between sound features and lip movement parameters, the lip movements in the video image (or video without audio) are adjusted; the adjusted image (or video without audio) is then fused with the target speech to generate the target video.
[0084] Different vocal characteristics, such as pitch and speech rate, affect the speed and amplitude of lip movements. For example, when the pitch of a speech increases, the mouth may open wider; when the speech rate is fast, the lip movements will also change more rapidly. The mapping relationship between vocal characteristics and lip movement parameters can be established in advance. After determining the vocal characteristics of the target speech, the lip movement parameters can be determined based on the mapping relationship. Then, the images in the video can be rendered based on the lip movement parameters to obtain the adjusted images.
[0085] In some embodiments, lip-shape parameters for each character are generated based on the sound features of each character in the target speech, the mapping relationship between sound features and lip-shape parameters, and the lip-shape parameters of each character in the target speech. Based on the correspondence between multiple images (or video segments) in the video and multiple speech segments in the target speech, the characters in the multiple images, and the lip-shape parameters of each character, the lip-shape in the images of the video (or video without audio) is adjusted.
[0086] Lip rendering and adjustments can be performed separately for each character. First, a facial model for each character can be created. This model includes the geometry and texture of oral cavity components such as lips, teeth, and tongue. 3D modeling technology allows for precise description of the shapes of these components and definition of their movement relationships. For example, the degree of lip opening and closing, and the upward or downward tilt of the corners of the mouth, can be modeled in detail. Furthermore, a parametric model of the lip shape can be created for each facial model. Key parameters are typically used to describe changes in lip shape, such as lip opening (measured by the distance between the lips), lip protrusion, and the position of the corners of the mouth. These parameters effectively control the shape of the lip and can be dynamically adjusted based on vocal characteristics.
[0087] Establishing a mapping relationship between voice features and lip-shape parameters can be achieved using a fourth machine learning model. This fourth machine learning model can be a neural network. For example, a Recurrent Neural Network (RNN) or a Long Short-Term Memory (LSTM) network can be used to process the time-series features of speech and map them to sequences of changes in lip-shape parameters. During training, a large amount of speech and lip-shape data pairs are needed to learn this mapping relationship. This data can be manually labeled or extracted from real videos. By continuously adjusting the weights of the neural network, it can accurately predict changes in lip-shape parameters based on voice features.
[0088] To make the lip-syncing more accurate for each character, a fourth machine learning model can be trained for each character. The speech of each character in the target audio is input into the corresponding fourth machine learning model to obtain the lip-sync parameters for each character. Then, based on the correspondence between the speech and the image for each character, the lip-sync parameters for each character are divided into lip-sync parameters for different images or video segments. Finally, the images or video segments are rendered based on the lip-sync parameters to obtain the lip-sync-adjusted image.
[0089] Then, based on the correspondence between the time of the images in the video and the time of the target audio, the images in the video (or the video without audio) and the target audio can be clipped and edited to obtain the completed target video.
[0090] The method described in the above embodiments can match the lip movements of characters in a video with the target audio, allowing users to perceive that the characters are speaking in a second language, thus creating a more immersive viewing experience and enhancing the audiovisual experience. Furthermore, the lip movements in the target video are generated entirely automatically, improving generation efficiency.
[0091] Some other embodiments of the video processing method disclosed herein are described below with reference to Figure 2.
[0092] Figure 2 is a flowchart of some other embodiments of the video processing method of this disclosure. As shown in Figure 1, the interaction method of this embodiment includes steps S201 to S214.
[0093] In step S201, keyframes are extracted from the video.
[0094] In step S202, based on keyframes, one or more characters in the video are identified, and the identifier of each character and the facial image of each character are stored in the database.
[0095] In step S203, the audio files in the video are extracted, and the original speech and background sound are separated.
[0096] In step S204, the original speech is split into speech segments based on the first text.
[0097] In step S205, a reference voice segment for each character is determined based on the facial image of each character and the video segment corresponding to each voice segment, and the voice characteristics of each character are determined based on the reference voice segment of each character.
[0098] In step S206, the first text corresponding to the video is extracted.
[0099] In step S207, the first text is translated into the second text.
[0100] In step S208, the role corresponding to each statement in the second text is identified.
[0101] In step S209, the target speech is generated based on the character corresponding to each statement in the second text and the voice characteristics of each character.
[0102] In step S210, video files that do not contain audio are extracted from the video.
[0103] In step S211, the video file that does not contain audio is split into multiple video segments according to the first text.
[0104] In step S212, lip-sync parameters are generated based on the target speech, and multiple new video clips are generated based on the lip-sync parameters.
[0105] In step S213, the background audio of the video is obtained.
[0106] In step S214, the target speech, background sound, and multiple new video clips are spliced together and their timelines are aligned to generate the target video.
[0107] The method described in the above embodiments can automatically translate, dub, and match lip movements to generate a target video, improving the efficiency of target video generation. Furthermore, the target video can provide users with a good audio-visual experience, thereby increasing the dissemination efficiency and scope of the target video.
[0108] The generated target video can be distributed to different pages. In some embodiments, the target video is displayed on the target page; the target video is played in response to a user's playback operation.
[0109] The target video can be used as a video in an immersive video stream, for example, displaying a target page in response to a user's swipe action and playing the target video on the target page.
[0110] This disclosure also provides a video processing apparatus, which will be described below with reference to FIG3.
[0111] Figure 3 is a structural diagram of a video processing apparatus according to some embodiments of the present disclosure. As shown in Figure 3, the video processing apparatus 30 of this embodiment includes: an acquisition module 310, a translation module 320, a generation module 330, and a fusion module 340.
[0112] The acquisition module 310 is configured to acquire a first text corresponding to the original speech of one or more characters in the video, wherein the original speech and the first text are expressed in a first language.
[0113] The translation module 320 is configured to translate the first text into a second text expressed in a second language.
[0114] The generation module 330 is configured to generate target speech for one or more characters based on the voice features of each character in one or more characters, the correspondence between one or more characters and multiple statements in the second text, and the second text.
[0115] The fusion module 340 is configured to fuse the video images with the target speech to generate the target video.
[0116] In some embodiments, the translation module 320 is configured to use a first machine learning model to translate the sentences of the first text corresponding to each speech segment into sentences of the second text based on the duration of each speech segment in the original speech and the translation target, wherein the estimated duration of the translated sentences of the second text after conversion to speech is within a preset range as the difference between the estimated duration of the speech in the original speech and the duration of the speech in the original speech is within a preset range.
[0117] In some embodiments, the translation module 320 is configured to employ a first machine learning model to translate the sentences of the first text corresponding to each speech segment into sentences of the second text based on the duration of each speech segment in the original speech, the number of syllables of words in the second language, and the translation target. The estimated duration of the sentences of the second text after being converted into speech is determined based on the number of words in the sentences of the second text and the number of syllables of each word.
[0118] In some embodiments, the generation module 330 is configured to divide the second text into texts corresponding to each character based on the correspondence between one or more characters and multiple statements in the second text; and to generate the speech of each character in the target speech using a second machine learning model based on the corresponding text and sound features of each character.
[0119] In some embodiments, the generation module 330 is configured to determine the speech segment corresponding to each statement in the second text in the original speech; identify the emotional information of the speech segment corresponding to each statement in the second text as the emotional information of each statement in the second text; and use a second machine learning model to generate target speech for one or more characters based on the voice features of each character in one or more characters, the correspondence between one or more characters and multiple statements in the second text, the emotional information of each statement in the second text, and the second text, wherein the voice features of each character include at least one of acoustic features and prosodic features.
[0120] In some embodiments, the generation module 330 is configured to determine the correspondence between one or more characters and multiple statements in the first text based on at least one of the correspondences between multiple images in the video and multiple statements in the first text, and the correspondences between multiple statements in the first text and multiple speech segments in the original speech; and to determine the correspondence between one or more characters and multiple statements in the second text based on the correspondences between multiple statements in the first text and multiple statements in the second text, and the correspondences between one or more characters and multiple statements in the first text.
[0121] In some embodiments, the generation module 330 is configured to identify whether a speaking face image exists in each of the plurality of images; determine the speaking character based on the speaking face image and the face image of each character; and determine the correspondence between the speaking character and the plurality of statements in the first text based on the correspondence between the plurality of images and the plurality of statements in the first text, and the image in which the speaking character is located.
[0122] In some embodiments, the generation module 330 is configured to determine the speech segment corresponding to each image of a non-speaking face in the original speech; determine the role corresponding to each image of a non-speaking face based on the sound features of the speech segments corresponding to each image of a non-speaking face and the sound features of each role; and determine the correspondence between the role corresponding to each image of a non-speaking face and one or more statements in the first text based on the correspondence between each image of a non-speaking face and statements in the first text and the role corresponding to each image of a non-speaking face.
[0123] In some embodiments, the generation module 330 is configured to determine the speech segment in the original speech corresponding to each of the multiple images in the video; determine the character corresponding to each image based on the sound features of the speech segment in the original speech corresponding to each image and the sound features of each character; and determine the correspondence between one or more characters and multiple statements in the first text based on the correspondence between the multiple images and multiple statements in the first text and the character corresponding to each image.
[0124] In some embodiments, the video processing apparatus 30 further includes: a preprocessing module 350 configured to extract multiple keyframes from the video; identify multiple facial images in the multiple keyframes; remove duplicate facial images from the multiple facial images; and determine one or more characters based on the remaining one or more facial images.
[0125] In some embodiments, the preprocessing module 350 is further configured to determine a reference speech segment corresponding to each character based on multiple reference images of one or more characters speaking in the video; the correspondence between the multiple reference images and multiple speech segments in the original speech; and to determine the vocal features of each character based on the reference speech segment corresponding to each character.
[0126] In some embodiments, the fusion module 340 is further configured to determine the sound features of the target speech; adjust the lip movements in the video image according to the sound features of the target speech and the mapping relationship between the sound features and lip movement parameters; and fuse the adjusted image with the target speech to generate the target video.
[0127] In some embodiments, the fusion module 340 is configured to generate a lip shape image for each character based on the sound features of each character in the target speech, the mapping relationship between the sound features and lip shape parameters; and to adjust the lip shapes in the images of the video based on the correspondence between multiple images in the video and multiple speech segments in the target speech, the characters in the multiple images, and the lip shape image of each character.
[0128] In some embodiments, the acquisition module 310 is configured to acquire subtitles as first text in response to the video including subtitles; and to acquire raw speech from the video in response to the video not including subtitles, and to convert the raw speech into first text using a third machine learning model.
[0129] In some embodiments, the video processing apparatus 30 further includes a display module 360, configured to display a target video on a target page; and to play the target video in response to a user's playback operation on the target video.
[0130] This disclosure also provides an electronic device. Figure 4 shows a block diagram of an electronic device according to some embodiments of this disclosure.
[0131] Memory 41 is used to store one or more computer-readable instructions. Memory 41 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 41 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.
[0132] The processor 42 is configured to execute computer-readable instructions to implement the video processing method or the method described in any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments; repeated details will not be elaborated upon here.
[0133] Processor 42 can be configured to execute the steps shown in Figures 1 and 2. Processor 42 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc.
[0134] The processor 42 and the memory 41 can communicate with each other directly or indirectly. For example, the processor 42 and the memory 41 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 42 and the memory 41 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0135] It should be noted that the components of the electronic device 4 shown in Figure 4 are exemplary and not limiting. The electronic device 4 may have other components depending on the actual application requirements. The processor 42 can control other components in the electronic device 4 to perform the desired functions.
[0136] Electronic device 4 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.
[0137] Figure 5 shows a block diagram of an electronic device according to other embodiments of the present disclosure.
[0138] The electronic device 5 shown in Figure 5 can be a computer system with a dedicated hardware structure, capable of performing corresponding functions when relevant applications are installed.
[0139] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0140] As shown in Figure 5, the Central Processing Unit (CPU) 51 performs various processes based on programs stored in the Read-Only Memory (ROM) 52 or programs loaded from the storage section 58 into the Random Access Memory (RAM) 53. The RAM 53 stores data required as needed when the CPU 51 performs various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 52, RAM 53, and storage section 58 can be various forms of computer-readable storage media. It should be noted that although the ROM 52, RAM 53, and storage section 58 are shown separately in Figure 5, one or more of them can be combined or located in the same or different memories or storage modules.
[0141] CPU 51, ROM 52 and RAM 53 are interconnected via bus 54. Input / output interface 55 is also connected to bus 54.
[0142] The following components are connected to the input / output interface 55: input section 56, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 57, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 58, including hard disks, magnetic tapes, etc.; and communication section 59, including network interface cards such as LAN cards, modems, etc. The communication section 59 allows communication processing to be performed via a network such as the Internet. It is readily understood that although parts of the electronic device 5 shown in Figure 5 communicate via bus 54, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0143] As needed, drive 510 is also connected to input / output interface 55. Removable media 511, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on drive 510 as needed, so that computer programs read from them can be installed into storage section 58 as needed.
[0144] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 511.
[0145] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to perform the methods described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 59, or installed from storage section 58, or installed from ROM 52. When the computer program is executed by CPU 51, the methods of embodiments of this disclosure are performed.
[0146] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0147] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0148] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium that, when executed by a processor, implement the methods described in any of the foregoing embodiments.
[0149] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0150] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0151] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the video processing method described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.
[0152] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0154] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0155] According to some embodiments of this disclosure, a video processing method is provided, comprising: acquiring a first text corresponding to the original speech of one or more characters in a video, wherein the original speech and the first text are expressed in a first language; translating the first text into a second text expressed in a second language using a first machine learning model; generating target speech for one or more characters using a second machine learning model based on the voice features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text; and fusing the video images with the target speech to generate a target video.
[0156] In some embodiments, translating a first text into a second text expressed in a second language using a first machine learning model includes: using the first machine learning model to translate the sentences of the first text corresponding to each speech segment into sentences of the second text based on the duration of each speech segment in the original speech and the translation target, wherein the estimated duration of the translated sentences after conversion to speech and the duration of the corresponding speech in the original speech are within a preset range.
[0157] In some embodiments, employing a first machine learning model to translate the statement of the first text corresponding to each speech segment into a statement of the second text based on the duration of each speech segment in the original speech and the translation target includes: employing a first machine learning model to translate the statement of the first text corresponding to each speech segment into a statement of the second text based on the duration of each speech segment in the original speech, the number of syllables of words in the second language, and the translation target, wherein the estimated duration of the statement of the second text after being converted into speech is determined based on the number of words in the statement of the second text and the number of syllables of each word.
[0158] In some embodiments, using a second machine learning model to generate target speech for one or more characters based on the voice features of each character in one or more characters, the correspondence between one or more characters and multiple statements in a second text, and the second text, includes: dividing the second text into texts corresponding to each character based on the correspondence between one or more characters and multiple statements in the second text; and using the second machine learning model to generate speech for each character in the target speech based on the corresponding text and voice features of each character.
[0159] In some embodiments, using a second machine learning model to generate target speech for one or more characters based on the voice features of each character in one or more characters, the correspondence between one or more characters and multiple statements in a second text, and the second text, includes: determining the speech segment corresponding to each statement in the second text in the original speech; identifying the emotional information of the speech segment corresponding to each statement in the second text as the emotional information of each statement in the second text; and using the second machine learning model to generate target speech for one or more characters based on the voice features of each character in one or more characters, the correspondence between one or more characters and multiple statements in the second text, the emotional information of each statement in the second text, and the second text, wherein the voice features of each character include at least one of acoustic features and prosodic features.
[0160] In some embodiments, the video processing method further includes: determining a correspondence between one or more characters and multiple statements in the first text based on at least one of the correspondences between multiple images in the video and multiple statements in the first text, and the correspondences between multiple statements in the first text and multiple speech segments in the original speech; and determining a correspondence between one or more characters and multiple statements in the second text based on the correspondences between multiple statements in the first text and multiple statements in the second text, and the correspondences between one or more characters and multiple statements in the first text.
[0161] In some embodiments, determining the correspondence between one or more characters and multiple sentences in the first text based on at least one of a first correspondence between multiple images in a video and multiple sentences in a first text, and a second correspondence between multiple sentences in the first text and multiple speech segments in the original speech includes: identifying whether a speaking face image exists in each of the multiple images; determining the speaking character based on the speaking face image and the face image of each character; and determining the correspondence between the speaking character and the multiple sentences in the first text based on the correspondence between the multiple images and multiple sentences in the first text, and the image where the speaking character is located.
[0162] In some embodiments, determining the correspondence between one or more characters and multiple statements in the first text based on at least one of a first correspondence between multiple images in the video and multiple statements in the first text, and a second correspondence between multiple statements in the first text and multiple speech segments in the original speech, further includes: determining the speech segment corresponding to each image of a non-speaking face in the original speech; determining the character corresponding to each image of a non-speaking face based on the sound features of the speech segments corresponding to each image of a non-speaking face and the sound features of each character; and determining the correspondence between the character corresponding to each image of a non-speaking face and one or more statements in the first text based on the correspondence between each image of a non-speaking face and statements in the first text, and the character corresponding to each image of a non-speaking face.
[0163] In some embodiments, determining the correspondence between one or more characters and multiple statements in the first text based on at least one of a first correspondence between multiple images in the video and multiple statements in the first text, and a second correspondence between multiple statements in the first text and multiple speech segments in the original speech, includes: determining a speech segment in the original speech corresponding to each image in the video; determining the character corresponding to each image based on the sound features of the speech segment in the original speech corresponding to each image and the sound features of each character; and determining the correspondence between one or more characters and multiple statements in the first text based on the correspondence between multiple images and multiple statements in the first text and the character corresponding to each image.
[0164] In some embodiments, the video processing method further includes: extracting multiple keyframes from the video; identifying multiple facial images in the multiple keyframes; removing duplicate facial images from the multiple facial images; and determining one or more characters based on the remaining one or more facial images.
[0165] In some embodiments, the video processing method further includes: determining a reference speech segment corresponding to each character based on multiple reference images of one or more characters speaking in the video; the correspondence between the multiple reference images and multiple speech segments in the original speech; and determining the voice features of each character based on the reference speech segment corresponding to each character.
[0166] In some embodiments, fusing video images with target speech to generate a target video includes: determining the sound features of the target speech; adjusting the lip movements in the video images according to the sound features of the target speech and the mapping relationship between the sound features and lip movement parameters; and fusing the adjusted images with the target speech to generate the target video.
[0167] In some embodiments, adjusting the lip movements in the video images based on the mapping relationship between the sound features of the target speech and the sound features and lip movement parameters includes: generating a lip movement image for each character based on the sound features of the speech of each character in the target speech and the mapping relationship between the sound features and the lip movement parameters; and adjusting the lip movements in the video images based on the correspondence between multiple images in the video and multiple speech segments in the target speech, the characters in the multiple images, and the lip movement image of each character.
[0168] In some embodiments, obtaining the first text corresponding to the original speech of one or more characters in a video includes: obtaining the subtitles as the first text in response to the video including subtitles; and obtaining the original speech from the video in response to the video not including subtitles, and converting the original speech into the first text using a third machine learning model.
[0169] In some embodiments, the video processing method further includes: displaying a target video on a target page; and playing the target video in response to a user's playback operation on the target video.
[0170] According to some other embodiments of this disclosure, a video processing apparatus is provided, comprising: an acquisition module configured to acquire first text corresponding to the original speech of one or more characters in a video, wherein the original speech and the first text are expressed in a first language; a translation module configured to translate the first text into second text expressed in a second language; a generation module configured to generate target speech for one or more characters based on the voice features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text; and a fusion module configured to fuse the images of the video with the target speech to generate a target video.
[0171] According to further embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform a video processing method as described in any embodiment of the present disclosure based on instructions stored in the memory.
[0172] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the video processing method of any embodiment of the present disclosure.
[0173] According to some other embodiments of the present disclosure, a computer program is provided, comprising: instructions that, when executed by a processor, cause the processor to perform a video processing method according to any embodiment of the present disclosure.
[0174] According to further embodiments of the present disclosure, a computer program product is provided, including instructions that, when executed by a processor, implement the video processing method of any embodiment of the present disclosure.
[0175] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A video processing method, comprising: Obtain a first text corresponding to the original speech of one or more characters in a video, wherein the original speech and the first text are expressed in a first language; The first machine learning model is used to translate the first text into a second text expressed in the second language. A second machine learning model is used to generate the target speech of the one or more characters based on the voice features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text. The images from the video are fused with the target audio to generate the target video.
2. The video processing method according to claim 1, wherein, The step of translating the first text into a second text expressed in a second language using a first machine learning model includes: Using the first machine learning model, based on the duration of each speech segment in the original speech and the translation target, the sentence of the first text corresponding to each speech segment is translated into a sentence of the second text, wherein the translation target is that the difference between the estimated duration of the sentence of the second text after conversion to speech and the duration of the corresponding speech in the original speech is within a preset range.
3. The video processing method according to claim 2, wherein, The step of using the first machine learning model to translate the sentences of the first text corresponding to each speech segment in the original speech into sentences of the second text based on the duration and translation target of each speech segment includes: Using the first machine learning model, based on the duration of each speech segment in the original speech, the number of syllables in the words of the second language, and the translation target, the sentences of the first text corresponding to each speech segment are translated into sentences of the second text. The estimated duration of the sentences of the second text after being converted into speech is determined based on the number of words in the sentences of the second text and the number of syllables in each word.
4. The video processing method according to any one of claims 1-3, wherein, The step of employing a second machine learning model to generate the target speech for the one or more characters, based on the vocal features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text, includes: Based on the correspondence between the one or more roles and multiple statements in the second text, the second text is divided into texts corresponding to each role; Using the second machine learning model, the speech of each character in the target speech is generated based on the corresponding text and the voice features of each character.
5. The video processing method according to any one of claims 1-4, wherein, The step of employing a second machine learning model to generate the target speech for the one or more characters, based on the vocal features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text, includes: Determine the corresponding speech segment in the original speech for each statement in the second text; Identify the emotional information of the speech segments corresponding to each sentence in the second text, and use it as the emotional information of each sentence in the second text; Using the second machine learning model, the target speech of the one or more characters is generated based on the voice features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, the emotional information of each sentence in the second text, and the second text. The voice features of each character include at least one of acoustic features and prosodic features.
6. The video processing method according to any one of claims 1-5, further comprising: Based on at least one of the following correspondences—the correspondence between multiple images in the video and multiple sentences in the first text, and the correspondence between multiple sentences in the first text and multiple speech segments in the original audio—the correspondence between the one or more characters and multiple sentences in the first text is determined. Based on the correspondence between multiple statements in the first text and multiple statements in the second text, and the correspondence between the one or more roles and multiple statements in the first text, the correspondence between the one or more roles and multiple statements in the second text is determined.
7. The video processing method according to claim 6, wherein, The step of determining the correspondence between the one or more characters and the multiple statements in the first text based on at least one of the following correspondences: a first correspondence between multiple images in the video and multiple statements in the first text, and a second correspondence between multiple statements in the first text and multiple speech segments in the original audio. Identify whether each of the plurality of images contains a face that is speaking; Based on the image of the speaking face and the image of each character's face, determine the character who is speaking; Based on the correspondence between the multiple images and multiple sentences in the first text, and the image where the speaking character is located, the correspondence between the speaking character and multiple sentences in the first text is determined.
8. The video processing method according to claim 7, wherein, The step of determining the correspondence between the one or more characters and the multiple statements in the first text based on at least one of the following correspondences: a first correspondence between multiple images in the video and multiple statements in the first text, and a second correspondence between multiple statements in the first text and multiple speech segments in the original speech; further includes: Determine the corresponding speech segment in the original speech for each image that does not contain an image of a speaking face; Based on the sound features of the speech segments corresponding to each image of the non-speaking face image and the sound features of each character, determine the character corresponding to each image of the non-speaking face image; Based on the correspondence between each image of the non-speaking face and a statement in the first text, and the role corresponding to each image of the non-speaking face, determine the correspondence between the role corresponding to each image of the non-speaking face and one or more statements in the first text.
9. The video processing method according to claim 6, wherein, The step of determining the correspondence between the one or more characters and the multiple statements in the first text based on at least one of the following correspondences: a first correspondence between multiple images in the video and multiple statements in the first text, and a second correspondence between multiple statements in the first text and multiple speech segments in the original audio. Determine the speech segment in the original audio corresponding to each of the multiple images in the video; The character corresponding to each image is determined based on the sound features of the speech segments in the original speech corresponding to each image and the sound features of each character. Based on the correspondence between the multiple images and multiple sentences in the first text, and the role corresponding to each image, the correspondence between the one or more roles and multiple sentences in the first text is determined.
10. The video processing method according to any one of claims 1-9, further comprising: Extract multiple keyframes from the video; Identify multiple face images in the multiple keyframes; Duplicate face images are removed from the plurality of face images, and the one or more characters are determined based on the remaining one or more face images.
11. The video processing method according to claim 10, further comprising: Based on multiple reference images of one or more characters speaking in the video; The correspondence between the multiple reference images and multiple speech segments in the original speech is used to determine the reference speech segment corresponding to each character; The vocal characteristics of each character are determined based on the reference voice segments corresponding to each character.
12. The video processing method according to any one of claims 1-11, wherein, The step of fusing the images from the video with the target speech to generate the target video includes: Determine the acoustic features of the target speech; Based on the sound features of the target speech and the mapping relationship between the sound features and lip shape parameters, the lip shape in the video image is adjusted; The adjusted image is fused with the target speech to generate the target video.
13. The video processing method according to claim 12, wherein, The step of adjusting the lip movements in the video image based on the sound features of the target speech and the mapping relationship between the sound features and lip movement parameters includes: Based on the sound features of each character's speech in the target speech and the mapping relationship between the sound features and lip shape parameters, a lip shape image of each character is generated; Based on the correspondence between multiple images in the video and multiple speech segments in the target speech, the characters in the multiple images, and the lip-sync images of each character, the lip-sync images in the video are adjusted.
14. The video processing method according to any one of claims 1-13, wherein, The first text corresponding to the original voice of one or more characters in the video includes: In response to the video including subtitles, the subtitles are obtained as the first text; In response to the fact that the video does not include subtitles, the original speech is obtained from the video, and the original speech is converted into the first text using a third machine learning model.
15. The video processing method according to any one of claims 1-14, further comprising: Display the target video on the target page; In response to the user's playback operation on the target video, the target video is played.
16. A video processing apparatus, comprising: The acquisition module is configured to acquire a first text corresponding to the original speech of one or more characters in a video, wherein the original speech and the first text are expressed in a first language; The translation module is configured to translate the first text into a second text expressed in a second language; The generation module is configured to generate the target speech of the one or more characters based on the voice features of each character in the one or more characters, the correspondence between the one or more characters and multiple sentences in the second text, and the second text. The fusion module is configured to fuse the images of the video with the target speech to generate the target video.
17. An electronic device comprising: Memory; as well as A processor coupled to the memory, the processor being configured to perform the video processing method as described in any one of claims 1 to 15 based on instructions stored in the memory.
18. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video processing method of any one of claims 1 to 15.
19. A computer program product comprising: Instructions, wherein when executed by a processor, the instructions implement the video processing method according to any one of claims 1 to 15.
20. A computer program comprising: Instructions, wherein when executed by a processor, the instructions implement the video processing method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Video translation method, system and device and storage medium
CN112562721A
Dubbing method and device, electronic equipment and storage medium
CN115171645A
Video voice generation method and device, and storage medium
CN118678148A
Amplifier with common mode detection
KR1020200033190A
Video translation method, system and device, and storage medium
WO2022110354A1