Video processing method and device, electronic equipment, storage medium and program product
By identifying and repairing video text areas, and automatically translating and generating audio that keeps sound characteristics consistent, the problem of low video automatic translation and dubbing efficiency is solved, and efficient and natural video processing effects are achieved.
Patent Information
- Application Number
- CN202510803835.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-05
AI Technical Summary
The prior art is difficult to realize automatic translation and dubbing of videos in a short time, especially videos with subtitle files that cannot be obtained, and manual translation and dubbing are inefficient.
By identifying the text area in the video, removing the pixels corresponding to the text and repairing the background pixels, automatically translating the text and generating audio that keeps the sound characteristics consistent, automatic translation and dubbing of the video are achieved.
It improves the efficiency of video translation and dubbing, ensures that the picture is smooth and natural, does not affect the original content, the audience has no perception, and the audio and text match, improving the audio-visual effect.
Smart Images

Figure CN120602719A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a video processing method, device, electronic device, storage medium, and program product. Background Art
[0002] With the development of Internet technology, the dissemination of videos has become faster and more widespread. When videos are disseminated to other countries or regions, in order to provide local audiences with a better viewing experience, the language in the videos needs to be translated into the local language. Summary of the Invention
[0003] According to some embodiments of the present disclosure, a video processing method is provided, including: identifying a text area in a video, and determining a first text in the text area and pixels corresponding to the first text; removing pixels corresponding to the first text, and repairing background pixels corresponding to the first text; adding a second text translated based on the first text to the text area; replacing a first audio corresponding to the first text in the video with a second audio generated based on the second text, wherein the second audio maintains consistent sound characteristics of the same character in the first audio.
[0004] According to other embodiments of the present disclosure, a video processing device is provided, including: an identification module configured to identify a text area in a video and determine a first text in the text area and pixels corresponding to the first text; a repair module configured to remove pixels corresponding to the first text and repair background pixels corresponding to the first text; an adding module configured to add a second text translated based on the first text to the text area; and a replacement module configured to replace a first audio corresponding to the first text in the video with a second audio generated based on the second text, wherein the second audio maintains consistent sound characteristics with the same character in the first audio.
[0005] According to some further embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory coupled to the processor, for storing instructions, which, when executed by the processor, causes the processor to execute the video processing method of any embodiment of the present disclosure.
[0006] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the video processing method of any embodiment of the present disclosure is performed.
[0007] According to yet other embodiments of the present disclosure, a computer program product is provided, comprising: instructions, wherein when the instructions are executed by a processor, the processor is caused to perform the video processing method according to any one of the embodiments of the present disclosure.
[0008] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The following describes embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:
[0010] Figure 1 A schematic diagram showing a flow chart of a video processing method according to some embodiments of the present disclosure;
[0011] Figure 2 A schematic diagram showing a flow chart of a video processing method according to some other embodiments of the present disclosure;
[0012] Figure 3 A schematic diagram illustrating the architecture of a video processing system according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic diagram showing the structure of a video processing device according to some embodiments of the present disclosure is shown;
[0014] Figure 5 A schematic diagram showing the structure of an electronic device according to some embodiments of the present disclosure;
[0015] Figure 6 Schematic diagrams showing the structure of electronic devices according to other embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein.
[0017] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of the steps set forth in these embodiments should be interpreted as being merely exemplary and not limiting the scope of the present disclosure.
[0018] The term “including” and its variations used in the present disclosure are open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least in part based on.”
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0021] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0022] As short videos, skits, and other video formats become increasingly popular and shared, the demand for video translation and dubbing is increasing. Relying solely on manual translation and dubbing for large numbers of videos would be impossible in a short period of time. Automatic translation is possible using artificial intelligence (AI) technology, but this generally requires obtaining the video's subtitle file in advance, translating the subtitle file, and then combining it with the video. Automatic translation is difficult for videos without subtitle files. Video dubbing generally requires professional staff.
[0023] The present disclosure proposes a video processing method that can automatically identify a text area in a video, determine a first text in the text area and the pixels corresponding to the first text; remove the pixels corresponding to the first text, repair the background pixels corresponding to the first text, and then add the second text translated from the first text to the text area. In addition, the first audio corresponding to the first text in the video can be automatically replaced with a second audio generated based on the second text, and the second audio is consistent with the voice characteristics of the same character in the first audio. The video processing method disclosed in the present disclosure can realize automatic translation and automatic dubbing of text in a video, does not require a subtitle file, is applicable to any video, and improves the efficiency of video translation and dubbing. By identifying the text area, erasing the first text, and restoring the background pixels, the video picture can be made smoother and more natural without affecting the content of the picture itself. The addition of the second text will not produce a mismatch or discordant look, making the audience unaware of the modification of the text in the video, thereby improving the display effect of the video after adding the second text. In addition, the second audio is consistent with the voice characteristics of the same character in the first audio, more closely matching the image of the character, and cooperating with the second text to improve the overall audio-visual effect of the video.
[0024] The video processing method disclosed herein may be executed by a video processing device or an electronic device, and the video processing device may be implemented in software and / or hardware.
[0025] Figure 1 Flowcharts of some embodiments of the video processing method disclosed herein. Figure 1 As shown, the video processing method of this embodiment includes steps S102 to S108.
[0026] In step S102 , a text region in the video is identified, and a first text in the text region and pixels corresponding to the first text are determined.
[0027] The text region (text area) in a video can include subtitle areas, as well as other non-subtitle areas, such as text areas on signage in the video, without being limited to the examples listed. It is possible to identify only the subtitle area in the video. In this case, a recognition range can be set for the video image, and the text area can be identified within this range, reducing computational complexity and improving recognition efficiency. For example, only the lower half of the video image can be identified to determine whether a text area exists.
[0028] The first text may include a first subtitle text and may also include a first non-subtitle text. Based on the features of the text, a specific text in the text area, namely the first text, may be identified, and pixels corresponding to the first text in the video image may be determined.
[0029] In step S104 , the pixels corresponding to the first text are removed, and the background pixels corresponding to the first text are repaired.
[0030] The first text in the video is erased, and the image is reconstructed and repaired so that the image in the video appears to have no text added.
[0031] In step S106 , the second text obtained by translating the first text is added to the text area.
[0032] A machine learning model can be used to translate the first text to obtain a second text. The first subtitle text can be translated to obtain a second subtitle text, and the first non-subtitle text can be translated to obtain a second non-subtitle text. The second subtitle text and the second non-subtitle text can be added to the text area in different ways based on their relationship with the image in the video.
[0033] In step S108, the first audio corresponding to the first text in the video is replaced by the second audio generated based on the second text.
[0034] A second audio file can be generated based on the first audio file and the second text. The second audio file maintains the same voice characteristics as the character in the first audio file. A video may include one or more characters, each with different voice characteristics. The character corresponding to each speech segment (e.g., sentence) in the first audio file can be determined, and the voice characteristics of each character can be extracted. The second audio file is generated based on the voice characteristics of the character corresponding to each speech segment and the text corresponding to each speech segment in the second text file. The first audio file is removed from the video, and the second audio file is added to the video according to the time information of the first audio file to achieve dubbing of the video.
[0035] The video processing method of the above embodiment can automatically identify the text area in the video, determine the first text in the text area and the pixels corresponding to the first text; remove the pixels corresponding to the first text, and repair the background pixels corresponding to the first text, and then add the second text translated from the first text to the text area. In addition, the first audio corresponding to the first text in the video can be automatically replaced with the second audio generated based on the second text and the second audio is consistent with the sound characteristics of the same character in the first audio. The video processing method of the above embodiment can realize automatic translation and automatic dubbing of text in the video, without the need for subtitle files, and is applicable to any video, thereby improving the efficiency of video translation and dubbing. By identifying the text area, erasing the first text and restoring the background pixels, the video picture can be made smoother and more natural, without affecting the content of the picture itself. After adding the second text, there will be no mismatch or discordant feeling, so that the audience will not perceive the modification of the text in the video, thereby improving the display effect of the video after adding the second text. In addition, the second audio is consistent with the sound characteristics of the same character in the first audio, more closely matching the image of the character, and cooperating with the second text to improve the overall audio-visual effect of the video.
[0036] The following describes in detail how to identify a text region in a video and the first text in the text region.
[0037] In some embodiments, key frames are determined from multiple frames of a video, and a machine learning model is used to identify whether the key frames include text regions. For key frames that include text regions, optical character recognition (OCR) is used to identify the first text in the text regions. For example, the machine learning model can identify text regions based on an edge detection algorithm, but this is not limited to the examples given.
[0038] When recognizing only subtitles, the detection range can be determined based on the location of the text area in some key frames of the video. For the remaining key frames, the text area and the first text in the text area are identified within the detection range. If the detection range is part of a key frame, identifying the text area and the first text in the text area only within the detection range can improve recognition efficiency.
[0039] In order to further improve the accuracy of determining the first text, the OCR technology may be combined with the speech recognition technology to determine the first text.
[0040] In some embodiments, the text area includes the text area of each frame image in the multiple frames of the video, and determining the first text in the text area includes: using optical character recognition to identify the first reference text corresponding to the text area of each frame image; performing speech recognition on the video to determine the second reference text corresponding to each voice segment in the multiple voice segments; determining the correspondence between the multiple voice segments and the multiple frames of images based on the time information of each voice segment and the time information of each frame image; determining the first text based on the correspondence between the multiple voice segments and the multiple frames of images, the first reference text corresponding to the text area of each frame image, and the second reference text corresponding to each voice segment.
[0041] The first reference text corresponding to the text region in each keyframe can be identified. Based on the relationship between the keyframe and each image frame in the video, the first reference text corresponding to the text region in each image frame can be determined. The first reference text can include multiple text segments (e.g., sentences), and the multiple image frames corresponding to each text segment can be determined. Speech recognition, i.e., recognizing the first audio in the video, can determine the second reference text corresponding to each speech segment (e.g., sentence). By extracting the time information of the speech segments and the time information of each image frame in the video, the correspondence between multiple speech segments and multiple image frames can be determined. Furthermore, the correspondence between multiple text segments in the first reference text and the second reference text corresponding to the multiple speech segments can be determined. Combining the first and second reference texts can more accurately determine the first text. For example, ASR (Automatic Speech Recognition) can be used to determine the second reference text corresponding to each speech segment.
[0042] The method of the above embodiment combines OCR technology and speech recognition technology to recognize the first text in the text area of the video, thereby improving the accuracy of the first text recognition and thus improving the accuracy of subsequent translation and dubbing.
[0043] In some embodiments, determining the first text based on the correspondence between multiple speech segments and multiple frame images, the first reference text corresponding to the text area of each frame image and the second reference text corresponding to each speech segment includes: for each speech segment: for the multiple frame images corresponding to the speech segment, determining whether the first reference texts corresponding to the multiple frame images are the same, and determining whether the first reference texts corresponding to the multiple frame images are the same as the second reference text corresponding to the speech segment; in response to the first reference texts corresponding to the multiple frame images being different or the first reference texts corresponding to the multiple frame images being different from the second reference text corresponding to the speech segment, determining the text corresponding to the speech segment based on the semantic understanding of the context of the speech segment and / or the second reference text associated with the speech segment.
[0044] Based on the time information of each speech segment and the time information of each image frame, the correspondence between multiple speech segments and multiple image frames can be determined. For each speech segment, it can be determined whether the first reference text of the multiple image frames corresponding to the speech segment is the same. Typically, if the first reference text of multiple image frames is the same, or if the first reference text of most images in the multiple image frames is the same, text recognition errors may occur in some images, or speech and text may be out of sync. Therefore, if the first reference texts corresponding to multiple image frames are different, the first reference text with the highest probability of appearing in the first reference texts corresponding to the multiple image frames can be determined as the first target text. The first target text is compared with the second reference text corresponding to the speech segment. If they are different, the correct text corresponding to the speech segment can be determined based on the semantic understanding of the context of the speech segment and / or the second reference text associated with the speech segment. For example, a machine learning model can be used to understand at least one of the context of the speech segment and the context of the second reference text to determine the correct text. In addition, a machine learning model can also be used to understand the content of the video segment corresponding to the speech segment to determine the correct text.
[0045] In addition to subtitle areas, some images may also include non-subtitle areas. Based on the positions of the subtitle and non-subtitle areas in the image, the text in the subtitle area and the text in the non-subtitle area in the first reference text can be distinguished. The method of the above embodiment is used to determine the correct text for the text in the subtitle area. For each non-subtitle area, multiple texts corresponding to the non-subtitle area in multiple frames of the image can be compared, and the text with the highest probability of appearing can be selected as the text in the non-subtitle area.
[0046] The method of the above embodiment compares the first reference text obtained based on OCR recognition with the second reference text obtained based on speech recognition, and combines the understanding of the context to improve the accuracy of determining the text corresponding to each speech segment.
[0047] After the first text is recognized, it is necessary to erase the first text in the video and restore the image in the video.
[0048] In some embodiments, the pixels corresponding to the first text include pixels of the text in each frame of multiple frames of the video, and repairing the background pixels corresponding to the first text includes: using a first machine learning model to understand the content of each frame of the image, and repairing the background pixels corresponding to the text in each frame of the image based on the pixels surrounding the pixels of the text in each frame of the image.
[0049] For example, the first machine learning model is a video restoration model, which locates the text area through visual positioning or deep learning, and uses the surrounding pixel information to diffuse to the missing area to reconstruct the background after erasing the first text, which is not limited to the examples given.
[0050] By understanding the content of each frame of the image and reconstructing and repairing the background corresponding to the text in each frame of the image based on the pixels surrounding the pixels of the text in each frame of the image, the accuracy of the repair can be improved and the display effect of the video can be enhanced.
[0051] A machine learning model can be used to automatically translate the first text into a second text. If the first text includes a first subtitled text and a first non-subtitled text, each can be translated separately. The following describes a method for translating the first text into a second text, which can further improve translation accuracy.
[0052] In some embodiments, adding a second text obtained by translating the first text to the text area includes: using a second machine learning model to understand at least one of each voice segment in the video, the text corresponding to each voice segment in the first text, and the video segment corresponding to each voice segment, and determining the emotional information corresponding to each voice segment; using a fourth machine learning model to translate the text corresponding to each voice segment in the first text according to the emotional information corresponding to each voice segment, and obtain the text corresponding to each voice segment in the second text; adding the text corresponding to each voice segment in the second text to the subtitle area of the video, wherein the text area includes the subtitle area.
[0053] Different words can be translated for the same word under different emotions. Each speech segment (e.g., each sentence) can be understood based on at least one of its corresponding text, speech, or video. Furthermore, the emotional information corresponding to each speech segment can be determined by combining the context of at least one of the text, speech, or video corresponding to each speech segment. The fourth machine learning model can be a translation model, such as a large language model, which can translate the text corresponding to each speech segment in the first text based on the emotional information corresponding to each speech segment and the semantics of the text corresponding to each speech segment to obtain the text corresponding to each speech segment in the second text.
[0054] The text corresponding to each speech segment in the second text is the second subtitle text. The text corresponding to each speech segment in the second text can be added to the subtitle area of the corresponding image in the video based on the time information corresponding to each speech segment. A track (timeline) corresponding to the second subtitle text and a track (timeline) of the video without the first subtitle text can be automatically generated, and the second subtitle text and the video without the first subtitle text can be merged.
[0055] The method of the above embodiment understands at least one of each voice segment in the video, the text corresponding to each voice segment in the first text, and the video segment corresponding to each voice segment, and can more accurately determine the emotional information of each voice segment, and then translate the text corresponding to each voice segment according to the emotional information, which can improve the accuracy of the translation.
[0056] In some embodiments, using a fourth machine learning model to translate the text corresponding to each voice segment in the first text according to the emotional information corresponding to each voice segment includes: determining the subject type of the video; using the fourth machine learning model to translate the text corresponding to each voice segment in the first text according to the emotional information corresponding to each voice segment and the subject type of the video.
[0057] Videos of different subject matter (for example, finance, law, etc.) can correspond to certain specific words. Using the fourth machine learning model to refer to the emotional information corresponding to each voice clip and the subject matter of the video for translation can improve the accuracy of the translation and help viewers accurately understand the video content.
[0058] In some embodiments, the text area includes a subtitle area and a non-subtitle area, the first text includes a first subtitle text and a first non-subtitle text, the second text includes a second subtitle text and a second non-subtitle text, the first audio is the audio corresponding to the first subtitle text, and the second audio is the audio corresponding to the second subtitle text. Adding the second text translated based on the first text to the text area includes: generating an image of adding the second non-subtitle text in the non-subtitle area based on the second non-subtitle text and the image corresponding to the second non-subtitle text in the video.
[0059] The second subtitle text and the second audio can be synthesized with the video using tracks. For example, the track for the second subtitle text and the second audio are generated based on the time information corresponding to the second subtitle text and the second audio. The final video is synthesized using the track for the second subtitle text, the track for the second audio, and the track for the video after removing the first subtitle text and the first audio.
[0060] In order to make the fusion of the second non-subtitle text and the image more realistic and natural, an image generation model can be used to generate an image with the second non-subtitle text added based on the second non-subtitle text and the image corresponding to the second non-subtitle text in the video, thereby improving the display effect of the video.
[0061] The above embodiment describes how to process text in a video. The following describes how to process audio in a video.
[0062] In some embodiments, replacing the first audio corresponding to the first text in the video with the second audio generated based on the second text includes: separating the human voice and background sound in the first audio to obtain the first human voice audio and the first background audio; determining the text corresponding to the first human voice audio in the second text based on the text corresponding to the first human voice audio in the first text; generating the second human voice audio based on the first human voice audio and the text corresponding to the first human voice audio in the second text; and replacing the first human voice audio with the second human voice audio.
[0063] After acquiring the video, the audio in the video, namely the first audio, can be extracted. Since the first audio may contain human voice and background sound, the human voice and background sound must first be separated to obtain the first human voice audio and the first background audio. The first human voice audio can be divided into multiple voice segments, each of which can include one or more sentences. The character corresponding to each voice segment can be further determined, and based on the time information of each voice segment, the corresponding video segment (multiple frames) can be determined for each voice segment, completing the video preprocessing. Furthermore, after extracting the first text, the text corresponding to each voice segment in the first text can be determined. After translating the first text into the second text, the text corresponding to each voice segment in the second text can also be determined. These correspondences can then be used to generate the second human voice audio.
[0064] The method of the above embodiment separates the human voice and background sound in the first audio, generates a second human voice audio based on the first human voice audio and the text corresponding to the first human voice audio in the second text, and replaces the first human voice audio with the second human voice audio, which can improve the accuracy and effect of video dubbing.
[0065] In some embodiments, separating the human voice and background sound in the first audio to obtain the first human voice audio and the first background audio includes: transforming the first audio from the time domain to the frequency domain to obtain a spectrogram; extracting features of the human voice based on the spectrogram; determining a mask matrix based on the features of the human voice; and separating the human voice and background sound in the first audio based on the mask matrix and the spectrogram to obtain the first human voice audio and the first background audio.
[0066] For example, the first audio can be transformed from the time domain to the frequency domain using an STFT (Short-time Fourier Transform) to obtain a spectrogram, not limited to the examples listed. For example, a machine learning model can be used to extract features of human voices, such as a convolutional neural network or a Transformer model, not limited to the examples listed. For example, a machine learning model can be used to predict a mask matrix, such as a neural network model, not limited to the examples listed. The human voice and background sound in the first audio can be separated by multiplying the mask matrix with the spectrogram, and other operations.
[0067] The method of the above embodiment can improve the accuracy of separating human voice and background sound, thereby making the second human voice audio subsequently generated based on the first human voice audio more accurate, thereby improving the effect of dubbing the video.
[0068] In some embodiments, generating the second vocal audio based on the first vocal audio and the text corresponding to the first vocal audio in the second text includes: determining the role corresponding to each voice segment in multiple voice segments in the first vocal audio; determining the text corresponding to each role in the second text based on the role corresponding to each voice segment and the text corresponding to each voice segment in the second text; determining the sound features of each role based on the voice segment corresponding to each role in the first vocal audio; generating the second vocal audio based on the sound features of each role and the text corresponding to each role in the second text.
[0069] The character corresponding to each voice segment in the first vocal audio can be determined, and the character corresponding to each voice segment can be assigned a corresponding identifier to mark it. Furthermore, all audio segments corresponding to each character can be determined, and the voice features of each character can be extracted from all audio segments corresponding to each character. Based on the voice features of each character and the text corresponding to each character in the second text, vocal audio for each character in the second vocal audio can be generated.
[0070] The method in the above embodiment extracts each character's voice features from the corresponding voice clips in the first-person audio, and uses these features to generate each character's voiceover, resulting in the second-person audio. The second-person audio is consistent with the voice features of each character in the first-person audio, more closely matching the visual content of the video and providing a better audiovisual experience for the viewer.
[0071] In some embodiments, determining the role corresponding to each voice segment in the first human voice audio includes: extracting the voiceprint features of each voice segment; comparing the voiceprint features of multiple voice segments to determine one or more voice segments corresponding to each voiceprint feature; determining the role corresponding to each voice segment based on the one or more voice segments corresponding to each voiceprint feature, wherein each voiceprint feature corresponds to one role.
[0072] Each character's voice has its own unique characteristics, and voiceprint features can be used to distinguish different characters. Therefore, by comparing the voiceprint features of each voice segment in the first person's audio, it is possible to determine which ones belong to the same character.
[0073] The method of the above embodiment, based on the recognition of the voiceprint features of each voice segment in the first human voice audio, can automatically and accurately distinguish the characters corresponding to different voice segments, thereby providing a basis for subsequent more accurate generation of the second human voice audio.
[0074] The above embodiments describe how to separate the human voice from the background in the first audio and how to determine the role corresponding to each voice segment in the first audio. To further improve the accuracy of generating the second audio, the present disclosure also provides the following method for generating the second audio.
[0075] In some embodiments, generating the second human voice audio based on the sound features of each character and the text corresponding to each character in the second text includes: using a second machine learning model to understand at least one of each voice segment, the text corresponding to each voice segment in the first text, and the video segment corresponding to each voice segment, and determining the emotional information corresponding to each voice segment; using a third machine learning model to generate the second human voice audio based on the emotional information corresponding to each voice segment, the character corresponding to each voice segment, the sound features of each character, and the text corresponding to each voice segment in the second text.
[0076] The same sentence, spoken by the same person with different emotions, sounds different, and the audience's listening experience is also different. Therefore, at least one of each voice segment, the text corresponding to each voice segment in the first text, and the video segment corresponding to each voice segment can be first input into the second machine learning model to determine the emotional information corresponding to each voice segment. In order to improve the accuracy of the emotional information determination, at least one of the context of each voice segment, the context corresponding to each voice segment in the first text, and the context of the video segment corresponding to each voice segment can also be input into the second machine learning model to determine the emotional information corresponding to each voice segment. The second machine learning model can be a large multimodal model that can understand information in multiple modalities. Emotional labels can be added to each voice segment. For example, emotional labels can include calm, sad, happy, angry, surprised, disgusted, afraid, crying, etc., not limited to the examples given.
[0077] The third machine learning model can generate audio from the text, such as a TTS (Text to Speech) model. Using the third machine learning model, the text corresponding to each speech segment in the second text can be converted into the corresponding second human voice audio based on the correspondence between multiple speech segments, multiple characters, various emotional information, and multiple sound features.
[0078] The method of the above embodiment, in the process of generating the second human voice audio, not only refers to the sound characteristics of each character but also combines the emotional information of each voice clip, so that the generated second human voice audio is more natural and consistent with the real emotion, thereby improving the overall audio-visual effect of the video.
[0079] In some embodiments, determining the sound features of each character based on the speech segments corresponding to each character in the first person voice audio includes, for each character: determining, from the speech segments corresponding to the character in the first person voice audio, a speech segment corresponding to each emotional information in one or more emotional information; and determining, based on the speech segments corresponding to each emotional information, the timbre coding and rhythmic parameters corresponding to the character under each emotional information as the sound features of the character.
[0080] For each character, the corresponding speech segment for each emotional information can be determined, and then the corresponding timbre code and prosodic parameters for each emotional information can be determined as the voice characteristics of the character. Prosodic parameters include, but are not limited to, speech rate and pitch. The fifth machine learning model can be used to extract the timbre feature vector and prosodic parameters for each character under each emotional information to obtain the voice characteristics of the character.
[0081] The method of the above embodiment determines timbre coding and prosody parameters according to different emotional information of different characters as sound features, thereby improving the accuracy of the sound features of each character and further improving the accuracy of generating the second human voice audio.
[0082] The above embodiments implement the processing of human voice audio, and the following method can also be used to process background audio.
[0083] In some embodiments, replacing the first audio corresponding to the first text in the video with the second audio generated based on the second text also includes: detecting whether the first background audio includes a voice segment in a first language, wherein the first language is the language corresponding to the first text; in response to the first background audio including the voice segment in the first language, generating a voice segment in a second language based on the voice segment in the first language, wherein the second language is the language corresponding to the second text, and the sound features of the voice segment in the first language and the voice segment in the second language are consistent; generating second background audio based on the voice segment in the second language; and replacing the first background audio with the second background audio.
[0084] The background audio may include a speech segment in the first language, such as the sound of a person speaking in the background. The text corresponding to the speech segment in the first language can be determined using speech recognition technology. The text corresponding to the speech segment in the first language is translated into text in the second language. Referring to the method of the above embodiment, the sound features corresponding to the speech segment in the first language can be determined, and a speech segment in the second language is generated based on the text in the second language and the sound features corresponding to the speech segment in the first language, thereby generating the second background audio. If the background audio does not include a speech segment in the first language, the background audio may not be processed.
[0085] The second human voice audio and the second background audio are fused according to time information corresponding to the second human voice audio and the second background audio to obtain the second audio.
[0086] The method of the above embodiment can further translate and dub the human voice in the background audio when the separation between the human voice and the background is inaccurate, or when the background includes other human voices, so that the human voice in the background audio is also expressed in the second language, thereby improving the audio-visual effect of the video.
[0087] The following combination Figure 2 and 3 Describe the overall solution of the video processing method disclosed in the present invention.
[0088] Figure 2 Flowcharts of some embodiments of the video processing method disclosed herein. Figure 2 As shown, the video processing method of this embodiment includes steps S202 to S222.
[0089] In step S202, a first audio is extracted from the video, and the human voice and background sound in the first audio are separated to obtain a first human voice audio and a first background audio.
[0090] like Figure 3 As shown, the overall framework of video processing includes a module 301 for separating human voice and background sound, and a module 302 for determining multiple audio tracks of a first human voice audio and a first background audio to distinguish human voice from background sound.
[0091] In step S204, the role corresponding to each of the multiple voice segments in the first vocal audio is determined.
[0092] In step S206 , a text region in the video is identified, and a first text in the text region and pixels corresponding to the first text are determined.
[0093] In the process of determining the first text, voice recognition combined with OCR technology can be used to improve the accuracy of the first text recognition, that is, using Figure 3 The speech recognition module 303 and the OCR module 304 determine the first text.
[0094] In step S208 , the pixels corresponding to the first text are removed, and the background pixels corresponding to the first text are repaired.
[0095] The subtitle erasing module 305 can be used to remove pixels corresponding to the first subtitle text in the first text. If the first non-subtitle text needs to be removed, the non-subtitle erasing module can be used to remove the first non-subtitle text. The subtitle track generation module 307 can also be used to automatically generate a subtitle track (timeline) for the subsequent addition of the second subtitle text to the video.
[0096] In step S210, emotional information corresponding to each voice segment in the first human voice audio is determined.
[0097] like Figure 3 As shown, the emotion tag extraction module 306 may be used to determine the emotion information corresponding to each speech segment in the first human voice audio.
[0098] In step S212, the text corresponding to each speech segment in the first text is translated according to the emotion information corresponding to each speech segment, and the text corresponding to each speech segment in the second text is obtained as the second subtitle text.
[0099] In step S214, the first non-subtitle text in the first text is translated to obtain a second non-subtitle text.
[0100] like Figure 3 As shown, the corpus extraction module 308 is used to obtain the first subtitle text and the first non-subtitle text, and input them into the machine translation module 309 to obtain the second subtitle text and the second non-subtitle text. The translation evaluation module 310 can also be used to evaluate the translation of the second subtitle text and the second non-subtitle text.
[0101] In step S216, an image with the second non-subtitle text added to the non-subtitle area is generated based on the second non-subtitle text and the image corresponding to the second non-subtitle text in the video.
[0102] In step S218, the voice feature of each character is determined based on each voice segment in the first human voice audio and the character corresponding to each voice segment.
[0103] like Figure 3 As shown, the voice feature extraction module 311 can be used to extract the timbre code and rhythm parameters corresponding to each character under each emotion information as the voice feature of the character. The character and emotion module 312 can store the character and emotion information corresponding to each voice segment.
[0104] In step S220, the second vocal audio is generated based on the emotional information corresponding to each voice segment in the first vocal audio, the role corresponding to each voice segment, the sound characteristics of each role, and the text corresponding to each voice segment in the second text.
[0105] like Figure 3 As shown, the speech synthesis module 313 can synthesize the second human voice audio according to the sound characteristics of each character, the emotional information corresponding to each speech segment, and the character corresponding to each speech segment.
[0106] In step S222, a translated and dubbed video is generated according to the time information of the second subtitle text, the time information of the second human voice audio, the time information of the first background audio, and the time information of the video.
[0107] like Figure 3 As shown, module 314 is used to store the time information or track information of the second subtitle text, the second human voice audio, the first background audio and the video, and the video synthesis module 315 can be used to synthesize the video.
[0108] If the first background audio is processed according to the aforementioned embodiment to obtain the second background audio, a translated and dubbed video is generated based on the time information of the second subtitle text, the time information of the second vocal audio, the time information of the second background audio, and the time information of the video.
[0109] Through the method of the above embodiment, the efficiency of video translation and dubbing is improved, and the audio-visual effect of the video is improved, bringing a better experience to users.
[0110] The present disclosure also provides a video processing device, Figure 4 Provide a description.
[0111] Figure 4 FIG is a structural diagram of some embodiments of the video processing device disclosed herein. Figure 4 As shown, the video processing device 40 of this embodiment includes: an identification module 410 , a repair module 420 , an adding module 430 , and a replacing module 440 .
[0112] The recognition module 410 is configured to recognize a text region in a video, and determine a first text in the text region and pixels corresponding to the first text.
[0113] The repair module 420 is configured to remove pixels corresponding to the first text and repair background pixels corresponding to the first text.
[0114] The adding module 430 is configured to add the second text obtained by translating the first text to the text area.
[0115] The replacement module 440 is configured to replace the first audio corresponding to the first text in the video with the second audio generated based on the second text, wherein the second audio maintains the same voice characteristics as the same character in the first audio.
[0116] The video processing device of the above embodiment can automatically identify the text area in the video, determine the first text in the text area and the pixels corresponding to the first text; remove the pixels corresponding to the first text, and repair the background pixels corresponding to the first text, and then add the second text translated from the first text to the text area. In addition, it can also automatically replace the first audio corresponding to the first text in the video with the second audio generated based on the second text and make the second audio consistent with the voice characteristics of the same character in the first audio. The video processing device of the above embodiment can realize automatic translation and automatic dubbing of text in the video, without the need for subtitle files, and is applicable to any video, thereby improving the efficiency of video translation and dubbing. By identifying the text area, erasing the first text and restoring the background pixels, the video picture can be made smoother and more natural, without affecting the content of the picture itself. After adding the second text, there will be no mismatch or discordant feeling, so that the audience will not perceive the modification of the text in the video, and the display effect of the video after adding the second text is improved. In addition, the second audio is consistent with the voice characteristics of the same character in the first audio, more closely matching the image of the character, and cooperates with the second text to improve the overall audio-visual effect of the video.
[0117] In some embodiments, the pixels corresponding to the first text include the pixels of the text in each frame of multiple frames of the video, and the repair module 420 is configured to use a first machine learning model to understand the content of each frame of the image, and repair the background pixels corresponding to the text in each frame of the image based on the pixels around the pixels of the text in each frame of the image.
[0118] In some embodiments, the text area includes the text area of each frame image in the multiple frames of the video, and the recognition module 410 is configured to use optical character recognition to identify the first reference text corresponding to the text area of each frame image; perform speech recognition on the video to determine the second reference text corresponding to each voice segment in the multiple voice segments; determine the correspondence between the multiple voice segments and the multiple frames of images based on the time information of each voice segment and the time information of each frame image; determine the first text based on the correspondence between the multiple voice segments and the multiple frames of images, the first reference text corresponding to the text area of each frame image, and the second reference text corresponding to each voice segment.
[0119] In some embodiments, the recognition module 410 is configured to, for each speech segment: for multiple frames of images corresponding to the speech segment, determine whether the first reference texts corresponding to the multiple frames of images are the same, and determine whether the first reference texts corresponding to the multiple frames of images are the same as the second reference texts corresponding to the speech segment; in response to the first reference texts corresponding to the multiple frames of images being different or the first reference texts corresponding to the multiple frames of images being different from the second reference texts corresponding to the speech segment, determine the text corresponding to the speech segment based on the semantic understanding of the context of the speech segment and / or the second reference text associated with the speech segment.
[0120] In some embodiments, the replacement module 440 is configured to separate the human voice and background sound in the first audio to obtain the first human voice audio and the first background audio; determine the text corresponding to the first human voice audio in the second text based on the text corresponding to the first human voice audio in the first text; generate the second human voice audio based on the first human voice audio and the text corresponding to the first human voice audio in the second text; and replace the first human voice audio with the second human voice audio.
[0121] In some embodiments, the replacement module 440 is configured to detect whether the first background audio includes a speech segment in a first language, wherein the first language is the language corresponding to the first text; in response to the first background audio including the speech segment in the first language, generate a speech segment in a second language based on the speech segment in the first language, wherein the second language is the language corresponding to the second text, and the sound features of the speech segment in the first language are consistent with those of the speech segment in the second language; generate second background audio based on the speech segment in the second language; and replace the first background audio with the second background audio.
[0122] In some embodiments, the replacement module 440 is configured to transform the first audio from the time domain to the frequency domain to obtain a spectrogram; extract the characteristics of the human voice based on the spectrogram; determine the mask matrix based on the characteristics of the human voice; and separate the human voice and background sound in the first audio based on the mask matrix and the spectrogram to obtain the first human voice audio and the first background audio.
[0123] In some embodiments, the replacement module 440 is configured to determine the role corresponding to each voice segment in a plurality of voice segments in the first human voice audio; determine the text corresponding to each role in the second text based on the role corresponding to each voice segment and the text corresponding to each voice segment in the second text; determine the sound features of each role based on the voice segment corresponding to each role in the first human voice audio; and generate the second human voice audio based on the sound features of each role and the text corresponding to each role in the second text.
[0124] In some embodiments, the replacement module 440 is configured to extract the voiceprint features of each voice segment; compare the voiceprint features of multiple voice segments to determine one or more voice segments corresponding to each voiceprint feature; and determine the role corresponding to each voice segment based on the one or more voice segments corresponding to each voiceprint feature, wherein each voiceprint feature corresponds to one role.
[0125] In some embodiments, the replacement module 440 is configured to use a second machine learning model to understand at least one of each voice segment, the text corresponding to each voice segment in the first text, and the video segment corresponding to each voice segment, and determine the emotional information corresponding to each voice segment; and use a third machine learning model to generate a second human voice audio based on the emotional information corresponding to each voice segment, the role corresponding to each voice segment, the sound characteristics of each role, and the text corresponding to each voice segment in the second text.
[0126] In some embodiments, the replacement module 440 is configured to, for each character: determine, from the voice segment corresponding to the character in the first human voice audio, a voice segment corresponding to each of one or more emotional information; and, based on the voice segment corresponding to each emotional information, determine the timbre coding and rhythmic parameters corresponding to the character under each emotional information as the voice feature of the character.
[0127] In some embodiments, the adding module 430 is configured to use a second machine learning model to understand at least one of each voice segment in the video, the text corresponding to each voice segment in the first text, and the video segment corresponding to each voice segment, and determine the emotional information corresponding to each voice segment; use a fourth machine learning model to translate the text corresponding to each voice segment in the first text according to the emotional information corresponding to each voice segment, and obtain the text corresponding to each voice segment in the second text; add the text corresponding to each voice segment in the second text to the subtitle area of the video, wherein the text area includes the subtitle area.
[0128] In some embodiments, the adding module 430 is configured to determine the subject type of the video; and use the fourth machine learning model to translate the text corresponding to each voice segment in the first text according to the emotional information corresponding to each voice segment and the subject type of the video.
[0129] In some embodiments, the text area includes a subtitle area and a non-subtitle area, the first text includes a first subtitle text and a first non-subtitle text, the second text includes a second subtitle text and a second non-subtitle text, the first audio is the audio corresponding to the first subtitle text, and the second audio is the audio corresponding to the second subtitle text. The adding module 430 is configured to generate an image with the second non-subtitle text added to the non-subtitle area based on the second non-subtitle text and the image corresponding to the second non-subtitle text in the video.
[0130] The present disclosure also provides an electronic device, Figure 5 and 6 Provide a description. Figure 5 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0131] like Figure 5 As shown, the electronic device 5 includes a processor 52; and a memory 51 coupled to the processor 52, for storing instructions, which, when executed by the processor 52, enable the processor 52 to execute the video processing method of any embodiment of the present disclosure.
[0132] The electronic device of the above embodiment can automatically identify the text area in the video, determine the first text in the text area and the pixels corresponding to the first text; remove the pixels corresponding to the first text, and repair the background pixels corresponding to the first text, and then add the second text translated from the first text to the text area. In addition, it can also automatically replace the first audio corresponding to the first text in the video with the second audio generated based on the second text and make the second audio consistent with the voice characteristics of the same character in the first audio. The electronic device of the above embodiment can realize automatic translation and automatic dubbing of text in the video, without the need for subtitle files, and is applicable to any video, thereby improving the efficiency of video translation and dubbing. By identifying the text area, erasing the first text and restoring the background pixels, the video picture can be made smoother and more natural, without affecting the content of the picture itself. After adding the second text, there will be no mismatch or discordant feeling, so that the audience will not perceive the modification of the text in the video, and the display effect of the video after adding the second text is improved. In addition, the second audio is consistent with the voice characteristics of the same character in the first audio, more closely matching the image of the character, and cooperating with the second text to improve the overall audio-visual effect of the video.
[0133] Memory 51 is used to store one or more computer-readable instructions. Memory 51 may include any combination of various forms of computer-readable storage media, such as volatile and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 51 may store, for example, an operating system, applications, a boot loader, databases, and other programs, as well as various applications and data.
[0134] The processor 52 is configured to execute computer-readable instructions to implement the method described in any of the aforementioned embodiments. Detailed implementations of the various steps of the method can be found in the aforementioned embodiments, and repetitive details are omitted here.
[0135] The processor 52 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) may be of X86 or ARM architecture, etc.
[0136] The processor 52 and the memory 51 can communicate with each other directly or indirectly. For example, the processor 52 and the memory 51 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 52 and the memory 51 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0137] It should be noted that Figure 5 The components of the electronic device 5 shown are merely exemplary and non-limiting. The electronic device 5 may further include other components according to actual application requirements. The processor 52 may control the other components in the electronic device 5 to perform desired functions.
[0138] The electronic device 5 may be implemented by software, firmware and / or hardware, and may be integrated into a device installed with relevant application programs.
[0139] Figure 6 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.
[0140] Figure 6 The electronic device 6 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.
[0141] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital televisions and desktop computers, etc.
[0142] like Figure 6As shown, the central processing unit (CPU) 61 executes various processes according to the program stored in the read-only memory (ROM) 62 or the program loaded from the storage unit 68 to the random access memory (RAM) 63. In the RAM 63, data required when the CPU 61 executes various processes is stored as needed. The central processing unit is merely an example, and it may also be other types of processors, such as the various processors described above. The ROM 62, RAM 63 and storage unit 68 may be various forms of computer-readable storage media. It should be noted that although Figure 6 ROM 62, RAM 63 and storage portion 68 are shown separately in FIG, but one or more of them may be combined or located in the same or different memory or storage modules.
[0143] The CPU 61, the ROM 62, and the RAM 63 are connected to one another via a bus 64. To the bus 64, an input / output interface 65 is also connected.
[0144] The following components are connected to the input / output interface 65: an input section 66 such as a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output section 67 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage section 68 including a hard disk, a magnetic tape, etc.; and a communication section 69 including a network interface card such as a LAN card, a modem, etc. The communication section 69 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 6 Some of the electronic devices 6 are shown to communicate via a bus 64, but they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of wireless networks and wired networks.
[0145] A drive 610 is also connected to the input / output interface 65 as needed. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 610 as needed so that a computer program read therefrom is installed in the storage section 68 as needed.
[0146] When the series of processing described above is implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 611 .
[0147] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. The present disclosure also provides a computer program product, comprising: instructions, wherein when the instructions are executed by a processor, the processor executes the video processing method of any embodiment of the present disclosure.
[0148] The computer program product of the above embodiment can automatically identify the text area in the video, determine the first text in the text area and the pixels corresponding to the first text; remove the pixels corresponding to the first text, and repair the background pixels corresponding to the first text, and then add the second text translated from the first text to the text area. In addition, the first audio corresponding to the first text in the video can be automatically replaced with the second audio generated based on the second text and the second audio is consistent with the sound characteristics of the same character in the first audio. The computer program product of the above embodiment can realize the automatic translation and automatic dubbing of the text in the video, without the need for subtitle files, and is applicable to any video, thereby improving the efficiency of video translation and dubbing. By identifying the text area, erasing the first text and restoring the background pixels, the video picture can be made smoother and more natural, without affecting the content of the picture itself. After adding the second text, there will be no mismatch or discordant feeling, so that the audience will not perceive the modification of the text in the video, and the display effect of the video after adding the second text is improved. In addition, the second audio is consistent with the sound characteristics of the same character in the first audio, which is closer to the image of the character, and cooperates with the second text to improve the overall audio-visual effect of the video.
[0149] For example, some embodiments of the present disclosure include a computer program product that, when executed on a computer, causes the computer to implement the method described in any of the aforementioned embodiments. The computer program product includes computer instructions carried on a computer-readable medium, including program code for executing the method shown in the flowchart. In such embodiments, the computer instructions can be downloaded and installed from a network via the communication section 69, or installed from the storage section 68, or installed from the ROM 62. When the computer program is executed by the CPU 61, the method of the embodiment of the present disclosure is performed.
[0150] The present disclosure further provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the processor is caused to execute the video processing method according to any embodiment of the present disclosure.
[0151] The computer-readable storage medium of the above embodiment can automatically identify the text area in the video, determine the first text in the text area and the pixels corresponding to the first text; remove the pixels corresponding to the first text, and repair the background pixels corresponding to the first text, and then add the second text translated from the first text to the text area. In addition, the first audio corresponding to the first text in the video can be automatically replaced with the second audio generated based on the second text and the second audio is consistent with the voice characteristics of the same character in the first audio. The computer-readable storage medium of the above embodiment can realize automatic translation and automatic dubbing of text in the video, without the need for subtitle files, and is applicable to any video, thereby improving the efficiency of video translation and dubbing. By identifying the text area, erasing the first text and restoring the background pixels, the video picture can be made smoother and more natural, without affecting the content of the picture itself. After adding the second text, there will be no mismatch or discordant feeling, so that the audience will not perceive the modification of the text in the video, and the display effect of the video after adding the second text is improved. In addition, the second audio is consistent with the voice characteristics of the same character in the first audio, more closely matching the image of the character, and cooperating with the second text to improve the overall audio-visual effect of the video.
[0152] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.
[0153] The computer readable medium may be a computer readable storage medium, or a computer readable signal medium, or any combination of the two.
[0154] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. Computer instructions are stored on the computer-readable storage medium, and when the instructions are executed by the processor, the method described in any of the aforementioned embodiments is implemented.
[0155] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0156] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0157] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the method described in any of the aforementioned embodiments. For example, the instructions may be embodied as computer program codes.
[0158] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In situations involving a remote computer, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0160] The functions described above may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0161] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A video processing method, comprising: Identifying a text region in a video, and determining a first text in the text region and pixels corresponding to the first text; removing pixels corresponding to the first text and repairing background pixels corresponding to the first text; adding a second text obtained by translating the first text into the text area; The first audio corresponding to the first text in the video is replaced with the second audio generated based on the second text, wherein the second audio is consistent with the voice characteristics of the same character in the first audio.
2. The video processing method according to claim 1, wherein: The pixels corresponding to the first text include pixels of the text in each frame of the multiple frames of the video, and the repairing of background pixels corresponding to the first text includes: The first machine learning model is used to understand the content of each frame of the image, and the background pixels corresponding to the text in each frame of the image are repaired based on the pixels surrounding the pixels of the text in each frame of the image.
3. The video processing method according to claim 1, wherein: The text region includes a text region of each frame image in the multiple frames of the video, and determining the first text in the text region includes: Using optical character recognition, identifying a first reference text corresponding to the text area of each frame of image; Performing speech recognition on the video to determine a second reference text corresponding to each speech segment in a plurality of speech segments; Determining a correspondence between the plurality of speech segments and the plurality of image frames according to the time information of each speech segment and the time information of each image frame; The first text is determined according to the correspondence between the multiple voice segments and the multiple frames of images, the first reference text corresponding to the text area of each frame of image, and the second reference text corresponding to each voice segment.
4. The video processing method according to claim 3, wherein: The determining of the first text according to the correspondence between the plurality of speech segments and the plurality of image frames, the first reference text corresponding to the text region of each image frame, and the second reference text corresponding to each speech segment includes, for each speech segment: determining, for a plurality of image frames corresponding to the speech segment, whether first reference texts corresponding to the plurality of image frames are identical, and determining whether the first reference texts corresponding to the plurality of image frames are identical to a second reference text corresponding to the speech segment; In response to the first reference text corresponding to the multiple-frame images being different or the first reference text corresponding to the multiple-frame images being different from the second reference text corresponding to the speech segment, the text corresponding to the speech segment is determined based on the semantic understanding of the context of the speech segment associated with the speech segment and / or the second reference text.
5. The video processing method according to claim 1, wherein: Replacing the first audio corresponding to the first text in the video with the second audio generated based on the second text includes: Separating the human voice and the background sound in the first audio to obtain a first human voice audio and a first background audio; Determining, based on the text corresponding to the first human voice audio in the first text, the text corresponding to the first human voice audio in the second text; generating a second human voice audio according to the first human voice audio and the text corresponding to the first human voice audio in the second text; The first vocal audio is replaced with the second vocal audio.
6. The video processing method according to claim 5, wherein: The step of replacing the first audio corresponding to the first text in the video with the second audio generated based on the second text further includes: detecting whether the first background audio includes a speech segment in a first language, wherein the first language is a language corresponding to the first text; In response to the first background audio including a speech segment in the first language, generating a speech segment in a second language based on the speech segment in the first language, wherein the second language is the language corresponding to the second text, and the speech segment in the first language and the speech segment in the second language have consistent sound features; generating a second background audio according to the speech segment of the second language; The first background audio is replaced by the second background audio.
7. The video processing method according to claim 5, wherein: The separating the human voice and the background sound in the first audio to obtain the first human voice audio and the first background audio includes: transforming the first audio from the time domain to the frequency domain to obtain a spectrogram; Extracting features of the human voice according to the spectrogram; Determining a mask matrix according to the characteristics of the human voice; The human voice and the background sound in the first audio are separated according to the mask matrix and the spectrogram to obtain a first human voice audio and a first background audio.
8. The video processing method according to claim 5, wherein: The generating the second human voice audio according to the first human voice audio and the text corresponding to the first human voice audio in the second text includes: Determining a role corresponding to each of the plurality of voice segments in the first human voice audio; Determining a text corresponding to each character in the second text according to the character corresponding to each voice segment and the text corresponding to each voice segment in the second text; Determining a voice feature of each character according to a voice segment corresponding to each character in the first human voice audio; The second human voice audio is generated according to the voice features of each character and the text corresponding to each character in the second text.
9. The video processing method according to claim 8, wherein: Determining the role corresponding to each voice segment in the first human voice audio includes: Extracting voiceprint features of each speech segment; Comparing the voiceprint features of the multiple voice segments to determine one or more voice segments corresponding to each voiceprint feature; According to one or more voice segments corresponding to each voiceprint feature, the role corresponding to each voice segment is determined, wherein each voiceprint feature corresponds to one role.
10. The video processing method according to claim 8, wherein: Generating the second human voice audio according to the voice feature of each character and the text corresponding to each character in the second text includes: Using a second machine learning model, understanding at least one of each of the voice segments, the text corresponding to each of the voice segments in the first text, and the video segment corresponding to each of the voice segments, to determine emotional information corresponding to each of the voice segments; Using a third machine learning model, the second human voice audio is generated based on the emotional information corresponding to each voice segment, the role corresponding to each voice segment, the sound characteristics of each role, and the text corresponding to each voice segment in the second text. The video processing method according to claim 10 , wherein: Determining the voice feature of each character according to the voice segment corresponding to each character in the first human voice audio includes, for each character: Determining, from the voice segment corresponding to the character in the first vocal audio, a voice segment corresponding to each of the one or more types of emotional information; According to the speech segments corresponding to each type of emotional information, the timbre coding and rhythm parameters corresponding to the character under each type of emotional information are determined as the voice features of the character.
12. The video processing method according to any one of claims 1 to 11, wherein: The adding the second text obtained by translating the first text into the text area includes: Using a second machine learning model, understand at least one of each voice segment in the video, the text corresponding to each voice segment in the first text, and the video segment corresponding to each voice segment, and determine the emotional information corresponding to each voice segment; Using a fourth machine learning model, translating the text corresponding to each speech segment in the first text according to the emotional information corresponding to each speech segment, to obtain the text corresponding to each speech segment in the second text; Adding text corresponding to each voice segment in the second text to a subtitle area of the video, wherein the text area includes the subtitle area.
13. The video processing method according to claim 12, wherein: The translating, using the fourth machine learning model and based on the emotion information corresponding to each speech segment, the text corresponding to each speech segment in the first text includes: Determining the subject matter type of the video; The fourth machine learning model is used to translate the text corresponding to each voice segment in the first text according to the emotional information corresponding to each voice segment and the subject matter type of the video.
14. The video processing method according to any one of claims 1 to 11, wherein: The text area includes a subtitle area and a non-subtitle area, the first text includes a first subtitle text and a first non-subtitle text, the second text includes a second subtitle text and a second non-subtitle text, the first audio is the audio corresponding to the first subtitle text, and the second audio is the audio corresponding to the second subtitle text, and adding the second text obtained by translating the first text to the text area includes: An image in which the second non-subtitle text is added to the non-subtitle area is generated according to the second non-subtitle text and an image corresponding to the second non-subtitle text in the video.
15. A video processing device, comprising: a recognition module configured to recognize a text region in a video and determine a first text in the text region and pixels corresponding to the first text; a repair module, configured to remove pixels corresponding to the first text and repair background pixels corresponding to the first text; an adding module configured to add a second text obtained by translating the first text into the text area; The replacement module is configured to replace the first audio corresponding to the first text in the video with the second audio generated based on the second text, wherein the second audio is consistent with the voice characteristics of the same character in the first audio.
16. An electronic device comprising: processor; as well as A memory coupled to the processor is used to store instructions, and when the instructions are executed by the processor, the processor is caused to perform the video processing method according to any one of claims 1 to 14.
17. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the processor is caused to execute the video processing method according to any one of claims 1 to 14.
18. A computer program product comprising: The instruction, wherein when the instruction is executed by a processor, the processor is caused to perform the video processing method according to any one of claims 1 to 14.