Video dubbing method and device, electronic equipment and storage medium
By correcting the difference between the audio duration and the subtitle display duration during video dubbing, a third audio is generated, which solves the problem of dubbing duration mismatch, achieves high-quality audio-visual synchronization, and avoids video file corruption.
Patent Information
- Application Number
- CN202511097706.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-28
AI Technical Summary
In existing video dubbing technology, the dubbing duration does not match the duration of the original video clip, resulting in audio-visual asynchrony. Furthermore, existing frame interpolation methods may damage the original video file or cause unnatural, prolonged silences between the dubbing and the video, affecting the viewing experience.
By separating the audio and generating subtitle text with timecode, the subtitle text is regularized according to the difference between the audio duration and the subtitle display duration, generating a third subtitle text, and generating a third audio through a speech synthesis engine, and finally synthesizing the target video to ensure audio-visual synchronization.
This avoids damaging the original video file, improves the quality and synchronization of the video dubbing, and ensures a natural and smooth audio-visual experience.
Smart Images

Figure CN120856931A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a video dubbing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Video dubbing refers to replacing dialogue in the source language of a video with dialogue in the target language to facilitate understanding by audiences from different language backgrounds. However, because different languages often differ in text length and pronunciation duration when expressing the same meaning, the dubbing duration may not match the original video's allocation of time for that dialogue segment, leading to audio-visual desynchronization.
[0003] Currently, to address the issue of mismatch between the dubbing duration and the original video clip duration during video dubbing, especially when the dubbing duration exceeds the original duration, the common practice is to attempt to align the original video by interpolating frames.
[0004] Using existing video frame interpolation methods can damage the original video file, and the computational logic is complex and the running efficiency is low. In addition, it does not handle situations where the dubbing duration is shorter than the original video segment duration, resulting in unnatural long silences between the dubbing and the picture, which affects the viewing experience. Summary of the Invention
[0005] This invention provides a video dubbing method, apparatus, electronic device, and storage medium to address the shortcomings of low efficiency and low quality in existing video dubbing technologies.
[0006] This invention provides a video dubbing method, comprising the following steps: After separating the first audio from the video to be dubbed, the first audio is converted into text sentence by sentence to obtain the first subtitle text with timecode; Each of the first subtitle texts is translated into a second subtitle text, and the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text. Obtain the second audio corresponding to each second subtitle text; based on the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text, normalize the second subtitle text to obtain the third subtitle text, wherein the subtitle display duration is determined based on the timecode of the second subtitle text; Generate the third audio corresponding to each of the third subtitle texts; The final target video is obtained by combining all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed.
[0007] According to a video dubbing method provided by the present invention, a third subtitle text is obtained by normalizing the second subtitle text based on the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text, including: If the speech rate normalization condition is met, obtain the ratio coefficient between the audio duration and the subtitle display duration; When the ratio coefficient is determined to be greater than 1, a summary and sentence reduction operation is performed on the second subtitle text to obtain the third subtitle text; When the ratio coefficient is determined to be less than 1, a word filling operation is performed on the second subtitle text to obtain the third subtitle text.
[0008] According to a video dubbing method provided by the present invention, the step of performing a summary and sentence reduction operation on the second subtitle text to obtain the third subtitle text specifically includes: Based on the ratio coefficient and the number of characters in the second subtitle text, determine the number of characters to be reduced when performing the summary reduction operation; A first prompt instruction is generated based on the reduced word count; The first prompt instruction and the second subtitle text are input into the text generation model to obtain the third subtitle text output by the text generation model.
[0009] According to a video dubbing method provided by the present invention, a word-filling operation is performed on the second subtitle text to obtain the third subtitle text, specifically including: The number of characters to be filled is determined based on the ratio coefficient and the number of characters in the second subtitle text; A second prompt instruction is generated based on the number of characters to be filled; The second prompt instruction and the second subtitle text are input into the text generation model to obtain the third subtitle text output by the text generation model.
[0010] According to a video dubbing method provided by the present invention, the speech rate normalization conditions include: Determine that the audio duration of the valid audio segment corresponding to the second audio is greater than the subtitle display duration of the second subtitle text; or, The audio duration of the effective audio segment is determined to be less than the subtitle display duration of the second subtitle text, and the difference between the two is greater than a preset threshold.
[0011] According to a video dubbing method provided by the present invention, generating a third audio corresponding to each of the third subtitle texts includes: Obtain the current speech rate coefficient for audio conversion of each of the third subtitle texts, as recorded in the speech synthesis engine; The target speech rate coefficient is obtained by adjusting the current speech rate coefficient based on the speech rate normalization coefficient. The current speech rate coefficient in the speech synthesis engine is adjusted to the target speech rate coefficient, so as to use the speech synthesis engine to perform audio conversion on the third subtitle text and generate the third audio. The speech rate normalization coefficient is determined based on the ratio coefficient between the audio duration and the subtitle display duration.
[0012] According to a video dubbing method provided by the present invention, for any target third subtitle text, after adjusting the current speech rate coefficient based on the speech rate normalization coefficient to obtain the target speech rate coefficient, the method further includes: Obtain the first speech rate of the third audio corresponding to the preceding third subtitle text adjacent to the target third subtitle text in the timecode; Based on the timecode of the target third subtitle text, determine the dialogue audio segment of the first audio within the duration range of the timecode, so as to obtain the second speech rate of the dialogue audio segment; Based on the first speech rate and the second speech rate, a soft speech rate range is determined; When it is determined that the target speech rate for audio conversion of the target third subtitle text based on the target speech rate coefficient is outside the soft speech rate range, a prompt message is generated to guide the user to perform manual optimization. The target third subtitle text is any one of all the third subtitle texts.
[0013] According to a video dubbing method provided by the present invention, the step of aligning the second subtitle text with a time window based on the audio duration of the second audio to obtain the third subtitle text includes: Based on the difference between the audio duration of the second audio and the display duration of the second subtitle text, the second subtitle text is normalized to obtain the third subtitle text, and the display duration of the subtitle text is determined based on the timecode of the second subtitle text.
[0014] According to a video dubbing method provided by the present invention, for any target third subtitle text, a third audio corresponding to the target third subtitle text is generated, including: Determine the sentiment tag of the third audio corresponding to the target third subtitle text; The emotion tag and the target third subtitle text are input into the speech synthesis engine to generate the third audio.
[0015] According to a video dubbing method provided by the present invention, determining the emotion tag of the third audio corresponding to the target third subtitle text includes: Identify the target dialogue audio segment in the first audio, wherein the timecode of the target dialogue audio segment is the same as the timecode of the target third subtitle text; After identifying the target video segment in the video to be dubbed, the target video segment is subjected to frame extraction processing to obtain multiple video screenshots; the timecode of the target video segment is the same as the timecode of the target third subtitle text; The emotion tag of the third audio is determined based on at least one of the target dialogue audio segment, the multi-frame video screenshots, and the third subtitle text.
[0016] According to a video dubbing method provided by the present invention, determining the emotion tag of the third audio based on at least one of the target dialogue audio segment, the multi-frame video screenshots, and the third subtitle text includes: The target dialogue audio segment is input into the speech emotion recognition model to obtain the first emotion recognition result output by the speech emotion recognition model. The video screenshot is input into the image emotion recognition model to obtain the second emotion recognition result output by the image emotion recognition model; The target third subtitle text is input into the text emotion recognition model to obtain the third emotion recognition result output by the text emotion recognition model; The emotion label is determined based on at least one of the first emotion recognition result, the second emotion recognition result, and the third emotion recognition result.
[0017] According to a video dubbing method provided by the present invention, for any target third subtitle text, generating a third audio corresponding to the target third subtitle text, the method further includes: Determine the audio feature parameters of the third audio corresponding to the target third subtitle text, the audio feature parameters including volume parameters and / or pitch parameters; The audio feature parameters and the target third subtitle text are input into the speech synthesis engine to generate the third audio. The target third subtitle text is any one of all the third subtitle texts.
[0018] According to a video dubbing method provided by the present invention, the volume parameter of the third audio corresponding to the target third subtitle text is determined based on the following steps: Based on the waveform of the first audio, determine the first average volume of the first audio. Based on the timecode of the target third subtitle text, determine the dialogue audio segment of the first audio within the duration interval of the timecode, and determine the second average volume of the dialogue audio segment based on the waveform of the dialogue audio segment; The volume parameters of the third audio corresponding to the target third subtitle text are determined based on the volume weighting coefficient between the second average volume and the first average volume, and the preset volume level in the speech synthesis engine.
[0019] According to a video dubbing method provided by the present invention, the pitch parameters of the third audio corresponding to the target third subtitle text are determined based on the following steps: Obtain the spectrogram of the first audio, which is obtained by performing a Fourier transform on the global waveform of the first audio. The average first pitch of the first audio frequency is determined based on the spectrogram. Based on the timecode of the target third subtitle text, determine the dialogue audio segment of the first audio within the duration interval of the timecode, and determine the second pitch average of the dialogue audio segment based on the local spectrogram of the dialogue audio segment; Based on the pitch weight coefficient between the second pitch average and the first pitch average, and the preset pitch level in the speech synthesis engine, the pitch quantity parameter of the third audio corresponding to the target third subtitle text is determined.
[0020] According to a video dubbing method provided by the present invention, for any target third subtitle text, generating a third audio corresponding to the target third subtitle text, the method further includes: Identify the target dialogue audio segment in the first audio, wherein the timecode of the target dialogue audio segment is the same as the timecode of the target third subtitle text; The target dialogue audio segment is input into the timbre recognition model to obtain the gender and age information output by the timbre recognition model; The target speaker information is obtained by matching the gender and age information from the intelligent agent library, which stores different speaker information corresponding to different gender and age information. The target speaker information and the third subtitle text are input into the speech synthesis engine to generate the third audio. The target third subtitle text is any one of all the third subtitle texts.
[0021] According to a video dubbing method provided by the present invention, the video dubbing method further includes a step of providing a dubbing preview, specifically including: Step 1: Divide all the third subtitle texts into multiple subtitle batches; Step 2: In response to the user's listening request, the current subtitle batch is processed by speech synthesis to generate the current dialogue audio batch, and the current dialogue audio batch is sent to the client for playback; Step 3: During the playback of the current dialogue audio batch on the client, speech synthesis is performed on the adjacent next subtitle batch to generate the next dialogue audio batch. Step 4: After confirming that the current dialogue audio has finished playing, send the next batch of dialogue audio to the client for playback in the order of timecodes. Repeat steps 3 to 4 until all subtitle batches have been traversed or a user's stop listening instruction is received.
[0022] The present invention also provides a video dubbing device, comprising the following modules: The video parsing unit is used to separate the first audio from the video to be dubbed and then perform sentence-by-sentence text conversion on the first audio to obtain the first subtitle text with timecode; The subtitle translation unit is used to translate each of the first subtitle texts into a second subtitle text, wherein the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text. The text normalization unit is used to obtain the second audio corresponding to each second subtitle text, and normalize the second subtitle text to obtain the third subtitle text based on the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text. The subtitle display duration is determined based on the timecode of the second subtitle text. An audio generation unit is used to generate a third audio corresponding to each of the third subtitle texts; The video integration unit is used to synthesize all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed to obtain the final target video.
[0023] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the video dubbing method described above.
[0024] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video dubbing method as described above.
[0025] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the video dubbing method as described above.
[0026] In cases where the audio duration differs from the subtitle display duration during video dubbing, this invention standardizes the subtitle text, thereby avoiding damage to the original video file and ensuring high-quality video dubbing. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0028] Figure 1 This is a flowchart illustrating the video dubbing method provided by the present invention.
[0029] Figure 2 This is a schematic diagram of the process for generating third audio based on third subtitle text provided by the present invention.
[0030] Figure 3 This is a flowchart illustrating the speech rate softening algorithm provided by the present invention.
[0031] Figure 4 This is one of the flowcharts for determining the emotion label of a third audio file provided by the present invention.
[0032] Figure 5 This is the second flowchart illustrating the process of determining the emotion label of a third audio file provided by the present invention.
[0033] Figure 6 This is a schematic diagram of the process for determining the volume parameters of a third audio signal provided by the present invention.
[0034] Figure 7 This is a flowchart illustrating the process of determining the pitch parameters of a third audio signal provided by the present invention.
[0035] Figure 8 This is a flowchart illustrating the process of determining the timbre parameters of a third audio signal, as provided by the present invention.
[0036] Figure 9 This is a flowchart illustrating the dubbing and listening method provided by the present invention.
[0037] Figure 10 This is a schematic diagram of the video dubbing device provided by the present invention.
[0038] Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0040] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Those skilled in the art will understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0041] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0042] The following is combined with Figures 1-11 This invention describes the video dubbing method, apparatus, electronic device, and storage medium provided by the present invention.
[0043] Figure 1 This is a flowchart illustrating the video dubbing method provided by the present invention, as shown below. Figure 1 As shown, the execution subject of the video dubbing method provided by the present invention can be a video dubbing device corresponding to the video dubbing method, or an electronic device such as a server or computer terminal with video dubbing program installed. Unless otherwise specified, the following embodiments will be described using a video dubbing device as an example.
[0044] As an optional embodiment, the video dubbing method provided by the present invention includes, but is not limited to, the following steps: Step 110: After separating the first audio from the video to be dubbed, perform sentence-by-sentence text conversion on the first audio to obtain the first subtitle text with timecode.
[0045] The video to be dubbed is a multimedia file containing video and original audio, such as an original English, Italian, or Spanish film requiring cross-language dubbing. For ease of explanation, the following embodiments will use the dubbing of an original English film into Chinese as an example, which is not intended to limit the scope of protection of this invention.
[0046] Audio extraction techniques are used to extract the original audio contained in a video to be dubbed. For example, multimedia processing tools (such as FFmpeg) can be used to completely extract the original audio from the video to be dubbed, forming a separate original audio file (such as WAV format).
[0047] Considering that the original audio usually contains mixed human dialogue audio and background audio, in order to process the human dialogue audio independently (e.g., speech recognition, translation) and retain the background audio in the original audio file when synthesizing the final target video, this invention will separate the audio source of the original audio file.
[0048] Furthermore, an audio source separation model (such as the DemucsV4 model) can be used to process the original audio file to accurately separate it into two or more independent audio tracks. One audio track contains only the part with human voice dialogue, which is the first audio in this embodiment; the other audio track contains only the part with background music, environmental sound effects, and other non-human voice dialogue, which is the background audio in this embodiment.
[0049] It should be noted that the background audio will be saved independently and then merged with the newly generated dubbing audio and silent video during the final target video synthesis process, thereby ensuring that the auditory atmosphere of the target video is natural and consistent.
[0050] Automatic speech recognition (ASR) technology can be used, for example, by employing an automatic speech recognition engine to convert the first audio into first subtitle text with timecode.
[0051] An automatic speech recognition engine can be a deep learning model based on Transformer, mainly consisting of a speech recognition module and a speech alignment module. The speech recognition module can convert the input audio into text, while the speech alignment module can detect and identify the start and end times of each human dialogue in the audio and generate its corresponding timecode.
[0052] Specifically, the first audio can be input into an automatic speech recognition engine, which analyzes the input first audio to generate the text content corresponding to each sentence of human dialogue, i.e., the first subtitle text.
[0053] It should be emphasized that during the process of text recognition in the present invention, the voice alignment module of the automatic speech recognition engine will detect and identify the start time and end time of each human voice dialogue, and generate time codes corresponding to the first subtitle text of each human voice dialogue.
[0054] For example, the first subtitle text with time codes can be represented in the form of "[00:00:10.500 - 00:00:11.300], Hello". Among them, "[00:00:10.500 - 00:00:11.300]" is the time code, which is used to mark the start time and end time of the first subtitle text "Hello" on the video timeline.
[0055] Step 120: Translate each first subtitle text into a second subtitle text respectively, and the time code of each second subtitle text is the same as that of the corresponding first subtitle text.
[0056] As an optional embodiment, first select and train a suitable text translation model. By inputting the first subtitle text into the text translation model, the second subtitle text output by the text translation model can be quickly obtained.
[0057] Then, associate the second subtitle text generated by the text translation model with the first subtitle text through time codes to ensure that each second subtitle text is temporally aligned with its corresponding first subtitle text. Optionally, the text translation model is a deep learning model based on the Transformer architecture.
[0058] Assume that the time code associated with the first subtitle text "Hello" is [00:00:10.500 - 00:00:11.300]. After the text translation model translates it into the second subtitle text "你好", it will set the above time code [00:00:10.500 - 00:00:11.300] as the time code of the second subtitle text "你好".
[0059] Step 130: Obtain the second audio corresponding to each second subtitle text, and perform time window alignment on the second subtitle text according to the audio duration of the second audio to obtain the third subtitle text.
[0060] Considering that although the first subtitle text and the second subtitle text express the same semantics as different languages, there are usually differences in the text length and pronunciation duration between the two. If dubbing is directly based on the second subtitle text, problems such as out-of-sync audio and video may occur. Therefore, the present invention proposes to perform time window alignment on the second subtitle text according to the audio duration of the second audio, so that its content can adapt to the corresponding subtitle display duration requirements on the premise of ensuring the core semantics remain unchanged.
[0061] As an optional embodiment, the second subtitle text is time-window aligned according to the audio duration of the second audio to obtain the third subtitle text. This can be implemented in the following way: Based on the difference between the audio duration of the second audio and the display duration of the second subtitle text, the second subtitle text is normalized to obtain the third subtitle text.
[0062] The duration of the second subtitle text can be determined based on the timecode of the second subtitle text.
[0063] Specifically, the subtitle display duration is determined based on the timecode of the second subtitle text, and can be calculated by subtracting the start time from the end time of that timecode. For example, a second subtitle text associated with the timecode [00:00:10.500-00:00:13.500] will have a subtitle display duration of 3 seconds (13.500 seconds - 10.500 seconds).
[0064] To estimate the dubbing duration, an audio pre-synthesis operation can be performed on each second subtitle text to generate a corresponding second audio. The audio duration of the second audio can then be compared with the subtitle display duration to determine if audio-visual desynchronization has occurred. If audio-visual desynchronization is confirmed (the audio duration of the second audio is greater than or less than the subtitle display duration), the second subtitle text can be normalized to generate a third subtitle text of a more suitable length.
[0065] For example, a second subtitle text has a display duration of 3 seconds and its content is "Hello, I am very happy to meet all the guests here today." The pre-synthesized second audio has a duration of 4.5 seconds. Since the audio duration is longer than the subtitle display duration, audio-visual desynchronization will inevitably occur at this point during video dubbing. Based on the video dubbing method provided by this invention, the second subtitle text is normalized (without compromising its semantics), resulting in a third subtitle text with the content "Hello, it's a pleasure to meet you all."
[0066] Conversely, if a second subtitle text also has a display duration of 3 seconds and its content is "Thank you," the audio duration of the second audio obtained by pre-synthesizing the second subtitle text is only 0.8 seconds. Since the audio duration of 0.8 seconds is less than the subtitle display duration of 3 seconds, a normalization operation can be provided for the second subtitle text, and the resulting third subtitle text could contain the content "Oh, thank you so much."
[0067] Step 140: Generate the third audio corresponding to each third subtitle text.
[0068] By inputting each third subtitle text into the speech synthesis engine, you can obtain the third audio generated by the speech synthesis engine through the speech conversion of the third subtitle text.
[0069] As an optional embodiment, some audio synthesis parameters of the speech synthesis engine can be set to control the emotion, speech rate, volume or tone of the speech synthesis engine when converting the third subtitle text into speech, so that the output third audio has a specific expressiveness and thus better fits the scene atmosphere of the original video.
[0070] Step 150: Combine all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed to obtain the final target video.
[0071] Prepare all necessary media elements for compositing, including but not limited to the silent video of the video to be dubbed, the background audio of the video to be dubbed, and all third-party audio. The silent video of the video to be dubbed refers to the portion of the video footage remaining after all audio information (including dialogue and background audio) has been extracted from the original video to be dubbed. The background audio of the video to be dubbed refers to the audio obtained through audio source separation technology, containing only ambient sound effects and background music. All third-party audio refers to the collection of third-party audio generated sequentially from each third-party subtitle text along the timeline.
[0072] All third-level audio tracks are spliced together to form a complete dialogue audio track aligned with the total video duration. Alternatively, a new blank audio track can be created, and then, based on the timecode of the third-level subtitle text corresponding to each third-level audio track, the audio track is precisely positioned at its corresponding time point within that blank audio track.
[0073] The final compositing operation is performed using a multimedia compositing tool (such as FFmpeg). This operation uses the silent video of the video to be dubbed as the video track, and the target language dialogue audio track integrated in the previous step and the background audio of the video to be dubbed as two separate audio tracks. These tracks are then mixed and encapsulated to output a complete multimedia file containing the new dubbing, which is the final target video.
[0074] The video dubbing method provided by this invention avoids damage to the original video file and ensures high-quality video dubbing when there is a difference between the audio duration and the subtitle display duration during the video dubbing process.
[0075] In another embodiment of the present invention, the second subtitle text is normalized to obtain the third subtitle text based on the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text, including: obtaining the ratio coefficient between the audio duration and the subtitle display duration when it is determined that the speech rate normalization condition is met; performing a summary reduction operation on the second subtitle text when the ratio coefficient is determined to be greater than 1 to obtain the third subtitle text; and performing a word filling operation on the second subtitle text when the ratio coefficient is determined to be less than 1 to obtain the third subtitle text.
[0076] The ratio coefficient is an indicator used to quantify the degree of deviation between the audio duration and the subtitle display duration. It is determined by the ratio between the audio duration of the second audio corresponding to the second subtitle text and the subtitle display duration of the second subtitle text.
[0077] When it is determined that the second subtitle text needs to be normalized, the present invention decides on the specific normalization method based on the ratio coefficient.
[0078] If the ratio is determined to be greater than 1, it indicates that the duration of the second audio exceeds the duration of the second subtitle text. In this case, a summary reduction operation needs to be performed on the second subtitle text.
[0079] The abstract reduction operation refers to generating a third subtitle text while retaining the core semantics of the second subtitle text. The third subtitle text is generally shorter and more concise than the second subtitle text.
[0080] For example, a second subtitle text has a display duration of 2 seconds and its content is "Considering the weather forecast says it might rain heavily tomorrow, we'd better cancel our planned outdoor picnic." Its corresponding second audio clip has a duration of 4 seconds, resulting in a ratio coefficient of 2 (4 seconds / 2 seconds). In this case, a summary and sentence reduction operation needs to be performed on the second subtitle text, such as generating a third subtitle text with the content "It might rain tomorrow, let's cancel the picnic."
[0081] If the ratio coefficient is determined to be less than 1, it indicates that the audio duration of the second audio is too short relative to the subtitle display duration of the second subtitle text, which may cause unnatural pauses during dubbing. In this case, word filling operation needs to be performed on the second subtitle text.
[0082] Word filling refers to generating third subtitle text while preserving the core semantics of the second subtitle text. The third subtitle text generally has more characters and a more complex expression than the second subtitle text.
[0083] For example, if a second subtitle text has a display duration of 3 seconds and its content is "OK", and its corresponding second audio has a duration of 0.5 seconds, the calculated ratio coefficient is approximately 0.167 (0.5 seconds / 3 seconds). In this case, a word-filling operation needs to be performed on the second subtitle text, such as generating a third subtitle text with the content "Hmm, okay, I think this arrangement is very reasonable".
[0084] The video dubbing method provided by this invention introduces a ratio coefficient, which transforms the original conceptual judgment of duration difference into a quantifiable indicator that can guide subsequent operations. Based on the magnitude of this indicator, the corresponding text modification strategy is automatically selected, thereby solving the problem of dubbing duration mismatch more efficiently and accurately, and further improving the synchronization quality and naturalness of the final dubbed video.
[0085] In another embodiment of the present invention, a summary reduction operation is performed on the second subtitle text to obtain the third subtitle text. Specifically, this includes: determining the number of characters to be reduced when performing the summary reduction operation based on the ratio coefficient and the number of characters in the second subtitle text; generating a first prompt instruction based on the number of characters reduced; inputting the first prompt instruction and the second subtitle text into a text generation model to obtain the third subtitle text output by the text generation model.
[0086] As an optional embodiment, when the obtained ratio coefficient is greater than 1, a summary reduction operation is performed, which includes, but is not limited to, the following steps: Based on the ratio coefficient and the number of characters in the second subtitle text, the target number of characters to be reduced is determined by a preset calculation formula. This target number of characters to be reduced can be determined according to formula (1): (1) in, This indicates the goal is to reduce the word count. This indicates the original number of characters in the second subtitle text. This represents the ratio coefficient.
[0087] Based on this objective, a first prompt instruction is generated to reduce the number of words. This first prompt instruction is designed for the text generation model and contains explicit task constraints to guide the text generation model to perform a specific degree of text reduction.
[0088] The text generation model can identify the required text editing operation type (such as summary reduction, word filling, etc.) based on the input prompts, and optimize or adjust the content or structure of the input second subtitle text while preserving the original semantics and contextual coherence, so as to generate a third subtitle text that meets the target word count requirement.
[0089] Optionally, the text generation model can be a deep learning model based on the Transformer architecture, such as T5 (Text-to-Text Transfer Transformer), BART (Bidirectional and Auto-Regressive Transformer), etc.
[0090] If a second subtitle text has 50 characters and the calculated reduction is 15 characters, then the generated first prompt instruction can be: "Please reduce the following text by approximately 15 characters while preserving the core semantics."
[0091] Then, the first prompt instruction and the second subtitle text are input into a text generation model. The text generation model will perform a summary and sentence reduction operation on the second subtitle text, and finally output a more concise text, which is the third subtitle text.
[0092] The video dubbing method provided by this invention introduces the calculation of word reduction and the generation of a first prompt instruction with quantitative constraints, which makes the adjustment of text length more precise and avoids the problem of insufficient or excessive reduction caused by the free play of the text generation model. This improves the automation level of text regularization and the accuracy of the final dubbing duration.
[0093] In another embodiment of the present invention, a word-filling operation is performed on the second subtitle text to obtain the third subtitle text. Specifically, this includes: determining the number of words to be filled when performing the word-filling operation based on the ratio coefficient and the number of words in the second subtitle text; generating a second prompt instruction based on the number of words to be filled; inputting the second prompt instruction and the second subtitle text into the text generation model to obtain the third subtitle text output by the text generation model.
[0094] Based on the ratio coefficient and the number of characters in the second subtitle text, the target number of characters to be filled is determined by a preset calculation formula. This target number of characters to be filled can be determined according to formula (2): (2) in, Indicates the target number of characters to fill in. This indicates the original number of characters in the second subtitle text. This represents the ratio coefficient.
[0095] Based on the target word count, a second prompt instruction is generated. This second prompt instruction is designed for the text generation model and contains explicit task constraints to guide the text generation model to perform a specific degree of text expansion.
[0096] If a second subtitle text has 10 characters and the calculated number of fill characters is 15, then the generated second prompt instruction can be: "Please add approximately 15 characters to the following text to make it more complete and natural without changing the core semantics."
[0097] Then, the second prompt instruction and the second subtitle text are input into a text generation model. The text generation model performs controlled word-filling operations on the second subtitle text, such as adding modifiers, explanatory words, or expanding the meaning, and finally outputs a more comprehensive text, which is the third subtitle text.
[0098] The video dubbing method provided by this invention transforms the ambiguous text expansion task into a precisely controllable text generation task by introducing the calculation of the number of fill words and generating a second prompt instruction with quantified constraints. This approach allows for more precise adjustment of text length, avoiding problems such as insufficient fill or content redundancy caused by the text generation model's unchecked discretion, thereby further improving the automation level of text regularization and the natural fluency of the final dubbing.
[0099] In another embodiment provided by the present invention, the speech rate normalization condition includes: determining that the audio duration of the effective audio segment corresponding to the second audio is greater than the subtitle display duration of the second subtitle text; or, determining that the audio duration of the effective audio segment is less than the subtitle display duration of the second subtitle text, and the difference between the two is greater than a preset threshold.
[0100] A valid audio segment refers to the audio segment containing actual speech content that remains after the silent or near-silent parts at the beginning and end of the second audio have been removed.
[0101] For example, a sound amplitude threshold can be set, and audio segments with amplitudes consistently below the threshold can be treated as blank sounds and discarded to obtain the effective audio segment. The duration of the effective audio segment is more representative of the time required for the actual spoken content.
[0102] After obtaining the audio duration of the valid audio segment, determine whether the speech rate normalization condition is triggered by meeting any of the following conditions: The first scenario: The audio duration of the valid audio segment is longer than the display duration of the second subtitle text. This scenario means that without processing the second subtitle text, the dubbing will not be able to finish playing within the given time window.
[0103] The second scenario: The audio duration of the effective audio segment is less than the subtitle display duration of the second subtitle text, and the difference between the two is greater than a preset threshold. This preset threshold is used to distinguish between natural pauses and unnatural long silences; it can be a fixed duration value or a dynamic value related to the subtitle display duration.
[0104] For example, the preset threshold can be set to one-third of the subtitle display duration. Under this setting, if a subtitle display duration is 3 seconds, then its preset threshold is 1 second. If the effective audio segment duration is 2.5 seconds, and the difference between it and the subtitle display duration is 0.5 seconds, which is less than the preset threshold, then the pause is considered natural and no normalization operation is triggered. Conversely, if the effective audio segment duration is 1.8 seconds, and the difference between it and the subtitle display duration is 1.2 seconds, which is greater than the preset threshold, then the silence is considered too long and unnatural, and a normalization operation will be triggered.
[0105] The video dubbing method provided by this invention can accurately identify the time limit problem that must be dealt with by finely limiting the speech rate. At the same time, by introducing threshold judgment, it avoids unnecessary modification of text with time difference within a reasonable range, and preserves the natural pauses and rhythm in the dialogue, so that the final dubbing is more fluent and natural to listen to while ensuring synchronization.
[0106] Figure 2 This is a schematic diagram of the process for generating third audio based on third subtitle text provided by the present invention, as shown below. Figure 2 As shown, as another optional embodiment provided by the present invention, generating the third audio corresponding to each third subtitle text includes, but is not limited to, the following steps: Step 210: Obtain the current speech rate coefficient for audio conversion of each third subtitle text, as recorded in the speech synthesis engine.
[0107] The current speech rate coefficient is a global speech rate setting value, which can be a value preset by the user for the overall style of the entire dubbed video, or a default value provided by the speech synthesis engine.
[0108] For example, the current speech rate coefficient can be a speech rate parameter that represents the overall dubbing style. This speech rate parameter The overall style of the dubbed video is defined as either calm, standard, or lively. Optionally, the standard speaking speed is set to 1.0. A value greater than 1.0 indicates that the global speaking speed is faster than the standard speaking speed; if... A value less than 1.0 indicates that the global speaking speed is slower than the standard speaking speed.
[0109] Step 220: Adjust the current speech rate coefficient based on the speech rate normalization coefficient to obtain the target speech rate coefficient. The speech rate normalization coefficient is determined based on the ratio coefficient between the audio duration and the subtitle display duration.
[0110] The speech rate normalization coefficient is determined based on the ratio between the audio duration of the second audio and the subtitle display duration of the second subtitle text. The calculation method of the speech rate normalization coefficient is as shown in formula (3): (3) in, This indicates the duration of the second audio clip. This indicates the duration of the second subtitle text. This is the speech rate normalization coefficient.
[0111] The target speech rate coefficient is calculated by combining the current speech rate coefficient with the speech rate normalization coefficient. The calculation method of the target speech rate coefficient is shown in formula (4): (4) in, For the target speech rate coefficient, This represents the current speech rate coefficient. This is the speech rate normalization coefficient.
[0112] Assuming the current speech rate coefficient The duration of the second subtitle text is 1.1. Tsub The duration is 2 seconds, which corresponds to the audio length of the second audio track. The speech rate is 3 seconds. At this time, the calculated speech rate normalization coefficient is 1.5 (3 / 2). Then the target speech rate coefficient is calculated to be 1.65 (1.1×1.5).
[0113] Step 230: Adjust the current speech rate coefficient in the speech synthesis engine to the target speech rate coefficient, so as to use the speech synthesis engine to convert the third subtitle text into audio and generate the third audio.
[0114] As an optional embodiment, the third subtitle text, along with the target speech rate coefficient, is passed as input to the speech synthesis engine.
[0115] Upon receiving input, the speech synthesis engine uses the target speech rate coefficient to replace the current speech rate coefficient, thereby precisely controlling the speaking rate of the output audio. Specifically, the acoustic model and vocoder inside the speech synthesis engine can generate third audio based on the third subtitle text and the target speech rate coefficient.
[0116] For example, for the target speech rate coefficient The third subtitle text is set to 1.65. This third subtitle text, along with the target speech rate coefficient, is used as input to the speech synthesis engine. The speech synthesis engine generates the third audio at a rate 65% faster than the standard speech rate, so that content that would normally take 3 seconds to finish can be completed within a subtitle display duration of nearly 2 seconds.
[0117] The video dubbing method provided by this invention introduces a speech rate normalization coefficient calculated based on the duration ratio, on top of the global current speech rate coefficient, and dynamically calculates a target speech rate coefficient for each subtitle. Then, using this target speech rate coefficient for audio conversion, the duration of the final audio can be precisely controlled by directly adjusting the speech synthesis rate without changing the third subtitle text.
[0118] Figure 3 This is a flowchart illustrating the speech rate softening algorithm provided by the present invention, as shown below. Figure 3 As shown, considering that in actual dialogue, if the change in speaking speed between adjacent sentences is too drastic, it will produce an abrupt and unnatural listening experience, this invention provides a speaking speed smoothing algorithm.
[0119] As another optional embodiment provided by the present invention, for any target third subtitle text, after adjusting the current speech rate coefficient based on the speech rate normalization coefficient to obtain the target speech rate coefficient, the following steps are included but not limited to: Step 310: Obtain the first speech rate of the third audio corresponding to the previous third subtitle text adjacent to the target third subtitle text in the timecode, wherein the target third subtitle text is any one of all third subtitle texts.
[0120] The target third subtitle text refers to the third subtitle text that is currently being prepared for audio conversion. The first speech rate can be obtained through methods including but not limited to the following steps: Get the third subtitle text preceding the target third subtitle text, and get its corresponding third audio.
[0121] The third audio is processed to remove the blank sounds at the beginning and end to obtain the effective speech duration.
[0122] The ratio of the number of characters in the preceding third subtitle text to the effective speech duration is determined as the first speech rate.
[0123] If the content of the previous third subtitle text is "It might rain tomorrow, let's cancel the picnic" (11 characters), and the effective speech duration of its corresponding third audio after removing blank sounds is 2.2 seconds, then the calculated first speech rate is 5 characters / second (11 / 2.2).
[0124] Step 320: Based on the timecode of the target third subtitle text, determine the dialogue audio segment within the duration range of the timecode of the first audio to obtain the second speech rate of the dialogue audio segment.
[0125] The methods for obtaining the second speaking speed include, but are not limited to, the following steps: Obtain the timecode associated with the target third subtitle text. This timecode corresponds to a duration interval including the start and end times. Then, extract the dialogue audio segments within this duration interval from the first audio.
[0126] The second speech rate of the dialogue audio segment is obtained. The second speech rate can be the ratio of the actual duration of human voice speech to the duration interval. The acquisition process can be achieved using Voice Activity Detection (VAD) technology. This VAD technology can analyze the dialogue audio segment and accurately calculate the actual duration of human voice speech within the duration interval.
[0127] The ratio of the actual duration of the human voice to the duration interval is defined as the second speech rate.
[0128] For example, if the duration of a target third subtitle text is 3 seconds, and after analyzing its corresponding first audio segment using speech activity detection technology, the total actual duration of human voice speech is determined to be 1.5 seconds. In this case, the calculated second speech rate is 0.5 (1.5 / 3).
[0129] Step 330: Based on the first speech rate and the second speech rate, determine the soft speech rate range. When the target speech rate for audio conversion of the target third subtitle text based on the target speech rate coefficient is outside the soft speech rate range, generate a prompt message to guide the user to perform manual optimization.
[0130] The gentle speaking speed range refers to a range of speaking speeds that sound natural and are in harmony with the rhythm of the context.
[0131] Optionally, the average of the first and second speaking speeds can be used as a reference speaking speed, and a range of fluctuation can be set based on this reference speaking speed. This range is the soft speaking speed range. For example, if the reference speaking speed is 5 words / second and the range of fluctuation is 5%, then the soft speaking speed range is [4.75 words / second, 5.25 words / second].
[0132] Target speech rate refers to the actual speaking rate when the target speech rate coefficient is applied to the audio conversion of the target third subtitle text. It can be obtained by multiplying the target speech rate coefficient by the standard speech rate.
[0133] Compare the calculated target speech rate with the soft speech rate range, and perform the corresponding operation based on the comparison result: If the target speech rate falls within the range of the gentle speech rate, then the target speech rate is considered appropriate, and the current speech rate adjustment scheme does not require intervention and can be directly used for audio synthesis.
[0134] If the target speaking speed is outside the gentle speaking speed range, it indicates that the speaking speed of the current sentence may be too fast or too slow compared to the rhythm of the context, which will produce an abrupt listening experience. At this time, a prompt message is generated to guide the user to manually optimize, such as: "The estimated speaking speed of the current sentence differs greatly from the rhythm of the context, which may lead to a disjointed listening experience. It is recommended to manually optimize the text or speaking speed parameters."
[0135] The video dubbing method provided by this invention dynamically constructs a smooth speech rate range by referencing the first speech rate of the previously synthesized third audio and the second speech rate of the first audio within the current time interval. Then, this smooth speech rate range is used to evaluate the reasonableness of the target speech rate to be applied. This effectively avoids the abrupt listening experience caused by excessively fast or slow changes in speech rate between adjacent sentences. Therefore, speech rate adjustment is no longer merely about meeting the duration constraints of a single sentence, but rather about optimizing from a higher dimension, ensuring the overall fluency and naturalness of the dialogue.
[0136] In another embodiment of the present invention, generating a third audio corresponding to any target third subtitle text includes: determining an emotion tag for generating the third audio corresponding to the target third subtitle text; inputting the emotion tag and the target third subtitle text into a speech synthesis engine to generate the third audio.
[0137] An emotion label is an identifier used to describe a speaker's emotional state, such as "happy," "sad," "angry," or "calm." The emotion label is derived by analyzing the third audio corresponding to the target third subtitle text. For example, the third audio corresponding to the target third subtitle text is input into a trained speech emotion recognition model to obtain the emotion label.
[0138] The target third subtitle text and the emotion tag are fed together as input to the speech synthesis engine. The speech synthesis engine will call its internal emotional speech library and acoustic model according to the received emotion tag, so that the final generated third audio also matches the emotion tag in terms of emotion expression.
[0139] If the identified emotion label is "happy" and the text content is "I passed the exam", the speech synthesis engine will generate a third audio clip with obvious joy and an upward intonation.
[0140] The video dubbing method provided by this invention introduces the identification and synthesis of emotion tags, so that the dubbing process is no longer a simple, emotionless reading of text, but a performance with emotional expression. This transfers the original speaker's emotions to the dubbing of the target language, greatly enhancing the dubbing's appeal, vividness, and immersiveness.
[0141] Figure 4This is one of the flowcharts for determining the emotion label of a third audio file provided by the present invention, such as... Figure 4 As shown, as another optional embodiment provided by the present invention, determining the emotion tag of the third audio corresponding to the target third subtitle text includes, but is not limited to, the following steps: Step 410: Determine the target dialogue audio segment in the first audio, where the timecode of the target dialogue audio segment is the same as the timecode of the target third subtitle text.
[0142] First, obtain the timecode associated with the target third subtitle text. This timecode corresponds to a duration interval including the start and end times. Then, extract the target dialogue audio segment within this duration interval from the first audio.
[0143] For example, if the timecode of the target third subtitle text is [00:01:05.200-00:01:07.800], then the audio segment from 1 minute 5 seconds 200 milliseconds to 1 minute 7 seconds 800 milliseconds will be extracted from the complete first audio and used as the target dialogue audio segment.
[0144] Step 420: After determining the target video segment in the video to be dubbed, perform video frame extraction on the target video segment to obtain multiple video screenshots; the timecode of the target video segment is the same as the timecode of the target third subtitle text.
[0145] Video frame extraction refers to extracting a series of static image frames from a target video segment at a preset frequency or according to a set of rules, in order to obtain a set of video screenshots containing multiple frames. For example, it can be set to extract one frame every 0.2 seconds.
[0146] To improve the efficiency of subsequent processing and avoid analyzing redundant information, the acquired multi-frame video screenshots can be deduplicated. For example, by comparing the similarity of two adjacent video screenshots, if the similarity is higher than a preset similarity threshold, it is considered that the content of these two frames is basically unchanged, and one of the frames can be discarded. After this deduplication process, a more representative set of key video screenshots that can reflect changes in the content of the images can be obtained.
[0147] For example, if the timecode of a target third-party subtitle text is [00:01:05.200-00:01:07.800], a 2.6-second segment of the target video to be dubbed will be extracted first. Next, this target video segment will undergo frame extraction, potentially yielding dozens of video screenshots. After further deduplication, only a few key video screenshots that clearly demonstrate changes in the speaker's facial expressions or scene transitions within that time period will be retained.
[0148] Step 430: Determine the emotion tag of the third audio based on at least one of the target dialogue audio segment, multi-frame video screenshots, and third caption text.
[0149] As an optional embodiment, the emotion tag determination process can be adaptively performed based on the available information sources: if only the target dialogue audio segment is available, the emotion tag is determined solely based on the target dialogue audio segment; if only multiple video screenshots and third-party subtitle text are available, the emotion tag is determined based on the multiple video screenshots and the third-party subtitle text.
[0150] When the emotional tendencies analyzed from three information sources (audio, video, and text) are consistent, a final emotional label is determined by combining these findings. For example, if the audio tone is identified as happy, the facial expression of the person in the video is smiling, and the text content is positive, then the emotional label is determined to be "happy".
[0151] When the analysis results from different information sources conflict, different judgment weights can be assigned to different information sources. For example, when judging sarcasm, the text content may be positive (such as "That's great"), but the tone of the audio and the character's facial expression may be negative. In this case, higher weights can be given to audio and visual information to determine the true complex emotional label, such as "sarcasm" or "dissatisfaction".
[0152] The video dubbing method provided by this invention constructs a multimodal information analysis framework, which comprehensively analyzes the original dialogue audio segments, video screenshots, and text content corresponding to the current subtitle timecode. This effectively solves the problems of ambiguity, conflict, or missing information that may exist when relying on a single information source for emotion judgment, thereby further ensuring the authenticity and accuracy of the emotion.
[0153] Figure 5 This is the second flowchart illustrating the process of determining the emotion label of a third audio file provided by the present invention, as follows: Figure 5 As shown, as another optional embodiment provided by the present invention, the emotion tag of the third audio is determined based on at least one of the target dialogue audio segment, multi-frame video screenshots, and third subtitle text, including but not limited to the following steps: Step 510: Input the target dialogue audio segment into the speech emotion recognition model and obtain the first emotion recognition result output by the speech emotion recognition model.
[0154] A speech emotion recognition model is a deep learning model trained on a certain amount of emotion-labeled speech data. Its function is to extract emotional information from audio signals. This speech emotion recognition model can be based on a Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), or Transformer architecture.
[0155] After receiving the target dialogue audio segment, the speech emotion recognition model first preprocesses it and extracts acoustic features, such as Mel-Frequency Cepstral Coefficients (MFCCs) or spectrograms.
[0156] Then, the extracted acoustic features are input into the speech emotion recognition model for classification, and finally the first emotion recognition result is output. The first emotion recognition result can be a single label representing the most likely emotion (e.g., "sadness"), or a list containing multiple emotions and their corresponding confidence scores (e.g., {sadness: 0.75, calm: 0.2, anger: 0.05}).
[0157] Step 520: Input the video screenshot into the image emotion recognition model to obtain the second emotion recognition result output by the image emotion recognition model.
[0158] Image emotion recognition models are deep learning models trained on a certain amount of image data with emotion labels. Their function is to extract emotional states from visual information.
[0159] The image emotion recognition model first performs face detection on the input multi-frame video screenshots to locate the facial regions of the people in the scene. Then, the facial regions are input into a facial expression recognition sub-network to analyze their expression features and determine the person's emotion, such as "happy", "sad" or "surprised".
[0160] Image emotion recognition models can also analyze the global image features of the entire video screenshot to determine the overall atmosphere of the scene. For example, by analyzing the color tone, lighting, and object composition of the image, the atmosphere of the scene can be identified as "suspenseful," "romantic," or "oppressive."
[0161] After receiving multiple video screenshots, the image emotion recognition model integrates the analysis from one or two of the aforementioned dimensions to output a second emotion recognition result. This second emotion recognition result can be a single label representing the most likely emotion, or a list containing multiple emotions and their corresponding confidence scores.
[0162] For example, for a video screenshot containing a person smiling with a bright background, the image emotion recognition model might output a second emotion recognition result {happy: 0.85, calm: 0.1, other: 0.05}.
[0163] Step 530: Input the target third subtitle text into the text emotion recognition model and obtain the third emotion recognition result output by the text emotion recognition model.
[0164] A text sentiment recognition model is a deep learning model trained on a certain amount of sentiment-labeled text data. Its function is to infer the emotions contained in the text content by analyzing its semantics. This text sentiment recognition model can be a large language model (LLM) with powerful natural language understanding capabilities.
[0165] After receiving the target third-party caption text, the text sentiment recognition model performs in-depth semantic and sentiment analysis on it, ultimately outputting a third-party sentiment recognition result. This result can be a single label representing the most likely emotion (e.g., "positive" or "negative"), or a list containing multiple emotions and their corresponding confidence scores. This third-party sentiment recognition result provides a basis for judgment from the semantic dimension of the text for subsequent multimodal fusion decisions.
[0166] For example, if the target third caption text is "Oh, that's great," the text emotion recognition model may determine the emotion as "happy" or "sarcastic" based on its internal semantic understanding.
[0167] Step 540: Determine an emotion label based on at least one of the first emotion recognition result, the second emotion recognition result, and the third emotion recognition result.
[0168] As an optional embodiment, the process of determining emotion labels can employ a fusion decision strategy, which can be preset as a set of rules or a small decision model to process and judge the recognition results from different modalities.
[0169] When only one recognition result is available, that result is directly adopted as the final emotion label. For example, if only the first emotion recognition result is available, it is used as the final emotion label.
[0170] When multiple recognition results are available, the following principles can be used to make a judgment: If the first emotion recognition result (from audio), the second emotion recognition result (from video), and the third emotion recognition result (from text) all point to the same emotion, such as "happy", then the final emotion label is determined to be "happy".
[0171] If the recognition results of different modalities are inconsistent, a preset weight can be assigned to the recognition results of each modality. For example, when judging whether there is sarcasm, the text content may be positive, but the tone of the audio and the facial expressions in the video may be negative. In this case, the recognition results of the audio and visual modalities can be given higher weights, so as to comprehensively determine the final emotion label as "sarcasm" or "dissatisfaction".
[0172] The video dubbing method provided by this invention configures dedicated speech emotion recognition models, image emotion recognition models, and text emotion recognition models for three different information modalities: audio, video, and text. It then performs a final fusion decision based on the independent emotion recognition results output by each model. This method concretizes the multimodal analysis framework, thereby further improving the reliability and consistency of the entire emotion transfer process while ensuring accurate emotion judgment.
[0173] In another embodiment of the present invention, generating a third audio corresponding to any target third subtitle text further includes: determining audio feature parameters for generating the third audio corresponding to the target third subtitle text, the audio feature parameters including volume parameters and / or pitch parameters; inputting the audio feature parameters and the target third subtitle text into a speech synthesis engine to generate the third audio; the target third subtitle text is any one of all third subtitle texts.
[0174] Considering that human natural speech is not monotonous, but contains rich volume and pitch variations, which are important carriers of tone and emotion, this invention inputs audio feature parameters and target third subtitle text into a speech synthesis engine to generate third audio.
[0175] By analyzing the first audio segment corresponding to the timecode of the target third subtitle text, the audio feature parameters required to generate the third audio are determined, including volume parameters and / or pitch parameters.
[0176] The audio feature parameters, along with the target third subtitle text, are passed as input to the speech synthesis engine. When generating the audio waveform, the speech synthesis engine uses the received audio feature parameters to adjust its output in real time.
[0177] For example, if the volume parameter in the audio feature parameters indicates that the third subtitle text should be louder than the average volume (such as an exclamation), and the pitch parameter in the audio feature parameters indicates that the pitch should rise (such as a question), then the third audio generated by the speech synthesis engine will have louder and higher pitch acoustic characteristics, thereby accurately reproducing the intonation changes of the original speaker.
[0178] The video dubbing method provided by this invention, by introducing the analysis and application of audio feature parameters, elevates the dubbing process from simple content copying to precise replication at the level of sound physical characteristics. This method successfully transfers the original speaker's intonation and rhythm to the target language dubbing, making the dubbing more vivid, natural, and expressive.
[0179] Figure 6 This is a flowchart illustrating the process of determining the volume parameters of a third audio signal provided by the present invention, as shown below. Figure 6 As shown, as another optional embodiment provided by the present invention, determining the volume parameters of the third audio corresponding to the target third subtitle text includes, but is not limited to, the following steps: Step 610: Based on the waveform of the first audio, determine the first average volume of the first audio.
[0180] The waveform of the first audio signal is a time-domain representation of the first audio signal. For example, Pulse-Code Modulation (PCM) data can intuitively show the relationship between the amplitude of the sound and time.
[0181] The first volume average can be determined by the root mean square (RMS) value of the global waveform of the first audio, and the calculation method of the first volume average can be shown in formula (5): (5) in, The average volume of the first volume. This indicates the total number of sampling points in the first audio waveform. Indicates the i Amplitude values at each sampling time point.
[0182] Step 620: Based on the timecode of the target third subtitle text, determine the dialogue audio segment within the duration interval of the timecode of the first audio, and determine the second average volume of the dialogue audio segment based on the waveform of the dialogue audio segment.
[0183] Obtain the timecode associated with the target third subtitle text. This timecode corresponds to a duration interval including the start and end times. Then, extract the dialogue audio segments within this duration interval from the first audio.
[0184] The waveform of the dialogue audio segment itself is analyzed to determine its second volume average. Optionally, the method for determining the second volume average is the same as that for determining the first volume average, i.e., by calculating the root mean square value of the waveform of the dialogue audio segment.
[0185] Step 630: Determine the volume parameters of the third audio corresponding to the target third subtitle text based on the volume weighting coefficient between the second volume average and the first volume average, and the preset volume level in the speech synthesis engine.
[0186] The volume weighting coefficient is the ratio of the average second volume to the average first volume, and its calculation method is shown in formula (6): (6) in, This is the volume weighting factor. This is the second average volume of the audio clip from the dialogue. It is the average volume of the first audio.
[0187] The final volume parameter is determined based on the volume weighting coefficient and the preset volume level in the speech synthesis engine. The preset volume level in the speech synthesis engine is a default volume value for the engine (e.g., within a range of 0 to 100, its default value is 50). This volume parameter can be determined using formula (7): (7) in, For volume parameters, This is the default volume value in the speech synthesis engine. This is the volume weighting factor.
[0188] For example, if the calculated first volume average RMS1 The second average volume of the dialogue audio segment is 2000. RMS2 If it is 3000, then the volume weight coefficient is... V The value is 1.5. This is the preset volume level for the speech synthesis engine. If the value is 50, then the volume parameter of the third audio corresponding to the target third subtitle text is determined. It is 75 (50 × 1.5).
[0189] The video dubbing method provided by this invention establishes a global first volume average as a benchmark, calculates the volume weight coefficient between the second volume average of each dialogue sentence and this benchmark, and then applies this coefficient to the preset volume of the speech synthesis engine. This method enables dubbing to fully preserve the dynamic range and natural intonation of the original speech, thereby further enhancing the vividness and realism of the dubbing while achieving volume feature transfer.
[0190] Figure 7 This is a flowchart illustrating the process of determining the pitch parameters of a third audio signal provided by the present invention, as shown below. Figure 7As shown, as another optional embodiment provided by the present invention, determining the pitch parameters of the third audio corresponding to the target third subtitle text includes, but is not limited to, the following steps: Step 710: Obtain the spectrogram of the first audio signal, which is obtained by performing a Fourier transform on the global waveform of the first audio signal.
[0191] Considering that the pitch in speech is mainly determined by the fundamental frequency of the sound, and the analysis of the fundamental frequency needs to be carried out in the frequency domain, it is necessary to convert the first audio from the time domain to the frequency domain.
[0192] The corresponding spectrogram is obtained by performing a Fourier transform on the global waveform of the first audio signal (i.e., the time-domain signal representation of the audio). This spectrogram is a graphical representation that visually displays the frequency components of the first audio signal, with the horizontal axis representing frequency and the vertical axis representing energy or amplitude at that frequency. To improve computational efficiency, this Fourier transform can be a Fast Fourier Transform (FFT).
[0193] Step 720: Determine the average first pitch of the first audio based on the spectrogram.
[0194] Considering that the pitch of a sound acoustically mainly corresponds to its fundamental frequency, the process of determining the average of the first pitch is equivalent to calculating the average fundamental frequency of the first audio audio. The process of calculating the average fundamental frequency of the first audio audio involves dividing the first audio audio into multiple audio frames, calculating the fundamental frequency value of each audio frame, and finally calculating the arithmetic mean of all valid fundamental frequency values to obtain the average of the first pitch of the first audio audio.
[0195] Specifically, the process of determining the fundamental frequency value of any audio frame includes, but is not limited to, the following steps: Set a frequency search range that conforms to the fundamental frequency characteristics of human speech, such as 85 Hz to 255 Hz, which can roughly cover the normal speech fundamental frequency of adult men and women.
[0196] Within the preset frequency search range of the local spectrogram corresponding to the audio frame, find the peak point with the highest energy.
[0197] Read the value corresponding to the peak point on the frequency axis; this value is the fundamental frequency value of the audio frame.
[0198] By traversing and analyzing all audio frames of the first audio file, a series of fundamental frequency values can be obtained. The arithmetic mean of all valid fundamental frequency values is then calculated, and the final result is the average of the first pitch of the first audio file.
[0199] Step 730: Based on the timecode of the target third subtitle text, determine the dialogue audio segment within the duration interval of the timecode of the first audio, and determine the second pitch average of the dialogue audio segment based on the local spectrogram of the dialogue audio segment.
[0200] Obtain the timecode associated with the target third subtitle text. This timecode corresponds to a duration interval including the start and end times. Then, extract the dialogue audio segments within this duration interval from the first audio.
[0201] Perform a Fourier transform on the dialogue audio segment to obtain its corresponding local spectrogram. Based on this local spectrogram, determine the second pitch mean of the dialogue audio segment.
[0202] Optionally, the process of determining the second pitch average of the dialogue audio segment is similar to the process of determining the first pitch average, that is, the dialogue audio segment is divided into multiple audio frames, the fundamental frequency value is identified from the spectrogram of each audio frame, and finally the arithmetic average of all valid fundamental frequency values is calculated, and the result is the second pitch average.
[0203] It should be noted that this second pitch average accurately reflects the actual pitch level used by the original speaker when uttering the specific sentence, providing crucial local pitch data for subsequent calculation of relative pitch weight.
[0204] Step 740: Determine the pitch quantity parameter of the third audio corresponding to the target third subtitle text based on the pitch weight coefficient between the second pitch average and the first pitch average, and the preset pitch level in the speech synthesis engine.
[0205] The pitch weighting coefficient is the ratio of the average of the second pitch to the average of the first pitch, and its calculation method is shown in formula (8): (8) in, This is the pitch weight coefficient. It is the average of the second tone in the audio segment of the dialogue. It is the average of the first pitch of the first audio.
[0206] The final pitch parameter is determined based on the pitch weight coefficient and the preset pitch level in the speech synthesis engine. The preset pitch level in the speech synthesis engine is a default pitch value (e.g., 50 within a relative pitch range of 0 to 100). This pitch parameter can be determined using formula (9): (9) in, For pitch parameters, This indicates the default pitch value preset in the speech synthesis engine. This is the pitch weighting coefficient.
[0207] For example, if the calculated average of the first pitch freq1 The average second tone of the dialogue audio clip is 150 Hz. freq2 If it is 180 Hz, then the pitch ratio is... The value is 1.2 (180 / 150). This is the default pitch value preset in the speech synthesis engine. If the value is 50, then this is the final pitch parameter determined for that sentence. The pitch parameter is set to 60 (50 × 1.2), and this pitch parameter will be passed to the speech synthesis engine to generate the third audio corresponding to the target third subtitle text.
[0208] The video dubbing method provided by this invention establishes a global first pitch average as a benchmark, calculates the pitch weight coefficient between the second pitch average of each dialogue sentence and this benchmark, and then applies this coefficient to the preset pitch of the speech synthesis engine. This method enables dubbing to completely preserve the intonation contour and natural pitch fluctuations of the original speech, thereby further improving the intonation realism and emotional accuracy of the dubbing while achieving pitch feature transfer.
[0209] Figure 8 This is a flowchart illustrating the process of determining the timbre parameters of a third audio signal provided by the present invention, as shown below. Figure 8 As shown, as another optional embodiment provided by the present invention, generating a third audio corresponding to any target third subtitle text further includes, but is not limited to, the following steps: Step 810: Determine the target dialogue audio segment in the first audio. The timecode of the target dialogue audio segment is the same as the timecode of the target third subtitle text. The target third subtitle text is any one of all third subtitle texts.
[0210] Considering that there may be multiple different speakers in a video, each with a unique timbre, this invention provides a matching mechanism based on timbre recognition in order to reflect this difference in roles during dubbing.
[0211] Parse the associated timecode from the target third subtitle text to be processed, and use the timecode to locate and extract the audio segment within the corresponding time period from the complete first audio. This audio segment is the target dialogue audio segment.
[0212] For example, if the timecode of a target third subtitle text is [00:02:15.100-00:02:18.300], then the audio segment from 2 minutes 15 seconds 100 milliseconds to 2 minutes 18 seconds 300 milliseconds will be extracted from the first audio and used as the audio segment of the target dialogue.
[0213] Step 820: Input the target dialogue audio segment into the voice recognition model and obtain the gender and age information output by the voice recognition model.
[0214] The timbre recognition model is a deep learning model specifically designed to analyze and identify the speaker's inherent acoustic features from speech signals.
[0215] This timbre recognition model can be trained on a certain amount of speech data labeled with the speaker's gender and age information, enabling it to learn the differences in voice characteristics among people of different genders and age groups, such as fundamental frequency range and formant structure.
[0216] The target dialogue audio segment is provided as input data to the voice recognition model, which analyzes the input data and outputs the original speaker's gender and age information. This gender and age information can be a set of descriptive labels, such as {gender: male, age group: middle-aged} or {gender: female, age group: youth}.
[0217] Step 830: Match the target speaker information from the intelligent agent library based on gender and age information. The intelligent agent library stores different speaker information corresponding to different gender and age information.
[0218] The intelligent agent library is a pre-set resource library storing a variety of synthesized speech timbres to choose from. In this library, each available timbre is encapsulated as an independent speaker, and each speaker's information is associated with specific gender and age information, thus establishing a clear mapping relationship. This speaker information may include a unique speaker identifier, a description of timbretic features (e.g., "sweet female voice" or "deep male voice"), and the relevant parameters required to invoke that speaker for speech synthesis.
[0219] Using the obtained gender and age information as query conditions, the system searches and matches within the intelligent agent database to find the speaker information that best matches the gender and age information. This speaker information is the target speaker information.
[0220] As an optional embodiment, when multiple matching speaker information is found, a unique target speaker information can be determined according to preset rules (e.g., selecting the default or the most frequently used one).
[0221] For example, if the obtained gender and age information is {gender: male, age group: middle-aged}, the system will search for speakers that match the label in the intelligent agent database. Ultimately, it may match a speaker with a deep and steady baritone voice and obtain the corresponding target speaker information.
[0222] Step 840: Input the target speaker information and the third caption text into the speech synthesis engine to generate the third audio.
[0223] The target speaker information, along with the third subtitle text, is passed as input to the speech synthesis engine. The speech synthesis engine then retrieves the corresponding timbre model and acoustic parameters from its internal timbre library based on the received target speaker information.
[0224] The speech synthesis engine then uses the selected timbre to perform audio conversion on the third subtitle text, ultimately generating a third audio clip with the specific character's timbre.
[0225] For example, if the target speaker's voice is a deep, steady baritone, and the third subtitle text is "We must find the truth," then the speech synthesis engine will use the baritone voice to generate the third audio for this sentence, thus making the dubbing fit the identity of a middle-aged male character.
[0226] The video dubbing method provided by this invention introduces a mechanism based on timbre recognition and speaker matching. This mechanism automatically analyzes the gender and age characteristics of the original speaker and selects a matching timbre from a pre-set intelligent voice library for dubbing. This method can assign unique dubbing timbres that match the identity characteristics of different speakers in multi-role videos, thereby further improving the character differentiation and expressiveness of the dubbing while ensuring audio-visual synchronization.
[0227] Figure 9 This is a flowchart illustrating the dubbing and listening method provided by the present invention, as follows: Figure 9 As shown, as another optional embodiment provided by the present invention, the video dubbing method further includes providing a dubbing preview, specifically including but not limited to the following steps: Step 910: Divide all third-party subtitle texts into multiple subtitle batches.
[0228] When responding to a user's request to listen to the audio, all third-party subtitle texts arranged in timecode order will be divided into multiple consecutive data units according to a preset number, and each data unit is a subtitle batch.
[0229] For example, if the preset quantity is 5, then among all the third subtitle texts arranged in timecode order, the 1st to 5th texts are divided into the first subtitle batch, the 6th to 10th texts are divided into the second subtitle batch, and so on, until all the third subtitle texts have been divided.
[0230] Step 920: In response to the user's listening request, the current subtitle batch is processed by speech synthesis to generate the current dialogue audio batch, and the current dialogue audio batch is sent to the client for playback.
[0231] Speech synthesis is performed on the current subtitle batch (e.g., the first subtitle batch). This speech synthesis process involves executing the audio generation process for each third subtitle text in the batch.
[0232] As an optional embodiment, the audio generation process may include determining emotion tags, audio feature parameters (such as volume and pitch), and target speaker information, and inputting these parameters along with third-party caption text into a speech synthesis engine to generate expressive, high-quality dialogue audio. After generating audio for all third-party caption texts in this batch, these generated audio segments together constitute the current dialogue audio batch.
[0233] The current audio batch (e.g., a set of audio files and their corresponding timecode information) is sent to the client over the network. Upon receiving the audio batch, the client's built-in player immediately begins playing the audio in the batch in timecode order.
[0234] Step 930: While the client is playing the current dialogue audio batch, speech synthesis is performed on the adjacent next subtitle batch to generate the next dialogue audio batch.
[0235] As an alternative embodiment, while the client begins playing the current batch of dialogue audio (e.g., the first batch), the next adjacent batch of subtitles (e.g., the second batch) is automatically and in parallel processed by the backend or a processor.
[0236] The process of speech synthesis for the next batch of subtitles is the same as that for the current batch of subtitles, that is, the audio generation process is executed one by one for each third subtitle text in the next batch of subtitles to generate the corresponding third audio.
[0237] Ultimately, the collection of all third audio tracks in the next subtitle batch constitutes the next dialogue audio batch. Once generated, this next dialogue audio batch is temporarily stored on the server or in the processing program, waiting to be called.
[0238] Step 940: After confirming that the current dialogue audio has finished playing, send the next batch of dialogue audio to the client for playback in the order of the timecodes.
[0239] As an optional implementation, once the client's player finishes playing the last audio segment in the current audio batch of the conversation, it can be determined that the batch has finished playing. At this time, the client can send a signal to the backend or processing program requesting the next batch of data.
[0240] Upon receiving this signal, the backend or processing program sends the pre-synthesized and prepared next batch of dialogue audio to the client. Upon receiving this batch, the client immediately adds it to the playback queue and, based on the timecode, seamlessly begins playing the next batch of dialogue audio, immediately following the previous batch.
[0241] Then, steps 930 to 940 are executed repeatedly. Specifically, the "next dialogue audio batch" that has just been sent to the client and started playing transforms into the new "current dialogue audio batch" in the new loop. At the same time, the backend or processing program immediately begins speech synthesis for the next batch of subtitles, preparing it to become the new "next dialogue audio batch".
[0242] This loop will continue until one of the following termination conditions is met: Once the last batch of dialogue audio corresponding to the last batch of subtitles has finished playing, the loop will automatically terminate because there are no more next batches of subtitles.
[0243] If, at any point during the loop, the user performs a stop operation on the client interface, the client will halt playback and notify the backend or processing program. At this point, the loop will immediately terminate, and any ongoing speech synthesis tasks will also be stopped to conserve computing resources.
[0244] The video dubbing method provided by this invention divides all third-party subtitle text into multiple batches and adopts a streaming listening mechanism that combines the first batch of synthesized playback with the subsequent batches being preloaded in parallel and seamlessly switched. This provides users with an instant and efficient interactive review capability while ensuring the continuity and smoothness of the listening process, thereby greatly improving the user's work efficiency in the dubbing adjustment and iteration process.
[0245] Figure 10 This is a structural schematic diagram of the video dubbing device provided by the present invention, as shown below. Figure 10 As shown, it mainly includes, but is not limited to: The video parsing unit 1010 is used to separate the first audio from the video to be dubbed, and then perform sentence-by-sentence text conversion on the first audio to obtain the first subtitle text with timecode.
[0246] The subtitle translation unit 1020 is used to translate each of the first subtitle texts into a second subtitle text, wherein the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text.
[0247] The text normalization unit 1030 is used to obtain the second audio corresponding to each second subtitle text, and normalize the second subtitle text to obtain the third subtitle text based on the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text. The subtitle display duration is determined based on the timecode of the second subtitle text.
[0248] The audio generation unit 1040 is used to generate a third audio corresponding to each of the third subtitle texts.
[0249] The video integration unit 1050 is used to synthesize all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed to obtain the final target video.
[0250] It should be noted that the video dubbing device provided by the present invention can execute the video dubbing method described in any of the above embodiments during specific operation, and this embodiment will not elaborate on this.
[0251] The video dubbing device provided by this invention avoids damage to the original video file and ensures high-quality video dubbing when there is a difference between the audio duration and the subtitle display duration during the video dubbing process.
[0252] Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 11As shown, the electronic device may include: a processor 1110, a communications interface 1120, a memory 1130, and a communication bus 1140, wherein the processor 1110, the communications interface 1120, and the memory 1130 communicate with each other through the communication bus 1140. The processor 1110 can call logical instructions in the memory 1130 to execute a video dubbing method, which includes: separating a first audio from the video to be dubbed; performing sentence-by-sentence text conversion on the first audio to obtain a first subtitle text with a timecode; translating each first subtitle text into a second subtitle text, wherein the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text; obtaining a second audio corresponding to each second subtitle text; normalizing the second subtitle text according to the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text to obtain a third subtitle text, wherein the subtitle display duration is determined based on the timecode of the second subtitle text; generating a third audio corresponding to each third subtitle text; and synthesizing all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed to obtain the final target video.
[0253] Furthermore, the logical instructions in the aforementioned memory 1130 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0254] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the video dubbing method provided in the above embodiments, the method comprising: separating a first audio from the video to be dubbed, performing sentence-by-sentence text conversion on the first audio to obtain a first subtitle text with a timecode; translating each first subtitle text into a second subtitle text, wherein the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text; obtaining a second audio corresponding to each second subtitle text, and, based on the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text, normalizing the second subtitle text to obtain a third subtitle text, wherein the subtitle display duration is determined based on the timecode of the second subtitle text; generating a third audio corresponding to each third subtitle text; and synthesizing all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed to obtain the final target video.
[0255] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the video dubbing method provided in the above embodiments. The method includes: separating a first audio from the video to be dubbed; performing sentence-by-sentence text conversion on the first audio to obtain a first subtitle text with a timecode; translating each first subtitle text into a second subtitle text, wherein the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text; obtaining a second audio corresponding to each second subtitle text; and, based on the difference between the audio duration of the second audio and the subtitle display duration of the second subtitle text, normalizing the second subtitle text to obtain a third subtitle text, wherein the subtitle display duration is determined based on the timecode of the second subtitle text; generating a third audio corresponding to each third subtitle text; and synthesizing all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed to obtain the final target video.
[0256] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0257] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0258] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A video dubbing method, characterized in that, include: After separating the first audio from the video to be dubbed, the first audio is converted into text sentence by sentence to obtain the first subtitle text with timecode; Each of the first subtitle texts is translated into a second subtitle text, and the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text. Obtain the second audio corresponding to each second subtitle text, and align the second subtitle text with a time window based on the audio duration of the second audio to obtain the third subtitle text; Generate the third dialogue audio corresponding to each of the third subtitle texts; The final target video is obtained by combining all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed.
2. The video dubbing method according to claim 1, characterized in that, Based on the difference between the audio duration of the second audio and the display duration of the second subtitle text, the second subtitle text is normalized to obtain the third subtitle text, which includes: If the speech rate normalization condition is met, obtain the ratio coefficient between the audio duration and the subtitle display duration; When the ratio coefficient is determined to be greater than 1, a summary and sentence reduction operation is performed on the second subtitle text to obtain the third subtitle text; When the ratio coefficient is determined to be less than 1, a word filling operation is performed on the second subtitle text to obtain the third subtitle text.
3. The video dubbing method according to claim 2, characterized in that, The process of performing a summary and sentence reduction operation on the second subtitle text to obtain the third subtitle text specifically includes: The number of words to be reduced when performing the summary reduction operation is determined based on the ratio coefficient and the number of words in the second subtitle text. A first prompt instruction is generated based on the reduced word count; The first prompt instruction and the second subtitle text are input into the text generation model to obtain the third subtitle text output by the text generation model.
4. The video dubbing method according to claim 2, characterized in that, Perform word-filling operation on the second subtitle text to obtain the third subtitle text, specifically including: The number of characters to be filled is determined based on the ratio coefficient and the number of characters in the second subtitle text; A second prompt instruction is generated based on the number of characters to be filled; The second prompt instruction and the second subtitle text are input into the text generation model to obtain the third subtitle text output by the text generation model.
5. The video dubbing method according to any one of claims 2-4, characterized in that, The speech rate normalization conditions include: Determine that the audio duration of the valid audio segment corresponding to the second audio is greater than the subtitle display duration of the second subtitle text; or, The audio duration of the effective audio segment is determined to be less than the subtitle display duration of the second subtitle text, and the difference between the two is greater than a preset threshold.
6. The video dubbing method according to any one of claims 2-4, characterized in that, The generation of the third audio corresponding to each of the third subtitle texts includes: Obtain the current speech rate coefficient for audio conversion of each of the third subtitle texts, as recorded in the speech synthesis engine; The target speech rate coefficient is obtained by adjusting the current speech rate coefficient based on the speech rate normalization coefficient. The current speech rate coefficient in the speech synthesis engine is adjusted to the target speech rate coefficient, so as to use the speech synthesis engine to perform audio conversion on the third subtitle text and generate the third audio. The speech rate normalization coefficient is determined based on the ratio coefficient between the audio duration and the subtitle display duration.
7. The video dubbing method according to claim 6, characterized in that, For any target third-level subtitle text, after adjusting the current speech rate coefficient based on the speech rate normalization coefficient to obtain the target speech rate coefficient, the process further includes: Obtain the first speech rate of the third audio corresponding to the preceding third subtitle text adjacent to the target third subtitle text in the timecode; Based on the timecode of the target third subtitle text, determine the dialogue audio segment of the first audio within the duration range of the timecode, so as to obtain the second speech rate of the dialogue audio segment; Based on the first speech rate and the second speech rate, a soft speech rate range is determined; When it is determined that the target speech rate for audio conversion of the target third subtitle text based on the target speech rate coefficient is outside the soft speech rate range, a prompt message is generated to guide the user to perform manual optimization. The target third subtitle text is any one of all the third subtitle texts.
8. The video dubbing method according to claim 1, characterized in that, The step of aligning the second subtitle text with a time window based on the audio duration of the second audio to obtain the third subtitle text includes: Based on the difference between the audio duration of the second audio and the display duration of the second subtitle text, the second subtitle text is normalized to obtain the third subtitle text, and the display duration of the subtitle text is determined based on the timecode of the second subtitle text.
9. The video dubbing method according to claim 1, characterized in that, For any target third subtitle text, generate the corresponding third audio, including: Determine the sentiment tag of the third audio corresponding to the target third subtitle text; The emotion tag and the target third subtitle text are input into the speech synthesis engine to generate the third audio.
10. The video dubbing method according to claim 9, characterized in that, The step of determining the emotion tag for the third audio corresponding to the target third subtitle text includes: Identify the target dialogue audio segment in the first audio, wherein the timecode of the target dialogue audio segment is the same as the timecode of the target third subtitle text; After identifying the target video segment in the video to be dubbed, the target video segment is subjected to frame extraction processing to obtain multiple video screenshots; the timecode of the target video segment is the same as the timecode of the target third subtitle text; The emotion tag of the third audio is determined based on at least one of the target dialogue audio segment, the multi-frame video screenshots, and the third subtitle text.
11. The video dubbing method according to claim 10, characterized in that, Determining the emotion tag of the third audio based on at least one of the target dialogue audio segment, the multi-frame video screenshots, and the third subtitle text includes: The target dialogue audio segment is input into the speech emotion recognition model to obtain the first emotion recognition result output by the speech emotion recognition model. The video screenshot is input into the image emotion recognition model to obtain the second emotion recognition result output by the image emotion recognition model; The target third subtitle text is input into the text emotion recognition model to obtain the third emotion recognition result output by the text emotion recognition model; The emotion label is determined based on at least one of the first emotion recognition result, the second emotion recognition result, and the third emotion recognition result.
12. The video dubbing method according to claim 1, characterized in that, For any target third subtitle text, generating the corresponding third audio for the target third subtitle text further includes: Determine the audio feature parameters of the third audio corresponding to the target third subtitle text, the audio feature parameters including volume parameters and / or pitch parameters; The audio feature parameters and the target third subtitle text are input into the speech synthesis engine to generate the third audio. The target third subtitle text is any one of all the third subtitle texts.
13. The video dubbing method according to claim 12, characterized in that, The volume parameter of the third audio corresponding to the target third subtitle text is determined based on the following steps: Based on the waveform of the first audio, determine the first average volume of the first audio. Based on the timecode of the target third subtitle text, determine the dialogue audio segment of the first audio within the duration interval of the timecode, and determine the second average volume of the dialogue audio segment based on the waveform of the dialogue audio segment; The volume parameters of the third audio corresponding to the target third subtitle text are determined based on the volume weighting coefficient between the second average volume and the first average volume, and the preset volume level in the speech synthesis engine.
14. The video dubbing method according to claim 12, characterized in that, The pitch parameters of the third audio corresponding to the target third subtitle text are determined based on the following steps: Obtain the spectrogram of the first audio, which is obtained by performing a Fourier transform on the global waveform of the first audio. The average first pitch of the first audio frequency is determined based on the spectrogram. Based on the timecode of the target third subtitle text, determine the dialogue audio segment of the first audio within the duration interval of the timecode, and determine the second pitch average of the dialogue audio segment based on the local spectrogram of the dialogue audio segment; Based on the pitch weight coefficient between the second pitch average and the first pitch average, and the preset pitch level in the speech synthesis engine, the pitch quantity parameter of the third audio corresponding to the target third subtitle text is determined.
15. The video dubbing method according to claim 1, characterized in that, For any target third subtitle text, generating the corresponding third audio for the target third subtitle text further includes: Identify the target dialogue audio segment in the first audio, wherein the timecode of the target dialogue audio segment is the same as the timecode of the target third subtitle text; The target dialogue audio segment is input into the timbre recognition model to obtain the gender and age information output by the timbre recognition model; The target speaker information is obtained by matching the gender and age information from the intelligent agent library, which stores different speaker information corresponding to different gender and age information. The target speaker information and the third subtitle text are input into the speech synthesis engine to generate the third audio. The target third subtitle text is any one of all the third subtitle texts.
16. The video dubbing method according to claim 1, characterized in that, The video dubbing method also includes a step of providing a dubbing preview, specifically including: Step 1: Divide all the third subtitle texts into multiple subtitle batches; Step 2: In response to the user's listening request, the current subtitle batch is processed by speech synthesis to generate the current dialogue audio batch, and the current dialogue audio batch is sent to the client for playback; Step 3: During the playback of the current dialogue audio batch on the client, speech synthesis is performed on the adjacent next subtitle batch to generate the next dialogue audio batch. Step 4: After confirming that the current dialogue audio has finished playing, send the next batch of dialogue audio to the client for playback in the order of timecodes. Repeat steps 3 to 4 until all subtitle batches have been traversed or a user's stop listening instruction is received.
17. A video dubbing device, characterized in that, include: The video parsing unit is used to separate the first audio from the video to be dubbed and then perform sentence-by-sentence text conversion on the first audio to obtain the first subtitle text with timecode; The subtitle translation unit is used to translate each of the first subtitle texts into a second subtitle text, wherein the timecode of each second subtitle text is the same as the timecode of the corresponding first subtitle text. The text normalization unit is used to obtain the second audio corresponding to each second subtitle text, and to perform time window alignment on the second subtitle text according to the audio duration of the second audio to obtain the third subtitle text; An audio generation unit is used to generate a third audio corresponding to each of the third subtitle texts; The video integration unit is used to synthesize all the third audio, the background audio of the video to be dubbed, and the silent video of the video to be dubbed to obtain the final target video.
18. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video dubbing method as described in any one of claims 1 to 16.
19. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video dubbing method as described in any one of claims 1 to 16.
20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video dubbing method as described in any one of claims 1 to 16.
Citation Information
Cited By
Video translation method and device
CN121814981A
Subtitle processing method and system, electronic equipment and storage medium
CN121815044A
A subtitle processing method, system, electronic device and storage medium
CN121815044B