Video decoding method and apparatus

CN121814981BActive Publication Date: 2026-06-23HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
Filing Date
2026-03-12
Publication Date
2026-06-23

Smart Images

  • Figure CN121814981B_ABST
    Figure CN121814981B_ABST
Patent Text Reader

Abstract

The application provides a video dubbing method and device. The method determines the corresponding human voice audio segment of each target subtitle text in the target language subtitle set based on the target language subtitle set and the audio file of the video to be dubbed. Then, the emotional category of the human voice audio segment corresponding to the target subtitle text and the long audio of the role to which the target subtitle text belongs are determined. After that, the target human voice audio segment corresponding to the target subtitle text is generated based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text and the long audio of the role to which the target subtitle text belongs. Finally, the video after dubbing is generated based on the target human voice audio segments corresponding to all target subtitle texts, which significantly improves the quality and availability of automatic dubbing and effectively improves the viewing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method and apparatus for translating video. Background Technology

[0002] Currently, short drama dubbing systems for overseas markets mainly rely on two approaches: fixed timbre libraries and single-sentence cloning. However, both have significant drawbacks: fixed timbre systems can only provide preset timbres and cannot reproduce the original character's voiceprint characteristics, resulting in a disconnect between the dubbing and the character's image; while single-sentence cloning systems support timbre reproduction, they are prone to timbre distortion in individual sentences due to excessively short input segments or noise interference, and lack correction functions for failed segments, leading to low-quality dubbing that is difficult to recover. In addition, the audio duration generated by both approaches is uncontrollable, easily causing the dubbing to become out of sync with the visuals and audio-visual synchronization, severely damaging the viewing experience. Summary of the Invention

[0003] In view of this, the present invention provides a video translation method and apparatus that significantly improves the quality and usability of automated dubbing and effectively enhances the viewing experience.

[0004] The first aspect of this invention provides a method for translating video, comprising:

[0005] Obtain the set of subtitles for the target language;

[0006] Based on the subtitle set in the target language and the audio file of the video to be translated, determine the human voice audio segment corresponding to each target subtitle text in the subtitle set in the target language;

[0007] Based on the target subtitle text and the corresponding human voice audio segment, determine the emotional category of the human voice audio segment corresponding to the target subtitle text;

[0008] Based on the human voice audio segment corresponding to the target subtitle text, determine the long audio of the character to whom the target subtitle text belongs;

[0009] Based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to whom the target subtitle text belongs, a target human voice audio segment corresponding to the target subtitle text is generated;

[0010] Based on the target human voice audio segments corresponding to all the target subtitle texts, a translated video is generated.

[0011] Optionally, obtaining the subtitle set for the target language includes:

[0012] The audio file of the video to be translated and the structured plot information are input into the speech-to-subtitle recognition model, and the output is a set of subtitle texts for the video to be translated. The structured plot information includes the overall plot information of the video to be translated and the plot information of the episode to which each audio segment belongs. For each audio segment, the speech-to-subtitle recognition model generates first text information and generates second text information based on the overall plot information of the video to be translated and the plot information of the episode to which the audio segment belongs. Based on the first text information and the second text information, the subtitle text corresponding to the audio segment is generated.

[0013] For each subtitle text in the set of subtitle texts of the video to be translated, the subtitle text is translated into the target language to obtain the target subtitle text.

[0014] Optionally, determining the human voice audio segment corresponding to each target subtitle text in the target language subtitle set based on the target language subtitle set and the audio file of the video to be translated includes:

[0015] Obtain the timeline order of the target subtitle text in the subtitle set of the target language, and use it as a timestamp sequence;

[0016] Speech activity detection is performed on the audio file of the video to be translated to obtain a speech segment sequence;

[0017] The time range of each target subtitle segment in the timestamp sequence is adjusted based on the speech segment sequence to obtain an adjusted timestamp sequence; wherein the time range of each target subtitle segment in the adjusted timestamp sequence is aligned with the speech segment in the speech segment sequence.

[0018] Optionally, determining the emotion category of the audio segment corresponding to the target subtitle text based on the target subtitle text and the corresponding human voice audio segment includes:

[0019] The emotion category of the audio segment corresponding to the target subtitle text is determined based on the posterior probability distribution of each emotion category of the target subtitle text and the posterior probability distribution of each emotion category of the corresponding human voice audio segment.

[0020] Optionally, determining the long audio of the character to whom the target subtitle text belongs based on the corresponding human voice audio segment includes:

[0021] The voiceprint features of the audio segment corresponding to the target subtitle text are identified to obtain the target voiceprint features;

[0022] Clustering is performed on all the target voiceprint features to obtain clustering results; wherein, the clustering results include multiple speakers; each speaker corresponds to at least one target voiceprint feature;

[0023] Based on the human voice audio segments corresponding to all target voiceprint features of the same speaker, a long audio of the speaker is generated;

[0024] Based on the speaker's long audio, determine the speaker's acoustic attribute information;

[0025] Based on the speaker's long audio, the speaker's acoustic attribute information, and the character's timbre profile, the long audio of the character to whom the target subtitle text belongs is determined.

[0026] Optionally, determining the long audio of the character to which the target subtitle text belongs based on the speaker's long audio, the speaker's acoustic attribute information, and the character's timbre profile includes:

[0027] A candidate set is determined based on the speaker's long audio and the long audio of the character in the character timbre file;

[0028] For each candidate role in the candidate set, the role corresponding to the speaker is determined based on the long audio of the candidate role, the acoustic role attribute information of the candidate role, the long audio of the speaker, and the acoustic attribute information of the speaker.

[0029] Based on the correspondence between the target subtitle text and the speaker's long audio, and the role corresponding to the speaker, the long audio of the role to which the target subtitle text belongs is determined.

[0030] Optionally, for each candidate role in the candidate set, based on the candidate role's long audio file, the candidate role's acoustic role attribute information, the speaker's long audio file, and the speaker's acoustic attribute information, the corresponding role of the speaker is determined, including:

[0031] Based on the long audio of the candidate role and the long audio of the speaker, the voiceprint channel score is determined;

[0032] The style channel score is determined based on the acoustic attribute information of the candidate role and the acoustic attribute information of the speaker;

[0033] Based on the voiceprint channel score and the style channel score, the dual-channel matching score between the candidate role and the speaker is determined; wherein, the candidate role with the largest dual-channel matching score is taken as the role corresponding to the speaker.

[0034] Optionally, generating the target voice audio segment corresponding to the target subtitle text based on the target subtitle text, the duration of the target subtitle text, the emotional category of the corresponding human voice audio segment, and the long audio of the character to whom the target subtitle text belongs includes:

[0035] The target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to which the target subtitle text belongs are input into the timbre replication model to obtain the target human voice audio segment corresponding to the target subtitle text.

[0036] Optionally, the timbre replication model includes: an audio quantization module, a text encoding module, a semantic encoding module, an attribute extraction module, and an audio dequantization module;

[0037] The audio quantization module quantizes the long audio of the character to whom the target subtitle text belongs into a discrete digital index sequence;

[0038] The text encoding module encodes the target subtitle text into a text feature vector sequence;

[0039] The attribute extraction module performs joint encoding based on the duration of the target subtitle text, the emotion category of the human voice audio segment corresponding to the target subtitle text, and the voiceprint features of the long audio of the character to which the target subtitle text belongs, to generate an attribute condition vector set.

[0040] The semantic encoding module processes the semantic feature vector sequence and the attribute condition vector set to obtain the audio quantization index sequence corresponding to the target subtitle text; wherein, the semantic feature vector sequence is obtained by temporally concatenating the discrete digital index sequence and the text feature vector sequence.

[0041] The audio dequantization module generates the target human voice audio segment corresponding to the target subtitle text based on the audio quantization index sequence.

[0042] Optionally, after generating the target voice audio segment corresponding to the target subtitle text based on the target subtitle text, the duration of the target subtitle text, the emotional category of the voice audio segment corresponding to the target subtitle text, and the long audio of the character to whom the target subtitle text belongs, the process further includes:

[0043] If the actual speech rate of the target human voice audio segment corresponding to the target subtitle text exceeds the speech rate threshold, then the target subtitle text and its context data are input into the subtitle simplification model, and the simplified target subtitle text is output.

[0044] Optionally, the subtitle simplification model includes a semantic encoder, a constraint encoder, and a decoder;

[0045] The semantic encoder generates a semantic representation based on the target subtitle text and its contextual data;

[0046] The constraint encoder generates constraint feature representations based on pseudo-target subtitle text and simplification ratio; wherein, the pseudo-target subtitle text is obtained by translating the target subtitle text, and the simplification ratio is the length ratio of the target subtitle text to the pseudo-target subtitle text;

[0047] The decoder generates a simplified target subtitle text based on the semantic representation and the constraint feature representation.

[0048] Optionally, generating the translated video based on the target human voice audio segments corresponding to all the target subtitle texts includes:

[0049] Align all the target audio segments corresponding to the target subtitle text and the target subtitle text according to the timeline of the subtitles in the video to be translated;

[0050] The target human voice audio track is subjected to dynamic range compression and loudness normalization to obtain the processed target human voice audio track.

[0051] The processed target human voice audio segment is mixed and synthesized with the original ambient audio track to obtain the translated video.

[0052] A second aspect of the present invention provides a video translation apparatus, comprising:

[0053] The acquisition unit is used to acquire a set of subtitles in the target language;

[0054] The human voice audio segment determination unit is used to determine the human voice audio segment corresponding to each target subtitle text in the target language subtitle set based on the target language subtitle set and the audio file of the video to be translated.

[0055] The emotion category determination unit is used to determine the emotion category of the human voice audio segment corresponding to the target subtitle text based on the target subtitle text and the human voice audio segment corresponding to the target subtitle text;

[0056] The character long audio determination unit is used to determine the long audio of the character to which the target subtitle text belongs based on the human voice audio segment corresponding to the target subtitle text;

[0057] The target human voice audio segment generation unit is used to generate the target human voice audio segment corresponding to the target subtitle text based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to which the target subtitle text belongs;

[0058] The translated video generation unit is used to generate translated videos based on the target human voice audio segments corresponding to all the target subtitle texts.

[0059] Optionally, the acquisition unit includes:

[0060] The first input unit is used to input the audio file of the video to be translated and the structured plot information into the speech-to-subtitle recognition model, and output a set of subtitle texts for the video to be translated; wherein, the structured plot information includes the overall plot information of the video to be translated and the plot information of the episode to which each audio segment belongs in the audio file, the speech-to-subtitle recognition model generates first text information for each audio segment, and generates second text information based on the overall plot information of the video to be translated and the plot information of the episode to which the audio segment belongs, and generates subtitle texts corresponding to the audio segment based on the first text information and the second text information;

[0061] The translation unit is used to translate each subtitle text in the set of subtitle texts of the video to be translated into the target language to obtain the target subtitle text.

[0062] Optionally, the human voice audio segment determination unit includes:

[0063] The timestamp sequence acquisition unit is used to acquire the timeline order of the target subtitle text in the subtitle set of the target language as a timestamp sequence;

[0064] The speech activity detection unit is used to perform speech activity detection on the audio file of the video to be translated to obtain a speech segment sequence;

[0065] A time range adjustment unit is used to adjust the time range of each target subtitle segment in the timestamp sequence based on the speech segment sequence to obtain an adjusted timestamp sequence; wherein the time range of each target subtitle segment in the adjusted timestamp sequence is aligned with the speech segment in the speech segment sequence.

[0066] Optionally, the emotion category determination unit includes:

[0067] The emotion category determination subunit is used to determine the emotion category of the human voice audio segment corresponding to the target subtitle text based on the posterior probability distribution of each emotion category of the target subtitle text and the posterior probability distribution of each emotion category of the human voice audio segment corresponding to the target subtitle text.

[0068] Optionally, the character long audio determination unit includes:

[0069] The voiceprint feature recognition unit is used to recognize the voiceprint features of the human voice audio segment corresponding to the target subtitle text, and obtain the target voiceprint features;

[0070] A clustering unit is used to cluster all the target voiceprint features to obtain a clustering result; wherein the clustering result includes multiple speakers; each speaker corresponds to at least one target voiceprint feature;

[0071] The speaker long audio generation unit is used to generate a long audio of the speaker based on the human voice audio segments corresponding to all target voiceprint features of the same speaker;

[0072] An acoustic attribute information determination unit is used to determine the acoustic attribute information of the speaker based on the speaker's long audio signal;

[0073] The character long audio determination subunit is used to determine the long audio of the character to which the target subtitle text belongs based on the speaker's long audio, the speaker's acoustic attribute information, and the character timbre profile.

[0074] Optionally, the character long audio determination subunit includes:

[0075] The candidate set determination subunit is used to determine the candidate set based on the speaker's long audio and the long audio of the character in the character timbre file;

[0076] The role determination subunit is used to determine the role corresponding to the speaker for each candidate role in the candidate set based on the long audio of the candidate role, the acoustic role attribute information of the candidate role, the long audio of the speaker, and the acoustic attribute information of the speaker.

[0077] The long audio determination subunit of the role to which the target subtitle text belongs is used to determine the long audio of the role to which the target subtitle text belongs based on the correspondence between the target subtitle text and the speaker's long audio and the role corresponding to the speaker.

[0078] Optionally, the role determination subunit includes:

[0079] The voiceprint channel score determination subunit is used to determine the voiceprint channel score based on the long audio of the candidate role and the long audio of the speaker;

[0080] The style channel score determination subunit is used to determine the style channel score based on the acoustic attribute information of the candidate role and the acoustic attribute information of the speaker;

[0081] The dual-channel matching score determination subunit is used to determine the dual-channel matching score between the candidate role and the speaker based on the voiceprint channel score and the style channel score; wherein, the candidate role with the largest dual-channel matching score is taken as the role corresponding to the speaker.

[0082] Optionally, the target human voice audio segment generation unit includes:

[0083] The second input unit is used to input the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to which the target subtitle text belongs into the timbre replication model to obtain the target human voice audio segment corresponding to the target subtitle text.

[0084] Optionally, the timbre replication model includes: an audio quantization module, a text encoding module, a semantic encoding module, an attribute extraction module, and an audio dequantization module;

[0085] The audio quantization module quantizes the long audio of the character to whom the target subtitle text belongs into a discrete digital index sequence;

[0086] The text encoding module encodes the target subtitle text into a text feature vector sequence;

[0087] The attribute extraction module performs joint encoding based on the duration of the target subtitle text, the emotion category of the human voice audio segment corresponding to the target subtitle text, and the voiceprint features of the long audio of the character to which the target subtitle text belongs, to generate an attribute condition vector set.

[0088] The semantic encoding module processes the semantic feature vector sequence and the attribute condition vector set to obtain the audio quantization index sequence corresponding to the target subtitle text; wherein, the semantic feature vector sequence is obtained by temporally concatenating the discrete digital index sequence and the text feature vector sequence.

[0089] The audio dequantization module generates the target human voice audio segment corresponding to the target subtitle text based on the audio quantization index sequence.

[0090] Optionally, the video translation device further includes:

[0091] The third input unit is used to input the target subtitle text and its context data into the subtitle simplification model if the actual speech rate of the target human voice audio segment corresponding to the target subtitle text exceeds the speech rate threshold, and output the simplified target subtitle text.

[0092] Optionally, the subtitle simplification model includes a semantic encoder, a constraint encoder, and a decoder;

[0093] The semantic encoder generates a semantic representation based on the target subtitle text and its contextual data;

[0094] The constraint encoder generates constraint feature representations based on pseudo-target subtitle text and simplification ratio; wherein, the pseudo-target subtitle text is obtained by translating the target subtitle text, and the simplification ratio is the length ratio of the target subtitle text to the pseudo-target subtitle text;

[0095] The decoder generates a simplified target subtitle text based on the semantic representation and the constraint feature representation.

[0096] Optionally, the translated video generation unit includes:

[0097] The alignment unit is used to align all the target audio segments corresponding to the target subtitle text and the target subtitle text according to the timeline of the subtitles in the video to be translated.

[0098] The audio segment processing unit is used to perform dynamic range compression and loudness normalization on the audio track of the target human voice audio segment to obtain the processed target human voice audio segment.

[0099] The synthesis unit is used to mix and synthesize the processed target human voice audio clip with the original ambient audio track to obtain the translated video.

[0100] A third aspect of the present invention provides an electronic device, comprising:

[0101] One or more processors;

[0102] A storage device on which one or more programs are stored;

[0103] When the one or more programs are executed by the one or more processors, the one or more processors implement the video translation method as described in any one of the first aspects.

[0104] A fourth aspect of the present invention provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the video translation method as described in any one of the first aspects.

[0105] As can be seen from the above solutions, the present invention provides a method and apparatus for translating videos. The method determines the human voice audio segment corresponding to each target subtitle text in the target language subtitle set based on a subtitle set in the target language and the audio file of the video to be translated. Then, it determines the emotional category of the human voice audio segment corresponding to the target subtitle text and the long audio of the character to whom the target subtitle text belongs. Subsequently, based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to whom the target subtitle text belongs, it generates the target human voice audio segment corresponding to the target subtitle text. Finally, based on the target human voice audio segments corresponding to all target subtitle texts, it generates the translated video, significantly improving the quality and usability of automated dubbing and effectively enhancing the viewing experience. Attached Figure Description

[0106] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0107] Figure 1 A detailed flowchart of a video translation method provided in an embodiment of the present invention;

[0108] Figure 2 A flowchart of a video translation method provided in another embodiment of the present invention;

[0109] Figure 3 A flowchart of a video translation method provided in another embodiment of the present invention;

[0110] Figure 4 A dual-path acoustic signature architecture is provided in another embodiment of the present invention;

[0111] Figure 5 A flowchart of a video translation method provided in another embodiment of the present invention;

[0112] Figure 6 An architectural diagram of a timbre replication model provided for another embodiment of the present invention;

[0113] Figure 7 This is a schematic diagram of a video translation device provided in another embodiment of the present invention. Detailed Implementation

[0114] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0115] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0116] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties.

[0117] It should be noted that the concepts of "first" and "second" mentioned in this invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0118] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0119] This invention provides a method for translating video, such as... Figure 1 As shown, the specific steps include:

[0120] S101. Obtain the set of subtitles for the target language.

[0121] It is understandable that the set of subtitles for the target language can be prepared in advance by the user, or it can be obtained by the user inputting the video to be translated, selecting the target language, and then processing the video. There is no limitation here.

[0122] Optionally, in another embodiment of the present invention, one implementation of step S101 specifically includes:

[0123] The audio file and structured plot information of the video to be translated are input into the speech subtitle recognition model, and the output is a set of subtitle texts for the video to be translated. For each subtitle text in the set of subtitle texts for the video to be translated, the subtitle text is translated into the target language to obtain the target subtitle text.

[0124] The structured plot information includes the overall plot information of the video to be translated and the plot information of each episode to which each audio segment belongs in the audio file. This structured plot information can be provided in advance by the user, obtained from relevant public websites, etc., and is not limited here. Plot information includes, but is not limited to, historical periods, geographical environments, world-building, and character information, etc., and is not limited here.

[0125] It is understandable that the audio file of the video to be translated can be prepared in advance by the user, or it can be obtained by processing the video after the user inputs it; there is no limitation here.

[0126] Specifically, a standard audio / video separation module can be used to decompose the input video stream into independent video streams and raw audio streams, ensuring the integrity and timing synchronization of the audio data.

[0127] In another embodiment of the present invention, to further improve the quality of the subtitle text set of the video to be translated, audio files can be processed using, but not limited to, deep learning-based audio separation techniques. This involves using spectral masking prediction to identify and separate human voice audio components from environmental audio components, thereby obtaining the human voice audio file of the video to be translated. The human voice audio file and structured plot information of the video to be translated are then input into the speech subtitle recognition model, which outputs the subtitle text set of the video to be translated.

[0128] Of course, the audio files of the videos to be translated can also be prepared by the user in advance; there are no restrictions here.

[0129] It should be noted that the audio slices in this embodiment are not the same as the human voice audio segments in subsequent embodiments. The audio slices only need to be sliced ​​at an approximate location, because here we only need to determine the approximate location of the subtitles. The human voice audio segments are at a more precise location. The main purpose is to ensure that the subtitle timeline fully covers the entire process from the start to the end of the pronunciation, without any speech interruption or redundant silence.

[0130] In the practical application of this invention, the speech-to-subtitle recognition model generates first text information for each audio segment, and generates second text information based on the overall plot information of the video to be translated and the plot information of the episode to which the audio segment belongs. Based on the first and second text information, the subtitle text corresponding to the audio segment is generated.

[0131] Specifically, the audio slices are quantized into discrete digital index sequences and mapped onto the text space through multiple projections to obtain the first text information. The overall plot information of the video to be translated and the plot information of the episode to which the audio segments belong are converted into a text feature vector sequence through a text embedding layer. The first text information is the second text information. The first text information and the second text information are concatenated to obtain the input sequence. The input sequence is then fed into a neural network layer for processing to obtain the subtitle text corresponding to the sliced ​​audio.

[0132] In the process of concatenating the first and second text information to obtain the input sequence, a modal delimiter can also be used to separate the first and second text information, resulting in the final input sequence. It can be No restrictions are imposed here.

[0133] It should be noted that the speech-to-text recognition model is trained in two stages. The first stage (mapping ability learning) freezes the main parameters of the large language model, optimizing only the projection layer and delimiter embedding, focusing on establishing the mapping relationship between the input audio segments and the text semantic space. The second stage (co-optimization) unfreezes all network parameters and performs end-to-end joint training, using the standard cross-entropy criterion as the loss function.

[0134] ;

[0135] Where Y is the predicted label, which is the actual subtitle text corresponding to the audio slice. This indicates that the model predicts the output at the current time step. The posterior probability.

[0136] It is understood that, in the actual application of the present invention, the implementation schemes in the above embodiments can also be applied to scenarios where the target language subtitles in the user-provided target language subtitle set are missing some target language subtitle texts. To fill in the missing target language subtitle texts, specifically, the position of the missing target language part is first determined, the audio slice at that position is obtained, the audio slice at that position is recognized using the speech subtitle recognition model proposed in the present invention, the missing subtitle text is obtained, and then the subtitle text is translated to complete the filling in of the missing target language subtitle texts.

[0137] Of course, in the practical application of this invention, the audio slices can be aligned first (for specific alignment methods, please refer to...). Figure 2After obtaining more accurate human voice audio segments (as shown in the embodiment), a voice subtitle recognition model is used to generate first text information for each human voice audio segment. Then, based on the overall plot information of the video to be translated and the plot information of the episode to which the human voice audio segment belongs, second text information is generated. Based on the first text information and the second text information, subtitle text corresponding to the human voice audio segment is generated.

[0138] The proposed speech-to-subtitle recognition model innovatively structures video plot-related information into a multi-dimensional prompt input. Simultaneously, it designs a lightweight projection mapping mechanism to convert the quantized speech digit index sequence into a text embedding space. These two information sources are concatenated at the input layer of a large language model, enabling the model to collaboratively utilize acoustic features and plot context to achieve accurate speech-to-subtitle recognition in film and television scenarios.

[0139] S102. Based on the subtitle set in the target language and the audio file of the video to be translated, determine the human voice audio segment corresponding to each target subtitle text in the subtitle set in the target language.

[0140] It is understandable that the audio file of the video to be translated can be prepared in advance by the user, or it can be obtained by processing the video after the user inputs it; there is no limitation here.

[0141] Specifically, a standard audio / video separation module can be used to decompose the input video stream into independent video streams and raw audio streams, ensuring the integrity and timing synchronization of the audio data.

[0142] In another embodiment of the present invention, to further improve the quality of the human voice audio segment corresponding to each target subtitle text, audio separation technology based on deep learning can be applied to the audio file, but is not limited to, using spectral mask prediction to identify and separate the human voice audio component from the environmental audio component, thereby obtaining the human voice audio file of the video to be translated. This is not limited here. Based on the subtitle set in the target language and the human voice audio file of the video to be translated, the human voice audio segment corresponding to each target subtitle text in the subtitle set in the target language is determined.

[0143] Of course, the audio files of the videos to be translated can also be prepared by the user in advance; there are no restrictions here.

[0144] To address the discrepancy between the subtitle timeline and the actual audio range (such as a delayed start of subtitles leading to missing vowels, or an premature end causing truncation of the final sound), in another embodiment of the present invention, one implementation of step S102 is as follows: Figure 2 As shown, it includes:

[0145] S201. Obtain the timeline order of the target subtitle text in the target language subtitle set as a timestamp sequence.

[0146] In the actual application of this invention, the timeline order of the target subtitle text can be provided by the subtitle collection of the target language, or it can be provided by the user, or an automatic timing tool can be called to timing the target subtitle text, etc., which is not limited here.

[0147] The time-axis sequence can be represented as: [[ , ],[ , ],...,[ , ]],in, For the first The start time of each subtitle segment This is the end time.

[0148] S202. Perform speech activity detection on the audio file of the video to be translated to obtain a speech segment sequence.

[0149] The speech segment sequence may, but is not limited to, use the Voice Activity Detection (VAD) algorithm to detect the human voice audio file of the video to be translated, and the detection result is used as the speech segment sequence, such as: [[ , ],[ , ],...,[ , ]],in, For the first The start time of each speech segment This is the end time.

[0150] S203. Adjust the time range of each target subtitle segment in the timestamp sequence based on the speech segment sequence to obtain the adjusted timestamp sequence.

[0151] In this process, the time range of each target subtitle segment in the adjusted timestamp sequence is aligned with the audio segment in the audio segment sequence.

[0152] Specifically, retrieve each target subtitle fragment in the traversal timestamp sequence. , Search the audio segment sequence for all audio segments that overlap with the time range of the current target subtitle segment; if overlapping audio segments exist, then move the target subtitle segment... The start time is adjusted to the minimum start time among all overlapping speech segments. ; Set the end time of the subtitles Adjusted to the maximum end time among all overlapping speech segments (max). If there are no overlapping audio segments, the time boundaries of the original target subtitle segment are preserved.

[0153] In the practical application of this invention, the rationality of the adjusted target subtitle segment time boundary can also be verified to ensure... Boundary conflict detection is performed on adjacent subtitle segments to avoid time overlap; minimum duration constraints are imposed on excessively short audio segments to ensure audio integrity, but no restrictions are imposed in this case.

[0154] Through the above-described embodiments, the present invention ensures that the time range of each target subtitle segment in the adjusted timestamp sequence is aligned with the audio segment in the audio segment sequence, thereby ensuring that the subtitle timeline fully covers the entire process from the start to the end of the pronunciation, without any audio truncation or redundant silence.

[0155] S103. Based on the target subtitle text and the corresponding human voice audio segment, determine the emotion category of the human voice audio segment corresponding to the target subtitle text.

[0156] In the practical application of this invention, the emotion category of the human voice audio segment corresponding to the target subtitle text can be determined based on the posterior probability distribution of each emotion category of the target subtitle text and the posterior probability distribution of each emotion category of the human voice audio segment corresponding to the target subtitle text.

[0157] Specifically, existing text sentiment recognition models can be used to predict the sentiment of target subtitle text, outputting the posterior probability distribution of sentiment categories 1 to D: The existing speech emotion recognition model is used to predict the emotion of the human voice audio segment corresponding to the target subtitle text, and the posterior probability distribution of emotion category 1 to D is output: The weighted fusion formula is used to calculate the overall posterior probability. : ,in For the preset text modality weights ( The emotional category of the audio segment corresponding to the final target subtitle text is the category with the highest comprehensive posterior probability.

[0158] S104. Based on the human voice audio segment corresponding to the target subtitle text, determine the long audio of the character to whom the target subtitle text belongs.

[0159] In the actual application of this invention, the character to which the target subtitle text belongs will be analyzed and judged, so that the character's audio can be more effectively replicated in subsequent steps based on the acoustic attribute information of the character to which the target subtitle text belongs.

[0160] Among them, acoustic attributes include, but are not limited to, age, gender, accent, speech rate, emotion, pitch, and pitch fluctuation, etc., which are not limited here.

[0161] It should be noted that S103 and S104 in this invention can be performed in any order, or step S104 can be executed first and then step S103, or steps S103 and S104 can be executed simultaneously. No limitation is made here.

[0162] Optionally, in another embodiment of the present invention, one implementation of step S104 is as follows: Figure 3 As shown, it includes:

[0163] S301. Identify the voiceprint features of the audio segment corresponding to the target subtitle text to obtain the target voiceprint features.

[0164] In the practical application of this invention, a dual-path voiceprint architecture is provided, such as... Figure 4 As shown, voiceprint feature extraction from audio segments is achieved through two voiceprint models (segment-level voiceprint model and character-level voiceprint model). This architecture contains two dedicated voiceprint models that share the underlying feature extraction network, but employ a differentiated dual-branch design in the top-level pooling layer.

[0165] 1) Segment-level voiceprint model: It adopts the statistical pooling mechanism, which is suitable for short speech segments whose acoustic features are non-stationary. Statistical pooling can capture the central trend and dynamic variation characteristics of timbre to generate timbre embeddings that are robust to local noise and speech rate changes. It is suitable for rapid grouping in the timbre clustering stage.

[0166] Specifically, after extracting acoustic feature sequences (dimension t×d) from human voice audio segments, the mean (dimension 1×d) and standard deviation (dimension 1×d) of the feature sequences are calculated, and the sequences are concatenated along the feature dimensions to output a timbre embedding vector (dimension 1×2d), which is the target voiceprint feature.

[0167] 2) Role-based voiceprint model: It adopts the attentional pooling mechanism, which is suitable for long audio. It can adaptively learn the weight distribution of each time step and dynamically focus on the key speech segments with the most speaker recognition. It is suitable for high-precision recognition in the role matching stage.

[0168] Specifically, for a long audio acoustic feature sequence (dimension t×d), the weight coefficients (dimension t×1) of each time step are adaptively calculated through the network layer, and the weighted aggregation is used to output a timbre embedding vector (dimension 1×d).

[0169] This invention addresses the scale mismatch problem in timbre representation between short-term audio segments and long-term character speech by constructing a short-term robust branch and a long-term stylization branch in a dual-path voiceprint architecture.

[0170] S302. Cluster all target voiceprint features to obtain clustering results.

[0171] The clustering results include multiple speakers; each speaker corresponds to at least one target voiceprint feature.

[0172] Since voiceprint features carry speaker information, clustering voiceprint features will group those with similar timbre into one category and determine them as the same speaker.

[0173] In the practical application of this invention, a clustering algorithm can be used to perform K-means clustering on all target voiceprint features, and a speaker label can be assigned to the human voice audio segment corresponding to each target voiceprint feature. Human voice audio segments of the same speaker are aggregated and spliced ​​to form a long audio segment of the role, and a mapping relationship between the human voice audio segments and the speaker's long audio segment is established.

[0174] Specifically, based on the K-means clustering results, the human voice audio segment corresponding to each target voiceprint feature is assigned to the corresponding cluster; according to the number of clusters formed by clustering, they are labeled as speaker 1, speaker 2, ..., speaker K in the order of cluster index.

[0175] S303. Generate a long audio file of the speaker based on the human voice audio segments corresponding to all target voiceprint features of the same speaker.

[0176] In this process, the audio segments of the voice corresponding to all target voiceprint features of the same speaker can be sorted in chronological order, or they can be sorted in terms of length. There is no limitation here.

[0177] S304. Determine the speaker's acoustic attribute information based on the speaker's long audio.

[0178] Specifically, the speaker's long audio is input into the attribute analysis model, and the speaker's acoustic attribute information is output.

[0179] The attribute analysis model consists of several specialized sub-models: the age, gender, accent, and emotion sub-models are pre-trained deep learning classification models; pitch attributes can be obtained by calculating the mean fundamental frequency using the pitch sub-model based on the fundamental frequency extraction algorithm; pitch fluctuation can be obtained by calculating the standard deviation or coefficient of variation of the fundamental frequency sequence using the pitch fluctuation sub-model; and speech rate attributes can be determined by the speech rate sub-model by statistically analyzing the number of speech units per unit time (estimated based on energy zero-crossing rate analysis), without further limitations. These sub-models work collaboratively to perform multi-dimensional acoustic attribute analysis on long audio clips at the character level.

[0180] S305. Based on the speaker's long audio, the speaker's acoustic attribute information, and the character's timbre profile, determine the long audio of the character to whom the target subtitle text belongs.

[0181] Optionally, in another embodiment of the present invention, one implementation of step S305 is as follows: Figure 5 As shown, it includes:

[0182] S501. Determine a candidate set based on the speaker's long audio and the long audio of the character in the character timbre profile.

[0183] Optionally, in another embodiment of the present invention, one implementation of step S501 specifically includes the following steps (steps A1 to A2):

[0184] Step A1: Extract the voiceprint features of the speaker's long audio.

[0185] The method for extracting the speaker's long audio voiceprint features can be implemented using the dual-path voiceprint architecture described in the above embodiments, which will not be elaborated here.

[0186] Step A2: Determine the candidate set based on the speaker's long audio voiceprint features, the voiceprint features of the characters' long audio voiceprints in the character timbre profile, and the voiceprint feature matching threshold.

[0187] The character voice profile includes voiceprint features and style feature vectors for long audio recordings of multiple characters.

[0188] Set up a collection of character voice files The speaker's long audio is The speaker's long audio recordings have the following voiceprint characteristics: .

[0189] Among them, the voiceprint feature matching threshold Pre-settings and modifications can be made by experts or authorized technical personnel; no restrictions are imposed here.

[0190] In the practical application of this invention, the speaker's long audio voiceprint features are first calculated based on the speaker's long audio voiceprint features and the long audio voiceprint features of the characters in the character timbre profile. Characters in the character voice profile Voiceprint characteristics of long audio frequencies Voiceprint channel fraction ,reserve The candidate set.

[0191] S502. For each candidate role in the candidate set, determine the role corresponding to the speaker based on the candidate role's long audio, the candidate role's acoustic role attribute information, the speaker's long audio, and the speaker's acoustic attribute information.

[0192] Optionally, in another embodiment of the present invention, one implementation of step S501 specifically includes the following steps (steps B1 to B3):

[0193] Step B1: Determine the voiceprint channel score based on the long audio of the candidate role and the speaker.

[0194] It should be noted that the implementation method of voiceprint channel fraction can be found in the corresponding part of the implementation method in step A2, and will not be repeated here.

[0195] Step B2: Determine the style channel score based on the acoustic attribute information of the candidate role and the speaker.

[0196] Style Channel Score The calculation method can be as follows:

[0197] ;in, This is the style feature vector of a character from the character's voice profile. The speaker's acoustic attributes are structured into a text description (e.g., "Age: Young, Gender: Male, Accent: None, Speech Rate: Fast, Emotion: Happy, Pitch: Medium, Pitch Fluctuation: Low"). The hidden vector of this text description is extracted using a large language model as the style feature vector. .

[0198] Step B3: Determine the dual-channel matching score between the candidate role and the speaker based on the voiceprint channel score and the style channel score.

[0199] Among them, the candidate role with the highest dual-channel matching score is taken as the role corresponding to the speaker.

[0200] Dual-channel matching score The calculation method can be as follows:

[0201] ;in, The fusion coefficient is preset and modified by experts or authorized technical personnel, and is not limited here.

[0202] It is understandable that if the matching method in the above embodiments fails to match the speaker's corresponding role, the speaker will be treated as a new role, and its long audio and acoustic attribute information will be stored in the timbre archive.

[0203] This invention constructs a dual-channel independent scoring mechanism for voiceprint and style, and achieves highly robust character matching by fusing the voiceprint channel score and style channel score with confidence weighting, effectively ensuring the consistency of character timbre and expression style across segments and episodes.

[0204] S503. Based on the correspondence between the target subtitle text and the speaker's long audio and the speaker's corresponding role, determine the long audio of the role to which the target subtitle text belongs.

[0205] It should be noted that in S302, a mapping relationship is established between the audio segments of human voices and the long audio recordings of the speaker. Here, the audio segment of human voice refers to the audio segment of human voice corresponding to the target text subtitle. In other words, the mapping relationship established in S302 between the audio segments of human voices and the long audio recordings of the speaker also reflects the correspondence between the target text subtitle and the long audio recording of the speaker. Since the correspondence between the speaker and the role is determined in step S502, the correspondence between the target text subtitle and the role can be determined. Each role has a long audio recording, so the long audio recording of the role to which the target subtitle text belongs can be determined.

[0206] S105. Based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to whom the target subtitle text belongs, generate the target human voice audio segment corresponding to the target subtitle text.

[0207] In the practical application of this invention, the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to which the target subtitle text belongs can be input into the timbre replication model to obtain the target human voice audio segment corresponding to the target subtitle text.

[0208] Among them, such as Figure 6 As shown, the timbre replication model includes: an audio quantization module, a text encoding module, a semantic encoding module, an attribute extraction module, and an audio dequantization module.

[0209] Specifically, the audio quantization module quantizes the long audio of the character to whom the target subtitle text belongs into a discrete digital index sequence. Each index corresponds to a high-dimensional quantized feature vector.

[0210] The text encoding module encodes the target subtitle text into a sequence of text feature vectors. .

[0211] Specifically, the target subtitle text is first converted into a phoneme sequence, and then each phoneme is mapped to a fixed-dimensional vector through an embedding network layer.

[0212] The attribute extraction module performs joint encoding based on the duration of the target subtitle text, the emotional category of the corresponding human voice audio segment, and the voiceprint features of the long audio clip of the character to whom the target subtitle text belongs, generating an attribute condition vector set. .in The dimension of the feature vector.

[0213] The semantic encoding module processes the semantic feature vector sequence and attribute condition vector set to obtain the audio quantization index sequence corresponding to the target subtitle text.

[0214] Among them, semantic feature vector sequence It is obtained by temporal concatenation of discrete numerical index sequences and text feature vector sequences. .

[0215] Specifically, the semantic encoding module performs high-dimensional information modeling on the semantic feature vector sequence and attribute condition vector set to predict the audio quantization index sequence corresponding to the target subtitle text.

[0216] The audio dequantization module generates the target human voice audio segment corresponding to the target subtitle text based on the audio quantization index sequence.

[0217] Specifically, the audio dequantization module restores the audio quantization index sequence into an audio signal, and uses the restored result as the target human voice audio segment corresponding to the target subtitle text.

[0218] In the practical application of this invention, the attribute extraction module performs refined encoding on the input three-dimensional attribute information (duration of the target subtitle text, emotional category of the human voice audio segment corresponding to the target subtitle text, and voiceprint features of the long audio of the character to whom the target subtitle text belongs) to generate an attribute feature vector sequence. The encoding process for the three attributes can be as follows:

[0219] 1) Duration Constraint Encoding: Encoding the duration of the target subtitle text The multi-band Fourier feature mapping is used to convert the feature representation into a continuous feature representation for each frequency band. Define the basic coding unit:

[0220] ;

[0221] in Using the reference frequency, all frequency band coding units are concatenated to form a duration-constrained coding vector. , dimension :

[0222] ;

[0223] 2) Sentiment Feature Encoding: The sentiment category of the human voice audio segment corresponding to the target subtitle text is encoded using a learnable embedding matrix. Mapped to the corresponding feature vector ,in Indicates the number of emotion categories.

[0224] 3) Voiceprint feature encoding: utilizing Figure 4 The dual-path voiceprint architecture in the text extracts the voiceprint features of the long audio recordings of the characters to whom the target subtitle text belongs. .

[0225] Finally, the three feature vectors are fused through a conditional normalization layer to obtain the attribute feature vector sequence. :

[0226] .

[0227] In the practical application of this invention, the semantic encoding module can adopt a multi-layer Transformer architecture, which is not limited here.

[0228] This invention designs an attribute-aware cross-attention mechanism in the first layer of the semantic encoding module. This mechanism uses the semantic feature vector sequence S as the query source and the attribute feature vector sequence... As keys and values, dynamic feature fusion under multiple attribute conditions is achieved. The specific calculation process is as follows:

[0229] 1) First, calculate the query vector matrix Q, the key vector matrix K, and the value vector matrix V:

[0230] ;

[0231] ;

[0232] ;

[0233] in, The length of the semantic feature vector sequence. It is a fully connected layer with learnable parameters.

[0234] 2) Computation of attribute-aware feature vector sequences :

[0235] ;

[0236] ;

[0237] in, express The Middle row (i.e., the first row) The feature vector at time (time). express The Middle Column feature vectors Calculate the attention score (degree of attention) of semantic features on the three attribute vectors. express The Middle The eigenvectors are obtained by multiplying the three scores by the value vector, resulting in the eigenvector of the first row. Feature vectors that fuse attribute information at different times ; It is then sent to the subsequent Transformer layer, where it continues to perform deep feature modeling using the standard self-attention mechanism.

[0238] In the practical application of this invention, the timbre replication model updates the network parameters by optimizing the quantization index prediction loss:

[0239] set up For the predicted audio quantization index sequence, For a real audio quantization index sequence, the prediction loss is defined as:

[0240] ;

[0241] in, This indicates that the model, given historical predictions and semantic feature vector sequences, and attribute feature vector sequence Below, it can correctly predict the first The conditional probability of each quantized index is used, and the loss is implemented through cross-entropy. Model parameters are updated through backpropagation.

[0242] S106. Generate the translated video based on the target human voice audio segments corresponding to all target subtitle texts.

[0243] In the practical application of this invention, all target subtitle texts and corresponding target human voice audio segments and target subtitle texts can be aligned with the timeline of the subtitles in the video to be translated; dynamic range compression and loudness normalization processing are performed on the audio track of the target human voice audio segments to obtain the processed target human voice audio segments; the processed target human voice audio segments are mixed and synthesized with the original ambient audio track to obtain the translated video.

[0244] Optionally, in another embodiment of the present invention, if the actual speech rate of the target human voice audio segment corresponding to the target subtitle text exceeds the speech rate threshold, an implementation manner of the video translation method further includes:

[0245] Input the target subtitle text and the context data of the target subtitle text into the subtitle simplification model, and output the simplified target subtitle text.

[0246] Among them, the calculation method of the actual speech rate can be: Given the actual duration T (unit: second) of the target human voice audio segment corresponding to the target subtitle text, first convert the target human voice audio segment corresponding to the target subtitle text into a phoneme sequence through a phoneme converter, count the total number of phonemes N, and calculate the actual speech rate as the ratio of the number of phonemes to the audio duration, that is, N / T (phonemes / second); Set the speech rate threshold (10 phonemes / second) and calculate the speech rate multiplier.

[0247] For example: The dubbing text "你好" is converted into a phoneme sequence "ni3 hao3" with a total of 4 phonemes, and the audio duration is 0.3 seconds. 4 / 0.3≈13.33 phonemes / second, and the standard speech rate is 10 phonemes / second. Then the speech rate multiplier K = 13.33 / 10 = 1.33 times the speed. When it exceeds the preset threshold, the optimization process is triggered.

[0248] Among them, the subtitle simplification model includes a semantic encoder, a constraint encoder, and a decoder.

[0249] The semantic encoder generates a semantic representation based on the target subtitle text and the context data of the target subtitle text.

[0250] The constraint encoder generates a constraint feature representation based on the pseudo-target subtitle text and the simplification ratio.

[0251] Among them, the pseudo-target subtitle text is obtained by translating the target subtitle text, and the simplification ratio is the length ratio of the target subtitle text to the pseudo-target subtitle text.

[0252] The decoder generates the simplified target subtitle text based on the semantic representation and the constraint feature representation.

[0253] In the actual application process of the present invention, the construction process of the subtitle simplification model includes the following steps (steps C1 to step C3):

[0254] Step C1 - Data preparation:

[0255] First-stage data: Construction of basic parallel corpus. Collect multilingual subtitle pairs from the platform's stock data to construct basic training samples. The basic training samples include: a single original subtitle + the context of the original subtitle (3 - 5 subtitles before and after) + the target language subtitle (label);

[0256] Second-stage data: Generation of course learning data, specifically including:

[0257] Pseudo-subtitle generation: A machine translation model is used to translate a single original subtitle to obtain pseudo-target subtitles;

[0258] Difficulty filtering: Only samples with pseudo-target subtitles longer than the original target subtitles are retained (constructed samples);

[0259] Simplification ratio calculation: Calculate the length ratio of the original target subtitle to the pseudo target subtitle as an indicator of simplification difficulty;

[0260] Course Levels: The data is divided into four difficulty levels based on the simplification ratio:

[0261] Level 1 (Beginner): Simplified ratio 0.8-1.0 (only slight simplification is needed);

[0262] Level 2 (Basic): Simplified by 0.6-0.8 (moderate simplification);

[0263] Level 3 (Advanced): Simplification ratio 0.4-0.6 (deep simplification);

[0264] Level 4 (Expert Level): Simplification ratio 0.2-0.4 (extreme simplification);

[0265] Step C2 - Model Architecture (Seq2Seq) Selection: A dual encoder-single decoder architecture can be adopted, but is not limited to, which includes a semantic encoder, a constraint encoder, and a decoder. The semantic encoder is used to process the original subtitles and context, capturing complete semantic information. The constraint encoder is used to process pseudo-target subtitles and simplification ratio, learning word count constraints. The decoder can use LSTM to generate the simplified target subtitles.

[0266] The training samples during the training process can be: [original subtitles, context, pseudo-target subtitles, simplified ratio]; the training sample labels are the target language subtitles.

[0267] Step C3 - Loss function design: Multi-objective balanced loss can be used, but is not limited to. No restrictions are imposed here.

[0268] Wherein, total loss = semantic loss + length loss, semantic loss A semantic similarity loss based on BERTScore can be used to ensure that the semantics remain unchanged after simplification. The calculation formula is as follows:

[0269] ;

[0270] ;

[0271] ;

[0272] ;

[0273] BERTScore calculates the semantic similarity between two text segments using a pre-trained BERT model, representing... Target subtitles This refers to the subtitles output by the subtitle simplification model. This represents the character sequence corresponding to the target subtitle. This indicates the character sequence corresponding to the output subtitle; Precision is used to... For each character in the text, calculate its position. Cosine similarity of BERT vectors of the most similar characters; Recall pairs For each character in the text, calculate its position. The cosine similarity of BERT vectors of the most similar characters.

[0274] Length loss The number of characters can be precisely controlled by calculating the absolute error between the shortened subtitle length and the target length.

[0275] ;

[0276] in, For character statistics functions, To simplify the proportions.

[0277] In the practical application of this invention, the training process of the simplified model can adopt, but is not limited to, a progressive difficulty increase mechanism, which is not limited here.

[0278] For example, training is divided into four stages. In the first stage (basic ability development (Level 1 data)), only introductory samples are used for training to allow the model to master basic semantic preservation capabilities. The loss function weights are: semantic preservation weight 90%, length control weight 10%.

[0279] In the second stage (capability expansion (Level 1+2 mixed data)), basic-level samples are added, gradually increasing the difficulty of simplification; loss function weights: semantic preservation weight 70%, length control weight 30%;

[0280] In the third stage (deeply simplified training (Level 1+2+3 mixed data)), advanced samples are added. The model needs to learn to maintain the core semantics under larger compression. The weights of the loss function are adjusted as follows: semantics are maintained at 40%, and length is controlled at 60%.

[0281] In the fourth stage (extreme simplification reinforcement (full mixed data)), expert-level samples are added to train the model to handle the most difficult simplification scenarios. The loss function weights are: semantics preserved at 10%, and length controlled at 90%.

[0282] In the actual application of this invention, the simplified target subtitle text can be regenerated with dubbing and the speech rate can be detected again. When the dubbing speech rate falls within a reasonable range or the number of simplification iterations reaches the preset upper limit, the process terminates; otherwise, the system will repeatedly execute the above optimization steps.

[0283] Of course, to further verify the quality of the target subtitle text, manual verification can also be performed, focusing on subjective quality dimensions such as the consistency of the character's voice and the accuracy of emotional expression. Any issues discovered can be fine-tuned by adjusting the dubbing parameters.

[0284] As can be seen from the above solution, the present invention provides a video translation method. Based on a set of subtitles in the target language and the audio file of the video to be translated, it determines the human voice audio segment corresponding to each target subtitle text in the target language subtitle set. Then, it determines the emotional category of the human voice audio segment corresponding to the target subtitle text and the long audio of the character to whom the target subtitle text belongs. Afterwards, based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to whom the target subtitle text belongs, it generates the target human voice audio segment corresponding to the target subtitle text. Finally, based on all the target human voice audio segments corresponding to the target subtitle text, it generates the translated video, significantly improving the quality and usability of automated dubbing and effectively enhancing the viewing experience.

[0285] Another embodiment of the present invention provides a video translation apparatus, such as... Figure 7 As shown, it specifically includes:

[0286] Acquisition unit 701 is used to acquire the set of subtitles for the target language.

[0287] The human voice audio segment determination unit 702 is used to determine the human voice audio segment corresponding to each target subtitle text in the target language subtitle set based on the target language subtitle set and the audio file of the video to be translated.

[0288] The emotion category determination unit 703 is used to determine the emotion category of the human voice audio segment corresponding to the target subtitle text based on the target subtitle text and the corresponding human voice audio segment.

[0289] The character long audio determination unit 704 is used to determine the long audio of the character to which the target subtitle text belongs based on the human voice audio segment corresponding to the target subtitle text.

[0290] The target human voice audio segment unit 705 is used to generate a target human voice audio segment corresponding to the target subtitle text based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to which the target subtitle text belongs.

[0291] The translated video generation unit 706 is used to generate translated videos based on the target human voice audio segments corresponding to all target subtitle texts.

[0292] For details on the specific operation of the units disclosed in the above embodiments of the present invention, please refer to the corresponding method embodiments, such as... Figure 1 As shown, it will not be elaborated further here.

[0293] As can be seen from the above solution, the present invention provides a video translation device. Based on a set of subtitles in the target language and the audio file of the video to be translated, it determines the human voice audio segment corresponding to each target subtitle text in the target language subtitle set. Then, it determines the emotional category of the human voice audio segment corresponding to the target subtitle text and the long audio of the character to whom the target subtitle text belongs. Afterwards, based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to whom the target subtitle text belongs, it generates the target human voice audio segment corresponding to the target subtitle text. Finally, based on all the target human voice audio segments corresponding to the target subtitle text, it generates the translated video, significantly improving the quality and usability of automated dubbing and effectively enhancing the viewing experience.

[0294] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0295] Another embodiment of the present invention provides an electronic device, comprising:

[0296] One or more processors.

[0297] A storage device on which one or more programs are stored.

[0298] When the one or more programs are executed by the one or more processors, the one or more processors implement the video translation method as described in the above embodiments.

[0299] Another embodiment of the present invention provides a computer storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the video translation method as described in the above embodiments.

[0300] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0301] It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0302] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0303] Another embodiment of the present invention provides a computer program product, which, when executed, is used to perform the above-described video translation method.

[0304] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined in the methods of the embodiments of the present invention.

[0305] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in this invention is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely exemplary forms for implementing the invention.

[0306] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0307] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to the specific combination of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with technical features of the present invention (but not limited to) that have similar functions.

Claims

1. A method for translating video, characterized in that, include: Obtain the set of subtitles for the target language; Based on the subtitle set in the target language and the audio file of the video to be translated, determine the human voice audio segment corresponding to each target subtitle text in the subtitle set in the target language; Based on the target subtitle text and the corresponding human voice audio segment, determine the emotional category of the human voice audio segment corresponding to the target subtitle text; Based on the human voice audio segment corresponding to the target subtitle text, determine the long audio of the character to whom the target subtitle text belongs; Based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to whom the target subtitle text belongs, a target human voice audio segment corresponding to the target subtitle text is generated; Based on the target human voice audio segments corresponding to all the target subtitle texts, a translated video is generated; The step of determining the emotion category of the audio segment corresponding to the target subtitle text based on the target subtitle text and the corresponding human voice audio segment includes: Based on the posterior probability distribution of each emotion category of the target subtitle text and the posterior probability distribution of each emotion category of the corresponding human voice audio segment, the emotion category of the corresponding human voice audio segment is determined.

2. The video translation method according to claim 1, characterized in that, The acquisition of the target language subtitle set includes: The audio file of the video to be translated and the structured plot information are input into the speech-to-subtitle recognition model, and the output is a set of subtitle texts for the video to be translated. The structured plot information includes the overall plot information of the video to be translated and the plot information of the episode to which each audio segment belongs. For each audio segment, the speech-to-subtitle recognition model generates first text information and generates second text information based on the overall plot information of the video to be translated and the plot information of the episode to which the audio segment belongs. Based on the first text information and the second text information, the subtitle text corresponding to the audio segment is generated. For each subtitle text in the set of subtitle texts of the video to be translated, the subtitle text is translated into the target language to obtain the target subtitle text.

3. The video translation method according to claim 1, characterized in that, The method of determining the human voice audio segment corresponding to each target subtitle text in the target language subtitle set and the audio file of the video to be translated, based on the target language subtitle set, includes: Obtain the timeline order of the target subtitle text in the subtitle set of the target language, and use it as a timestamp sequence; Speech activity detection is performed on the audio file of the video to be translated to obtain a speech segment sequence; The time range of each target subtitle segment in the timestamp sequence is adjusted based on the speech segment sequence to obtain an adjusted timestamp sequence; wherein the time range of each target subtitle segment in the adjusted timestamp sequence is aligned with the speech segment in the speech segment sequence.

4. The video translation method according to claim 1, characterized in that, The step of determining the long audio clip of the character to whom the target subtitle text belongs based on the corresponding human voice audio clip includes: The voiceprint features of the audio segment corresponding to the target subtitle text are identified to obtain the target voiceprint features; Clustering is performed on all the target voiceprint features to obtain clustering results; wherein, the clustering results include multiple speakers; each speaker corresponds to at least one target voiceprint feature; Based on the human voice audio segments corresponding to all target voiceprint features of the same speaker, a long audio of the speaker is generated; Based on the speaker's long audio, determine the speaker's acoustic attribute information; Based on the speaker's long audio, the speaker's acoustic attribute information, and the character's timbre profile, the long audio of the character to whom the target subtitle text belongs is determined.

5. The video translation method according to claim 4, characterized in that, The step of determining the long audio of the character to whom the target subtitle text belongs, based on the speaker's long audio, the speaker's acoustic attribute information, and the character's timbre profile, includes: A candidate set is determined based on the speaker's long audio and the long audio of the character in the character timbre file; For each candidate role in the candidate set, the role corresponding to the speaker is determined based on the long audio of the candidate role, the acoustic role attribute information of the candidate role, the long audio of the speaker, and the acoustic attribute information of the speaker. Based on the correspondence between the target subtitle text and the speaker's long audio, and the role corresponding to the speaker, the long audio of the role to which the target subtitle text belongs is determined.

6. The video translation method according to claim 5, characterized in that, For each candidate role in the candidate set, determining the role corresponding to the speaker based on the candidate role's long audio, the candidate role's acoustic role attribute information, the speaker's long audio, and the speaker's acoustic attribute information includes: Based on the long audio of the candidate role and the long audio of the speaker, the voiceprint channel score is determined; The style channel score is determined based on the acoustic attribute information of the candidate role and the acoustic attribute information of the speaker; Based on the voiceprint channel score and the style channel score, the dual-channel matching score between the candidate role and the speaker is determined; wherein, the candidate role with the largest dual-channel matching score is taken as the role corresponding to the speaker.

7. The video translation method according to claim 1, characterized in that, The process of generating the target voice audio segment corresponding to the target subtitle text based on the target subtitle text, the duration of the target subtitle text, the emotional category of the corresponding human voice audio segment, and the long audio of the character to whom the target subtitle text belongs includes: The target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to which the target subtitle text belongs are input into the timbre replication model to obtain the target human voice audio segment corresponding to the target subtitle text.

8. The video translation method according to claim 7, characterized in that, The timbre replication model includes: an audio quantization module, a text encoding module, a semantic encoding module, an attribute extraction module, and an audio dequantization module; The audio quantization module quantizes the long audio of the character to whom the target subtitle text belongs into a discrete digital index sequence; The text encoding module encodes the target subtitle text into a text feature vector sequence; The attribute extraction module performs joint encoding based on the duration of the target subtitle text, the emotion category of the human voice audio segment corresponding to the target subtitle text, and the voiceprint features of the long audio of the character to which the target subtitle text belongs, to generate an attribute condition vector set. The semantic encoding module processes the semantic feature vector sequence and the attribute condition vector set to obtain the audio quantization index sequence corresponding to the target subtitle text; wherein, the semantic feature vector sequence is obtained by temporally concatenating the discrete digital index sequence and the text feature vector sequence. The audio dequantization module generates the target human voice audio segment corresponding to the target subtitle text based on the audio quantization index sequence.

9. The video translation method according to claim 1, characterized in that, After generating the target voice audio segment corresponding to the target subtitle text based on the target subtitle text, the duration of the target subtitle text, the emotional category of the corresponding human voice audio segment, and the long audio of the character to whom the target subtitle text belongs, the process further includes: If the actual speech rate of the target human voice audio segment corresponding to the target subtitle text exceeds the speech rate threshold, then the target subtitle text and its context data are input into the subtitle simplification model, and the simplified target subtitle text is output.

10. The video translation method according to claim 9, characterized in that, The subtitle simplification model includes a semantic encoder, a constraint encoder, and a decoder; The semantic encoder generates a semantic representation based on the target subtitle text and its contextual data; The constraint encoder generates constraint feature representations based on pseudo-target subtitle text and simplification ratio; wherein, the pseudo-target subtitle text is obtained by translating the target subtitle text, and the simplification ratio is the length ratio of the target subtitle text to the pseudo-target subtitle text; The decoder generates a simplified target subtitle text based on the semantic representation and the constraint feature representation.

11. The video translation method according to claim 1, characterized in that, The process of generating a translated video based on the target human voice audio segments corresponding to all the target subtitle texts includes: Align all the target audio segments corresponding to the target subtitle text and the target subtitle text according to the timeline of the subtitles in the video to be translated; The target human voice audio track is subjected to dynamic range compression and loudness normalization to obtain the processed target human voice audio track. The processed target human voice audio segment is mixed and synthesized with the original ambient audio track to obtain the translated video.

12. A video translation device, characterized in that, include: The acquisition unit is used to acquire a set of subtitles in the target language; The human voice audio segment determination unit is used to determine the human voice audio segment corresponding to each target subtitle text in the target language subtitle set based on the target language subtitle set and the audio file of the video to be translated. The emotion category determination unit is used to determine the emotion category of the human voice audio segment corresponding to the target subtitle text based on the target subtitle text and the human voice audio segment corresponding to the target subtitle text; The character long audio determination unit is used to determine the long audio of the character to which the target subtitle text belongs based on the human voice audio segment corresponding to the target subtitle text; The target human voice audio segment unit is used to generate the target human voice audio segment corresponding to the target subtitle text based on the target subtitle text, the duration of the target subtitle text, the emotional category of the human voice audio segment corresponding to the target subtitle text, and the long audio of the character to which the target subtitle text belongs; The translated video generation unit is used to generate a translated video based on the target human voice audio segments corresponding to all the target subtitle texts; The translation device is further configured to determine the emotion category of the human voice audio segment corresponding to the target subtitle text based on the posterior probability distribution of each emotion category of the target subtitle text and the posterior probability distribution of each emotion category of the human voice audio segment corresponding to the target subtitle text.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment, storage medium and program product

    CN120199228A

  • Video dubbing language conversion method and system and related equipment

    CN120932629A