Multi-language audio inter-translation method and device
By extracting features from the source language audio and calculating the number of syllables, the alignment of audio duration and speech features in multilingual audio translation was achieved, solving the problem of duration mismatch in machine translation and improving the realism and naturalness of the translated audio.
Patent Information
- Application Number
- CN202511842866.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-17
AI Technical Summary
Existing machine multilingual audio translation methods rely on computational resources and model training, making it difficult to achieve natural alignment of audio durations between different languages, resulting in distortion and a strong mechanical feel in the translated audio.
By extracting the content features, timbre features, and emotional prosodic features of the source language audio through speech recognition, calculating the number of syllables, and applying controlled constraints to the target language text based on the number of syllables, a target language audio with the same timbre, emotion, and duration as the source language audio is generated.
The generated target language audio is semantically accurate, matches the duration, and highly reproduces the unique voice characteristics of the speaker in the source language audio, significantly improving the realism and naturalness of the output.
Smart Images

Figure CN121687010A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of audio processing, and particularly relates to a multi-language audio mutual translation method and device. BACKGROUND
[0002] Multi-language audio mutual translation aims to translate source language audio (such as English audio) into target language text sentence by sentence and paragraph by paragraph, and generate corresponding target language audio (such as Chinese audio).
[0003] At present, existing methods mainly include manual multi-language audio mutual translation and machine multi-language audio mutual translation. Manual multi-language mutual translation can deeply understand the semantics and cultural background of source language text, can translate accurate, natural and target language expression habit text, and can also take into account the emotion in the source language audio and the length of the translated audio. This translation method can flexibly handle the complexity and diversity of language, and can better maintain the accuracy of translation. However, the efficiency of manual translation is relatively low, time-consuming and resource-consuming. In the face of large-scale translation tasks, the speed of manual translation is difficult to meet the demand of rapid delivery.
[0004] Machine multi-language audio mutual translation performs well in capturing language context and complex structure, is good at processing long and difficult sentences and complex expressions, can generate quite natural and fluent translation, and effectively guarantees the consistency and coherence of semantics. However, this method needs to rely on powerful computing resources and a large amount of data set for pre-model training and fine-tuning.
[0005] Application Content
[0006] The purpose of the embodiments of the present application is to provide a multi-language audio mutual translation method and device to solve the defects of machine multi-language audio mutual translation method in the prior art which relies on computing resources and model training.
[0007] In order to solve the above technical problems, the present application is implemented as follows:
[0008] In a first aspect, a multi-language audio mutual translation method is provided, comprising the following steps:
[0009] Performing speech recognition on source language audio to obtain source language text, and extracting content features, timbre features and emotional prosody features from the source language audio;
[0010] Performing preprocessing on the source language text, and calculating the number of phonemes of the source language text;
[0011] Controlling and constraining the target language text generation process based on the number of phonemes, and generating target language text consistent with the semantics of the source language text and matching the number of phonemes;
[0012] Based on the target language text, the aforementioned content features, the timbre features, and the emotional prosody features, a target language audio that is consistent with the timbre, emotion, and duration of the source language audio is generated.
[0013] Secondly, a multilingual audio translation device is provided, including:
[0014] The recognition module is used to perform speech recognition on the source language audio to obtain the source language text, and to extract content features, timbre features and emotional prosody features from the source language audio.
[0015] The calculation module is used to preprocess the source language text and calculate the number of syllables in the source language text.
[0016] The first generation module is used to controllably constrain the target language text generation process based on the number of syllables, and generate target language text that is semantically consistent with the source language text and matches the number of syllables.
[0017] The second generation module is used to generate target language audio that is consistent with the timbre, emotion, and duration of the source language audio, based on the target language text, the aforementioned content features, the timbre features, and the emotional prosody features.
[0018] The target language audio generated by the embodiments of this application is not only semantically accurate and matched in duration, but also highly restores the unique voice characteristics of the speaker in the source language audio, thereby significantly improving the realism and naturalness of the output and effectively improving the overall performance and output quality. Attached Figure Description
[0019] Figure 1 This is a flowchart of a multilingual audio translation method provided in an embodiment of this application;
[0020] Figure 2 This is a specific implementation diagram of the multilingual audio translation method provided in the embodiments of this application;
[0021] Figure 3 This is a flowchart of syllable calculation provided in an embodiment of this application;
[0022] Figure 4 This is a flowchart of speech synthesis provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the structure of a multilingual audio translation device provided in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] This application aims to provide a method for multilingual audio translation that considers the speaker's timbre, emotion, and duration. In practical applications, due to differences in speech rates during natural reading of different languages, audio recordings expressing the same meaning often vary in length. This means that while existing methods can achieve accurate translation, they struggle to achieve true parallel alignment. Taking English and Chinese as examples, Chinese uses monosyllabic characters as its basic unit, with each syllable typically carrying independent semantics and exhibiting high information density. The average reading speed is approximately 2 to 3 syllables per second. English, on the other hand, uses more polysyllabic words. Although techniques such as weak forms and linking reduce syllable length to some extent (the average speaking speed is approximately 1.8 to 2.9 syllables per second), expressing the same content usually requires more syllables, resulting in an overall audio length that is often longer than that of Chinese. In everyday spoken text, English audio is typically 15% to 30% longer than its corresponding Chinese counterpart. If the audio is mechanically stretched or compressed in order to forcibly align the duration of the source language audio with the translated target language audio, it will distort the fundamental frequency and formant structure of the speech, severely damaging the naturalness and expressiveness of the speech, resulting in obvious mechanical sound and auditory distortion.
[0026] Specifically, this application proposes a multilingual audio translation method that balances speaker timbre, emotion, and duration consistency. It enables speech conversion between Chinese, English, Japanese, and Korean while maintaining consistency in the original speaker's timbre, emotion, and sentence duration. This method organically integrates multiple modules, including automatic speech recognition, syllable count calculation, controlled translation, timbre and prosodic feature extraction, and speech synthesis, to construct an integrated processing flow, thereby achieving an automated process from input audio to output translated audio. Compared to the information loss and coordination difficulties caused by the independent operation of each module in traditional solutions, this architecture effectively improves overall performance and output quality.
[0027] To achieve consistent audio duration, this application proposes a method for constraining audio duration by the number of syllables. Specifically, during the translation stage, the number of syllables in the source audio text is calculated and used as a condition (e.g., allowing the number of syllables in the translated text to fluctuate within ±10% of the number of syllables in the original text) to constrain the model in generating the target language translation. This method pre-aligns the audio duration during the text generation stage, fundamentally solving the duration mismatch problem caused by differences in speech rates between different languages.
[0028] In the audio synthesis stage, the core of this application's embodiments lies in achieving cross-language transfer of speaker identity features. Specifically, the system extracts the speaker's timbre, emotion, and other features from the source audio and directly injects them into the target language speech synthesis process. This ensures that the final generated target language audio is not only semantically accurate and duration-matched, but also highly reproduces the unique voice characteristics of the speaker in the source language audio, thereby significantly improving the realism and naturalness of the output.
[0029] The following description, in conjunction with the accompanying drawings, details a multilingual audio translation method provided in this application through specific embodiments and application scenarios.
[0030] like Figure 1 The diagram shown is a flowchart of a multilingual audio translation method provided in an embodiment of this application. The method includes the following steps:
[0031] Step 101: Perform speech recognition on the source language audio to obtain the source language text, and extract content features, timbre features, and emotional prosody features from the source language audio.
[0032] Specifically, an end-to-end speech recognition model can be used to generate source language text, and FSQ can be used to generate a content token sequence aligned with the audio frame as content features; a pre-trained CAM++ speaker encoder can be used to extract global timbre embedding vectors from the source language audio as timbre features; and an emotional prosodic encoder can be used to extract frame-level emotional prosodic feature sequences from the source language audio as emotional prosodic features.
[0033] Step 102: Preprocess the source language text and calculate the number of syllables in the source language text.
[0034] Specifically, punctuation marks, stop words, and abnormal characters in the source language text can be standardized, and the number of syllables in the source language text can be counted using linguistic rules or syllable counting tools.
[0035] Step 103: Based on the number of syllables, the target language text generation process is subject to controlled constraints to generate target language text that is semantically consistent with the source language text and has a matching number of syllables.
[0036] Specifically, when generating the target language text, the number of syllables in the target language text can be kept within a preset deviation range by using cue word constraints, template constraints, or language model decoding control strategies, thereby achieving consistency between the target language text and the source language audio duration.
[0037] Step 104: Based on the target language text, the aforementioned content features, the timbre features, and the emotional prosody features, generate a target language audio that is consistent with the timbre, emotion, and duration of the source language audio.
[0038] In this embodiment, before generating target language audio that is consistent with the timbre, emotion, and duration of the source language audio, the speech rate, pitch, and pause positions of the generated target language audio can be adjusted based on the emotional prosodic features, so that the overall expressive style of the target language audio is consistent with that of the source language audio.
[0039] The target language audio generated by the embodiments of this application is not only semantically accurate and matched in duration, but also highly restores the unique voice characteristics of the speaker in the source language audio, thereby significantly improving the realism and naturalness of the output and effectively improving the overall performance and output quality.
[0040] This application aims to establish a multilingual audio translation method that takes into account the consistency of speaker timbre, emotion, and duration. Its core objective is to achieve precise alignment of speaker timbre, emotion, translation, and duration. Specifically, this application mainly focuses on four aspects: speech recognition, syllable calculation, parallel translation, and audio synthesis. Figure 2 As shown.
[0041] In terms of speech recognition, the embodiments of this application mainly use Whisper for speech recognition, and at the same time extract Mel spectrum features for FSQ (Finite Scalar Quantization) to further extract token sequences with phoneme-level acoustic information.
[0042] Compared to traditional speech recognition methods and other deep learning models, Whisper demonstrates significant advantages across several key dimensions. Based on the advanced Transformer architecture, this model eliminates the need for cumbersome feature engineering and multi-stage processing in traditional pipelines. Its robustness to noise is particularly outstanding, effectively suppressing interference from background music and environmental noise, while maintaining high recognition accuracy even under non-ideal audio conditions such as low bitrates and telephone conversations.
[0043] To further extract more detailed acoustic information from the speech signal, this embodiment extracts an additional 128-dimensional Mel-spectral features on top of the Whisper front-end feature extraction and inputs them into the FSQ module for vector quantization. FSQ generates a compact and information-dense token sequence by mapping continuous Mel-spectral features to finite discrete symbolic representations. This sequence can effectively characterize phoneme-level speech attributes, including acoustic details such as stress intensity variations and pitch reflected by the fundamental frequency profile. These token sequences can serve as speech content representations, providing crucial support for subsequent speech synthesis.
[0044] In terms of syllable counting, methods such as character count matching or stretching the synthesized audio signal are commonly used to control the duration of translated audio. While these methods are easy to implement, they all have inherent drawbacks: character or word counting is often inaccurate due to the weak correlation between language units and duration, frequently leading to serious deviations in the prediction of the duration of the translated speech; while directly stretching or compressing the synthesized audio distorts the fundamental frequency and formants of the speech, severely sacrificing the naturalness and expressiveness of the speech, and producing obvious mechanical sounds and distortions.
[0045] In contrast, the embodiments of this application use the number of syllables as the core metric and constraint, which can effectively overcome the above-mentioned shortcomings. As the basic unit carrying the duration of speech, the total number of syllables is highly intrinsically correlated with the total duration of speech. By introducing syllable count constraints during the translation stage, active and precise control over the duration of speech is achieved while generating the translated text. This not only fundamentally avoids severe duration mismatch but also eliminates the need for subsequent destructive processing of the audio waveform. Thus, while ensuring that the duration of the target language audio and the source language audio are parallel, the natural fluency and sound quality fidelity of the synthesized speech are guaranteed to the greatest extent.
[0046] Currently, effective tools for syllable counting include Syllables, bigPhoney, and pyphen. This application uses bigPhoney as the tool for counting syllables, which has significant advantages over Syllables and pyphen: Syllables relies on pure phonetic rules, which can lead to misjudgments of some words, such as misclassifying "one" as two syllables. Pyphen primarily calculates syllables based on affix splitting rules (such as root and suffix division), relying more on word form than actual pronunciation, resulting in limited accuracy. Furthermore, Syllables and pyphen do not support numbers or times appearing in the text, recognizing them as one syllable. BigPhoney, on the other hand, queries the CMU pronunciation dictionary to count the number of vowel phonemes in the phonetic symbols. This method can more accurately handle a large number of pronunciation exceptions and complex derivatives, and the results best match the actual language sense of native speakers.
[0047] likeFigure 3 As shown, text preprocessing is required when calculating syllables. First, all punctuation marks must be completely removed, retaining only letters, numbers, and spaces. Simultaneously, all non-ASCII characters must be strictly filtered out to ensure the text contains only printable English characters and spaces. Furthermore, spaces are normalized by merging consecutive spaces into a single space and removing leading and trailing spaces to create a uniform and standardized text format. Finally, the cleaned text is directly input into the `syllable_count` method of the `bigPhoney` tool, which calculates the number of syllables based on its built-in English pronunciation rule model.
[0048] In parallel translation, the core objective is to establish accurate and equivalent communication bridges between different languages, ensuring that the source language can be completely and unbiasedly transmitted to the target language in terms of semantics, emotion, and logical structure, thereby meeting the fundamental need for "information equivalence" in cross-language communication. Currently, parallel translation can achieve high-quality semantic alignment and content synchronization. However, when expressing the same content in different languages, even when read by the same person, their natural audio durations differ inherently. If only text-level parallelism is considered, it is difficult to achieve audio duration synchronization. To maintain consistency between the source language audio duration and the translated target language audio duration, the current common approach is to translate first and then synthesize speech, that is, translate the source language text into the target language text, and then generate the translated speech through speech synthesis technology. To match the original audio duration, the system often performs simple time-axis stretching or compression on the synthesized speech. This method easily leads to speech distortion and abnormal intonation, seriously affecting the naturalness of hearing and the effectiveness of information transmission.
[0049] To address the aforementioned issues, this application's embodiments primarily constrain the translated text based on the number of syllables, indirectly achieving consistency in audio duration. Regarding model selection, while traditional neural network translation models offer high accuracy, they typically rely on large corpora for pre-training and struggle to achieve natural alignment of target language audio duration with source language audio duration. In contrast, large language models, with their powerful semantic understanding and generative flexibility, can precisely control the output through structured prompts. This application's embodiments use the number of syllables as a key control parameter. When the large language model performs translation tasks, a specific prompting mechanism constrains the total number of syllables in the translated text, making it as close as possible to the number of syllables in the original text. This mechanism, while ensuring translation accuracy, fundamentally solves the audio-video duration alignment problem, laying a solid foundation for generating natural and fluent cross-language speech in the subsequent speech synthesis stage.
[0050] To better achieve parallel translation tasks, this application compared mainstream open-source large language models and ultimately selected Qwen3 as the base model for the parallel translation part. This model supports the processing of multiple languages and dialects and possesses strong cross-language semantic capture capabilities. For example, in Chinese-English bilingual translation scenarios, Qwen3 can perform in-depth semantic analysis of expressions containing specific cultural connotations, such as Chinese idioms and proverbs, and accurately convert them into equivalent forms in the English translation, ensuring the accurate transmission of culturally loaded information. However, although Qwen3 performs excellently in general translation tasks, it still has significant limitations in the specific scenario of parallel translation. Although the model supports four languages—Chinese, English, Japanese, and Korean—there are differences in the balance of translation between the four languages, and the consistency of translation of specialized terminology needs improvement.
[0051] To address the aforementioned shortcomings, this application employs LoRA for targeted model fine-tuning. This application introduces the PAWS-X dataset as the training dataset for model fine-tuning. This dataset covers four target languages: Chinese, English, Japanese, and Korean, and includes various text types such as news reports, literary works, scientific documents, and everyday life scenarios, forming a rich translation corpus. Introducing the PAWS-X dataset into the model training process enhances the model's ability to learn semantic mapping relationships between the four target languages, enabling it to systematically master the expression paradigms and pragmatic habits of different languages. During the pre-training phase, statistical learning and feature extraction of large-scale parallel texts in the PAWS-X dataset continuously optimize the parameter configuration of the cross-lingual semantic understanding and generation modules, providing a high-quality parameter initialization foundation for subsequent fine-tuning.
[0052] The model training process uses the cross-entropy loss of the translation task as the optimization objective and employs the Adam optimizer for parameter updates. In the gradient calculation stage, only the low-rank matrices A and B are solved for gradients, while the original weight matrix of the base model remains frozen. This reduces the size of trainable parameters, effectively controlling computational overhead and mitigating the risk of overfitting on limited data. As training iterates, the low-rank matrices A and B gradually learn the characteristic patterns specific to parallel translation tasks. Taking Chinese-English translation as an example, through continuous optimization, the model gradually masters the precise correspondence between bilingual vocabulary, the transformation rules of syntactic structures, and the adaptation criteria of pragmatic scenarios. This targeted optimization enables the model to significantly improve its professional performance on parallel translation tasks while maintaining its original language understanding capabilities, achieving a synergistic improvement in translation quality and efficiency.
[0053] During the model optimization process, prompt design is a crucial link in ensuring translation quality. The embodiments of this application have formulated a systematic prompt optimization plan: First, clearly define the task objectives. The prompt needs to clearly express the core requirements of the translation task, including specifying the target language and the limit on the length of the translated text. Through an accurate task description, ensure that the model accurately understands the various constraints of the translation task. Regarding the translation requirements, the principle of the highest priority for accuracy is emphasized. In the prompt, it is elaborated in detail that "fully convey the original text information, without omitting key information details, without changing the original logical order, auxiliary words, adjectives, and modal particles can be omitted, and make the generated number of words within the specified range as much as possible. Maintain the original logical structure and semantic relationship. It is prohibited to add factual content or viewpoints that are not in the original text. It is prohibited to simplify or omit key information, and the core sentence meaning cannot be changed". To enable the model to better understand this principle, the embodiments of this application provide a large number of translation comparisons of different types of texts in the examples, allowing the model to master the skills of translation while ensuring accuracy by learning these examples.
[0054] For word count control, it is strictly stipulated in the prompt that "the actual number of words in the translated text (only Chinese characters, letters, and numbers) must fall within the range of 0.9 - 1.1 times the target number of words. If the initial translation exceeds this range, the translated text must be adjusted to meet the requirements. If the number of words after translation is not within the specified range, it needs to be readjusted. If the generated text has far more words than the target number of words, some auxiliary words, adjectives, and modal particles can be appropriately deleted, but the core sentence meaning and key information cannot be deleted". At the same time, to help the model better achieve word count control, specific word count adjustment strategy guidance is provided in the prompt. For example, when additional words are needed (close to the lower limit), prompt the model that "modifiers (adjectives, adverbs), modal particles, and supplementary short sentences that do not affect the core meaning can be preferentially considered for addition, or structural particles ('de', 'di', 'de') and dynamic particles ('zhe', 'le', 'guo') can be appropriately used to make the expression more fluent and natural. Avoid adding substantial new information"; when fewer words are needed (close to the upper limit), inform the model that "redundant expressions should be preferentially deleted, more concise words should be used, and synonymous content should be merged. Some auxiliary particles that do not affect the core semantics or tone can be omitted as appropriate, but ensure that the sentence is smooth and the original meaning is not changed". Through these detailed guidelines, enable the model to flexibly and reasonably adjust the length of the translated text during the translation process to meet the word count requirements.
[0055] Regarding the balance between accuracy and word count, the embodiments of this application provide a comprehensive translation strategy for the model in the prompt. In addition to the specific methods for adding and subtracting words mentioned above, it emphasizes that "all additions and deletions must be based on ensuring that the core meaning remains unchanged, the logic is clear, and the semantics are complete. Avoid making sentences awkward, stiff, or deviating from the original meaning simply to reach the word count." Numerous positive and negative examples are shown in the examples, allowing the model to learn how to skillfully adjust the word count while maintaining accuracy, so that the translation accurately conveys the original information while meeting the word count limit. For example, when translating the sentence "She smiled very happily," if the target word count is close to the lower limit, the model can appropriately add supplementary short phrases, translating it as "She was overjoyed, and a very happy smile bloomed on her face"; if the target word count is close to the upper limit, the model can simplify it to "She smiled happily." These examples help the model master the skill of balance.
[0056] To make the translated text more closely resemble the everyday expressions of the target language, this embodiment of the application pays special attention to adjusting tone and style in the prompt. The model is required to "accurately identify the emotions in the original text and translate them in a way that conforms to the everyday expressions of the target language, maintaining natural fluency, avoiding 'translationese,' and closely resembling the tone of conversation between friends." In the training data, this embodiment of the application includes a large number of texts with different tone styles, such as daily dialogues, movie lines, and social media posts, and provides corresponding translation requirements and examples for each type of text in the prompt, making the translation more engaging and natural, enabling the target language audience to better understand and appreciate the emotions and style of the original text.
[0057] In terms of synthesized audio, this application's embodiments are based on the CosyVoice2 model architecture with optimized design. While retaining its core advantages, it constructs a high-fidelity speech synthesis system that integrates fine-grained control over content, timbre, and emotion. As an end-to-end speech synthesis framework, CosyVoice2 achieves efficient text-to-speech mapping through discrete speech token modeling. Combined with the Flow decoder to generate high-fidelity Mel spectrograms, the overall architecture exhibits excellent performance in terms of synthesis naturalness and speech intelligibility, and eliminates the need for complex multi-stage pipeline design, significantly improving synthesis efficiency. However, the original model remains weak in emotional expression. Its emotion modeling relies on implicit contextual learning, lacking explicit controllability, resulting in synthesized speech exhibiting a single emotional expression and insufficient tension in different contexts. Simultaneously, its speaker representation mainly relies on global embedding, making it difficult to effectively capture individual personalized characteristics in prosody, spectral details, and pronunciation habits, especially in cross-linguistic scenarios, where problems such as intonation distortion easily occur.
[0058] like Figure 4As shown, this embodiment of the application optimizes input representation, model structure, and training mechanism in a coordinated manner. At the input layer, the direct speech recognition stage of this embodiment uses a 128-dimensional content token sequence generated by FSQ and precisely aligned with the source language audio as the core semantic guide for the synthesis process. Simultaneously, two independently extracted conditional signals are introduced: one is a 256-dimensional global timbre embedding obtained through a CAM++ pre-trained speaker encoder, used to stabilize the target speaker's vocal individuality; the other is a local emotion feature sequence output by a lightweight emotion prosodic encoder, which encodes emotion categories (such as joy, anger, sadness, and calm) and their intensity gradients at the frame-level granularity, ensuring that emotion changes are synchronized with speech rhythm. These three types of signals—content tokens, timbre embedding, and emotion sequence—are uniformly input into the improved synthesis network, forming a structured, resolvable, multi-dimensional control space. In the Flow decoder, this application's embodiment designs a "Timbre-Emotion Fusion Module." This module employs a multi-scale convolutional residual structure, fusing information from the language model output, timbre embedding, and emotion sequence. Through Conditional BatchNorm and feature reweighting mechanisms, it dynamically adjusts the local characteristics of the spectrum at each step of Mel spectrum generation. Specifically, this module can finely adjust acoustic dimensions strongly correlated with timbre and emotion, such as formant position, spectral slope, energy envelope, and noise distribution. This ensures that the timbre characteristics of the target speaker are not masked by language differences during cross-language conversion, while the prosodic changes of emotion (such as intonation fluctuations and pause rhythms) are realistically reflected in the spectral details, achieving a synthesis effect that more closely matches the speaker's emotions.
[0059] like Figure 5 The diagram shown is a structural schematic of a multilingual audio translation device provided in an embodiment of this application, comprising:
[0060] The recognition module 510 is used to perform speech recognition on the source language audio to obtain the source language text, and to extract content features, timbre features and emotional prosody features from the source language audio.
[0061] Specifically, the recognition module 510 is used to generate source language text using an end-to-end speech recognition model, and generate a content token sequence aligned with the audio frame as content features using FSQ; extract global timbre embedding vectors from the source language audio using a pre-trained CAM++ speaker encoder as timbre features; and extract frame-level emotional prosodic feature sequences from the source language audio using an emotional prosodic encoder as emotional prosodic features.
[0062] The calculation module 520 is used to preprocess the source language text and calculate the number of syllables in the source language text.
[0063] Specifically, the calculation module 520 is used to standardize punctuation marks, stop words and abnormal characters in the source language text, and to count the number of syllables in the source language text using linguistic rules or syllable calculation tools.
[0064] The first generation module 530 is used to subject the target language text generation process to controlled constraints based on the number of syllables, and generate target language text that is semantically consistent with the source language text and matches the number of syllables.
[0065] Specifically, the first generation module 530 is used to ensure that the number of syllables in the target language text does not exceed a preset deviation range when generating the target language text, through prompt word constraints, template constraints, or language model decoding control strategies, thereby achieving consistency between the target language text and the source language audio duration.
[0066] The second generation module 540 is used to generate a target language audio that is consistent with the timbre, emotion, and duration of the source language audio based on the target language text, the aforementioned content features, the timbre features, and the emotional prosody features.
[0067] Furthermore, the aforementioned device also includes:
[0068] The adjustment module is used to adjust the speech rate, pitch and pause positions of the generated target language audio based on the emotional prosodic features, so that the overall expressive style of the target language audio is consistent with that of the source language audio.
[0069] The target language audio generated by the embodiments of this application is not only semantically accurate and matched in duration, but also highly restores the unique voice characteristics of the speaker in the source language audio, thereby significantly improving the realism and naturalness of the output and effectively improving the overall performance and output quality.
[0070] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described multilingual audio translation method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0071] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0072] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0073] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for multi-lingual audio interpretation, the method comprising: The method comprises the following steps: speech recognition is performed on the source language audio to obtain a source language text, and content features, timbre features, and emotional prosody features are extracted from the source language audio; the source language text is preprocessed, and the number of syllables of the source language text is calculated; a target language text generation process is controlled and constrained based on the number of syllables, so that the target language text is semantically consistent with the source language text and the number of syllables is matched; target language audio that is consistent with the timbre, emotion, and duration of the source language audio is generated according to the target language text, the aforementioned content features, timbre features, and emotional prosody features.
2. The method of claim 1, wherein, The speech recognition performed on the source language audio to obtain a source language text, and the content features, timbre features, and emotional prosody features extracted from the source language audio specifically include: a source language text is generated by using an end-to-end speech recognition model, and a content token sequence aligned with audio frames is generated as content features by using FSQ; a global timbre embedding vector is extracted from the source language audio by using a pre-trained CAM++ speaker encoder as timbre features; a frame-level emotional prosody feature sequence is extracted from the source language audio by using an emotional prosody encoder as emotional prosody features.
3. The method of claim 1, wherein, The preprocessing of the source language text and the calculation of the number of syllables of the source language text specifically include: punctuation, stop words, and abnormal characters in the source language text are standardized, and the number of syllables of the source language text is counted by using linguistic rules or a syllable counting tool.
4. The method of claim 1, wherein, The controlled constraint of the target language text generation process based on the number of syllables, so that the target language text is semantically consistent with the source language text and the number of syllables is matched, specifically includes: When the target language text is generated, the number of syllables of the target language text is controlled to be within a preset deviation range by using a prompt word constraint, a template constraint, or a decoding control strategy of a language model, so that the target language text is consistent with the duration of the source language audio.
5. The method of claim 1, wherein, Before the target language audio that is consistent with the timbre, emotion, and duration of the source language audio is generated, the following step is further included: Based on the emotional prosody features, the speaking rate, pitch, and pause position of the generated target language audio are adjusted, so that the overall expression style of the target language audio is consistent with that of the source language audio.
6. A multi-lingual audio interpretation device, characterized by, The method comprises the following steps: a recognition module is configured to perform speech recognition on source language audio to obtain a source language text, and extract content features, timbre features, and emotional prosody features from the source language audio; a calculation module is configured to preprocess the source language text and calculate the number of syllables of the source language text; a first generation module is configured to control and constrain a target language text generation process based on the number of syllables, so that the target language text is semantically consistent with the source language text and the number of syllables is matched; a second generation module is configured to generate target language audio that is consistent with the timbre, emotion, and duration of the source language audio according to the target language text, the aforementioned content features, timbre features, and emotional prosody features.
7. The apparatus of claim 6, wherein The recognition module is specifically configured to generate source language text by using an end-to-end speech recognition model, and generate a content token sequence aligned with an audio frame as a content feature by using FSQ; and extract a global vocal color embedding vector as a vocal color feature from the source language audio by using a pre-trained CAM++ speaker encoder. A sentiment prosody encoder is used to extract a frame-level sentiment prosody feature sequence as a sentiment prosody feature from the source language audio.
8. The apparatus of claim 6, wherein, The calculation module is specifically configured to normalize punctuation marks, stop words and abnormal characters in the source language text, and count the number of syllables of the source language text by using linguistic rules or a syllable calculation tool.
9. The apparatus of claim 6, wherein, The first generation module is specifically configured to, when generating the target language text, control the number of syllables of the target language text to be within a preset deviation range by using a prompt word constraint, a template constraint or a decoding control strategy of a language model, so as to realize the consistency of the target language text and the source language audio in time length.
10. The apparatus of claim 6, wherein, Further comprising: An adjustment module is configured to adjust the speech rate, pitch and pause position of the generated target language audio based on the sentiment prosody feature, so that the overall expression style of the target language audio is consistent with that of the source language audio.
Citation Information
Patent Citations
Automatic translation using deep learning
CN112307776A
Voice processing method and device, computer equipment and storage medium
CN114220436A
Multi-speaker speech synthesis method and device, storage medium and computer equipment
CN120726991A
Time-length-controllable end-to-end speech translation method and translation system
CN120877707A
Voice content processing method and device, equipment and storage medium
CN120954391A