Multi-language speech translation method
Through the multilingual acoustic model, the allocation of language tags and dynamic switching orthography rules at the phoneme level is solved, and the recognition and translation errors of the existing speech translation system during in-word language switching is achieved, and high-accurate translation in a hybrid language environment is achieved, suitable for embedded devices and cloud servers.
Patent Information
- Application Number
- CN202510429265.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing pronunciation translation systems frequently recognize and translate errors when dealing with intra-word language switching, especially in a mixed language environment, it is difficult to accurately divide mixed words and dynamically adapt to translation needs.
Multilingual acoustic model is used to assign language tags at the phoneme level, combine end-to-end neural network architecture, dynamically switch language orthography rules, and optimize language tags through attention mechanisms and conditional random fields, and use multilingual neural machine translation model for transfer learning and data augmentation to ensure the correct transcription and translation of mixed words.
Improves translation accuracy in hybrid language environments, supports real-time voice translation and is suitable for embedded devices and cloud servers.
Abstract
Description
Technical Field , ,
[0009] ,
[0008]
[0001] The present invention relates to the technical field of speech recognition and machine translation, and specifically to a multilingual speech translation method. Background Art
[0002] In the context of globalization, multilingual mixed communication has become increasingly common. For example, English + Spanish (such as "breakfastar"), Chinese + English (such as "你好hello"), etc. Traditional speech translation systems usually assume that the input speech is in a single language, resulting in recognition and translation errors when dealing with in-word code-switching. For example: Existing ASR (Automatic Speech Recognition) systems are difficult to accurately segment mixed words (the English and Spanish parts in "comput-er-ita"); Existing MT (Machine Translation) systems usually rely on single-language input and cannot dynamically adapt to the translation requirements of mixed-language texts.
[0003] Existing technologies mainly focus on sentence-level code-switching (such as alternating between complete sentences in different languages), but lack support for in-word mixing (such as language switching within a word). Summary of the Invention
[0004] The purpose of the present invention is to provide a multilingual speech translation method to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A multilingual speech translation method, including the following steps: Receive a speech signal containing in-word code-switching; Perform phoneme-level recognition on the speech signal through a multilingual acoustic model and assign a language label to each phoneme; Group the phonemes according to the language labels and transcribe them according to the orthographic rules of the corresponding languages; Segment the transcribed text by the recognized source language and translate it into the target language; Combine the translated text to generate the final output.
[0006] Furthermore, the multilingual acoustic model adopts an end-to-end neural network architecture and uses speech data containing in-word code-switching during training.
[0007] Furthermore, the assignment of language labels is based on the acoustic features of phonemes and is optimized using an attention mechanism or a conditional random field.
[0008] Furthermore, in the transcription module, for mixed-language words, a method of dynamically switching language rules is adopted to ensure the correct transcription of different language parts.
[0009] Furthermore, the machine translation module adopts a multilingual neural machine translation model to support transfer learning or semi-supervised training for low-resource languages.
[0010] Furthermore, for low-resource languages, a pre-trained model of a high-resource language is used for initialization and fine-tuned through an adapter layer.
[0011] Furthermore, in the ASR module, phoneme-language joint modeling is adopted.
[0012] Furthermore, during the training process, data enhancement technology is used to simulate speech samples mixed with different languages.
[0013] Furthermore, in the final output stage, a language model is used to post-process the translation results to ensure semantic coherence.
[0014] Furthermore, the system supports real-time speech translation and can be deployed on embedded devices or cloud servers.
[0015] The present invention provides a multilingual speech translation method with the following beneficial effects: the present invention adopts a multilingual acoustic model and assigns language tags at the phoneme level rather than the traditional word or sentence level. The language tags dynamically switch the orthographic rules of different languages to ensure the correct transcription of mixed words. The transcribed text is translated in segments by language and then recombined into the final output, thereby improving translation accuracy. DETAILED DESCRIPTION
[0016] The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0017] A multilingual speech translation method includes the following modules: Input: Receive speech with intra-word code switching (e.g., English + Spanish mixed “breakfastar”); ASR module: Uses an end-to-end multilingual acoustic model (such as Conformer or Wav2Vec 2.0) to output phoneme sequences and language probability distribution; optimizes language label assignment through an attention mechanism or CRF.
[0018] Transcription module: Group phonemes according to language tags and dynamically switch spelling rules for languages such as English and Spanish.
[0019] For example, "breakfastar" is split into English "breakfast" + Spanish "ar", and transcribed in each language.
[0020] MT module: uses multilingual neural machine translation models (such as mBART) and supports transfer learning; Low-resource languages are fine-tuned with high-resource language models via an adapter layer (e.g., initializing a Spanish model for an indigenous Mexican language).
[0021] Output: The combined translated text, post-processed by a language model to ensure semantic coherence.
[0022] Training and optimization Data augmentation: Synthesize mixed-language speech samples (such as English phonemes + Spanish phonemes) to improve model generalization capabilities.
[0023] Joint modeling: The ASR module uses a joint phoneme-language loss function to optimize language label prediction.
[0024] Take the input voice "comput-er-ita" as an example: ASR recognition: Phoneme sequence: / k / / ɒ / / m / / p / / j / / u / / t / (English) + / e / / r / / i / / t / / a / (Spanish).
[0025] Language tags: The first 7 phonemes are labeled English, and the last 5 are labeled Spanish.
[0026] Transcription: English part → "comput", Spanish part → "erita".
[0027] Translation: English "comput" → target language "computing", Spanish "erita" → target language "small" (assuming the target language is Chinese).
[0028] Output: Combined into "compute small" or adjusted according to context.
[0029] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.
Claims
1. A multilingual speech translation method, characterized in that: The following steps are involved: receiving a speech signal including intra-word code switching; Perform phoneme-level recognition on speech signals through a multilingual acoustic model and assign a language label to each phoneme; Phonemes are grouped according to language labels and transcribed according to the orthographic rules of the corresponding language; Translate the transcribed text into the target language in segments according to the identified source language; The translated texts are combined to generate the final output.
2. A multilingual speech translation method according to claim 1, characterized in that: The multilingual acoustic model uses an end-to-end neural network architecture and is trained on speech data that includes intra-word code switching.
3. A multilingual speech translation method according to claim 2, characterized in that: The assignment of language labels is based on the acoustic features of phonemes and is optimized using attention mechanisms or conditional random fields.
4. A multilingual speech translation method according to claim 3, characterized in that: In the transcription module, for mixed-language words, dynamic switching of language rules is adopted to ensure the correct transcription of different language parts.
5. A multilingual speech translation method according to claim 4, characterized in that: The machine translation module uses a multilingual neural machine translation model and supports transfer learning or semi-supervised training for low-resource languages.
6. The multilingual speech translation method according to claim 5, wherein: For low-resource languages, a pre-trained model from a high-resource language is used for initialization and fine-tuned through an adapter layer.
7. A multilingual speech translation method according to claim 6, characterized in that: In the ASR module, phoneme-language joint modeling is adopted.
8. A multilingual speech translation method according to claim 7, characterized in that: During the training process, data enhancement technology is used to simulate speech samples mixed with different languages.
9. The multilingual speech translation method according to claim 8, wherein: In the final output stage, the language model is used to post-process the translation results to ensure semantic coherence.
10. The multilingual speech translation method according to claim 9, characterized in that: The system supports real-time speech translation and can be deployed on embedded devices or cloud servers.