Multi-language speech translation method

Through the multilingual acoustic model, the allocation of language tags and dynamic switching orthography rules at the phoneme level is solved, and the recognition and translation errors of the existing speech translation system during in-word language switching is achieved, and high-accurate translation in a hybrid language environment is achieved, suitable for embedded devices and cloud servers.

CN120509419AInactive Publication Date: 2025-08-19ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510429265.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing pronunciation translation systems frequently recognize and translate errors when dealing with intra-word language switching, especially in a mixed language environment, it is difficult to accurately divide mixed words and dynamically adapt to translation needs.

Method used

Multilingual acoustic model is used to assign language tags at the phoneme level, combine end-to-end neural network architecture, dynamically switch language orthography rules, and optimize language tags through attention mechanisms and conditional random fields, and use multilingual neural machine translation model for transfer learning and data augmentation to ensure the correct transcription and translation of mixed words.

Benefits of technology

Improves translation accuracy in hybrid language environments, supports real-time voice translation and is suitable for embedded devices and cloud servers.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a multi-language speech translation method, which relates to the technical field of speech recognition and machine translation, and comprises the following steps of: receiving a speech signal containing intra-word code switching; performing phoneme-level identification on the voice signal through a multi-language acoustic model, and allocating a language tag to each phoneme; grouping the phonemes according to the language tags, and performing transcription according to a normal character rule of a corresponding language; segmenting and translating the transcribed text into a target language according to the recognized source language; the translated texts are combined, final output is generated, and the multi-language acoustic model adopts an end-to-end neural network architecture. According to the multi-language speech translation method, a multi-language acoustic model is adopted, language tags are distributed at a phoneme level instead of a conventional word or sentence level, a normal character rule that the language tags dynamically switch different languages is adopted, correct transcription of mixed words is ensured, a transcribed text is translated in a segmented mode according to the languages, and then the translated text is combined into final output. And translation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field , ,

[0009] ,

[0008]

[0001] The present invention relates to the technical field of speech recognition and machine translation, and specifically to a multilingual speech translation method. Background Art

[0002] In the context of globalization, multilingual mixed communication has become increasingly common. For example, English + Spanish (such as "breakfastar"), Chinese + English (such as "你好hello"), etc. Traditional speech translation systems usually assume that the input speech is in a single language, resulting in recognition and translation errors when dealing with in-word code-switching. For example: Existing ASR (Automatic Speech Recognition) systems are difficult to accurately segment mixed words (the English and Spanish parts in "comput-er-ita"); Existing MT (Machine Translation) systems usually rely on single-language input and cannot dynamically adapt to the translation requirements of mixed-language texts.

[0003] Existing technologies mainly focus on sentence-level code-switching (such as alternating between complete sentences in different languages), but lack support for in-word mixing (such as language switching within a word). Summary of the Invention

[0004] The purpose of the present invention is to provide a multilingual speech translation method to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A multilingual speech translation method, including the following steps: Receive a speech signal containing in-word code-switching; Perform phoneme-level recognition on the speech signal through a multilingual acoustic model and assign a language label to each phoneme; Group the phonemes according to the language labels and transcribe them according to the orthographic rules of the corresponding languages; Segment the transcribed text by the recognized source language and translate it into the target language; Combine the translated text to generate the final output.

[0006] Furthermore, the multilingual acoustic model adopts an end-to-end neural network architecture and uses speech data containing in-word code-switching during training.

[0007] Furthermore, the assignment of language labels is based on the acoustic features of phonemes and is optimized using an attention mechanism or a conditional random field.

[0008] Furthermore, in the transcription module, for mixed-language words, a method of dynamically switching language rules is adopted to ensure the correct transcription of different language parts.

[0009] Furthermore, the machine translation module adopts a multilingual neural machine translation model to support transfer learning or semi-supervised training for low-resource languages.

[0010] Furthermore, for low-resource languages, a pre-trained model of a high-resource language is used for initialization and fine-tuned through an adapter layer.

[0011] Furthermore, in the ASR module, phoneme-language joint modeling is adopted.

[0012] Furthermore, during the training process, data enhancement technology is used to simulate speech samples mixed with different languages.

[0013] Furthermore, in the final output stage, a language model is used to post-process the translation results to ensure semantic coherence.

[0014] Furthermore, the system supports real-time speech translation and can be deployed on embedded devices or cloud servers.

[0015] The present invention provides a multilingual speech translation method with the following beneficial effects: the present invention adopts a multilingual acoustic model and assigns language tags at the phoneme level rather than the traditional word or sentence level. The language tags dynamically switch the orthographic rules of different languages to ensure the correct transcription of mixed words. The transcribed text is translated in segments by language and then recombined into the final output, thereby improving translation accuracy. DETAILED DESCRIPTION

[0016] The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0017] A multilingual speech translation method includes the following modules: Input: Receive speech with intra-word code switching (e.g., English + Spanish mixed “breakfastar”); ASR module: Uses an end-to-end multilingual acoustic model (such as Conformer or Wav2Vec 2.0) to output phoneme sequences and language probability distribution; optimizes language label assignment through an attention mechanism or CRF.

[0018] Transcription module: Group phonemes according to language tags and dynamically switch spelling rules for languages such as English and Spanish.

[0019] For example, "breakfastar" is split into English "breakfast" + Spanish "ar", and transcribed in each language.

[0020] MT module: uses multilingual neural machine translation models (such as mBART) and supports transfer learning; Low-resource languages are fine-tuned with high-resource language models via an adapter layer (e.g., initializing a Spanish model for an indigenous Mexican language).

[0021] Output: The combined translated text, post-processed by a language model to ensure semantic coherence.

[0022] Training and optimization Data augmentation: Synthesize mixed-language speech samples (such as English phonemes + Spanish phonemes) to improve model generalization capabilities.

[0023] Joint modeling: The ASR module uses a joint phoneme-language loss function to optimize language label prediction.

[0024] Take the input voice "comput-er-ita" as an example: ASR recognition: Phoneme sequence: / k / / ɒ / / m / / p / / j / / u / / t / (English) + / e / / r / / i / / t / / a / (Spanish).

[0025] Language tags: The first 7 phonemes are labeled English, and the last 5 are labeled Spanish.

[0026] Transcription: English part → "comput", Spanish part → "erita".

[0027] Translation: English "comput" → target language "computing", Spanish "erita" → target language "small" (assuming the target language is Chinese).

[0028] Output: Combined into "compute small" or adjusted according to context.

[0029] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.

Claims

1. A multilingual speech translation method, characterized in that: The following steps are involved: receiving a speech signal including intra-word code switching; Perform phoneme-level recognition on speech signals through a multilingual acoustic model and assign a language label to each phoneme; Phonemes are grouped according to language labels and transcribed according to the orthographic rules of the corresponding language; Translate the transcribed text into the target language in segments according to the identified source language; The translated texts are combined to generate the final output.

2. A multilingual speech translation method according to claim 1, characterized in that: The multilingual acoustic model uses an end-to-end neural network architecture and is trained on speech data that includes intra-word code switching.

3. A multilingual speech translation method according to claim 2, characterized in that: The assignment of language labels is based on the acoustic features of phonemes and is optimized using attention mechanisms or conditional random fields.

4. A multilingual speech translation method according to claim 3, characterized in that: In the transcription module, for mixed-language words, dynamic switching of language rules is adopted to ensure the correct transcription of different language parts.

5. A multilingual speech translation method according to claim 4, characterized in that: The machine translation module uses a multilingual neural machine translation model and supports transfer learning or semi-supervised training for low-resource languages.

6. The multilingual speech translation method according to claim 5, wherein: For low-resource languages, a pre-trained model from a high-resource language is used for initialization and fine-tuned through an adapter layer.

7. A multilingual speech translation method according to claim 6, characterized in that: In the ASR module, phoneme-language joint modeling is adopted.

8. A multilingual speech translation method according to claim 7, characterized in that: During the training process, data enhancement technology is used to simulate speech samples mixed with different languages.

9. The multilingual speech translation method according to claim 8, wherein: In the final output stage, the language model is used to post-process the translation results to ensure semantic coherence.

10. The multilingual speech translation method according to claim 9, characterized in that: The system supports real-time speech translation and can be deployed on embedded devices or cloud servers.