End-to-end Mongolian-Chinese speech translation method based on comparative learning

By adopting contrasting learning and progressive multi-task training methods in end-to-end speech translation, the translation problem of resource-scarce languages ​​is solved, the translation quality and efficiency are improved, and the generalization ability of the model is enhanced.

CN120183380APending Publication Date: 2025-06-20MINZU UNIVERSITY OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510247278.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing end-to-end speech translation technologies face huge challenges such as scarcity of data, differences in cross-modal feature spaces, and difficulty in model training when dealing with scarce resource languages ​​(such as Mongolian to Chinese).

Method used

Using the end-to-end Mongolian Chinese pronunciation translation method based on contrast learning, a corpus composed of a triple of speech transcription and translation is constructed, a pre-trained machine translation model is introduced, and an asymmetry multi-task strategy is used to alternately execute ASR, MT and ST tasks, and the total loss of fusion comparison learning is designed to optimize the model until convergence.

Benefits of technology

The performance of the model in speech translation tasks is improved, the ability to understand the relationship between speech and text is enhanced, the problems caused by data sparseness are reduced, the generalization ability of the model is improved, and the speech characteristics are better captured through comparative learning methods, improving the translation quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183380A_ABST
    Figure CN120183380A_ABST
Patent Text Reader

Abstract

The invention provides an end-to-end Mongolian-Chinese speech translation method and system based on comparative learning, a storage medium and electronic equipment, and relates to the technical field of speech translation. The embodiment of the invention provides an end-to-end Mongolian-Chinese speech translation method based on comparative learning in order to shorten the distance between Mongolian speech and Chinese text cross-modal data. Firstly, a positive and negative sample pair is constructed based on enhanced data by using a data enhancement method; then, contrast loss is introduced, the similarity between positive sample pairs is maximized, meanwhile, the similarity between negative samples is minimized, and therefore the model is encouraged to learn feature representation with higher discrimination; and finally, fusing the multi-task learning loss and designing the total loss of the model, so that the model can learn more effective speech-text representation mapping, and the accuracy of Mongolian-Chinese speech translation is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech translation, and particularly relates to an end-to-end Mongolian-Chinese speech translation method, system, storage medium and electronic device based on contrastive learning. Background Art

[0002] Speech translation technology aims to directly translate the speech of the source language into the text of the target language, overcome language barriers, and achieve cross-language communication. Traditional speech translation uses a cascaded method, that is, first perform automatic speech recognition (Automatic Speech Recognition, abbreviated as: ASR), and then perform text translation through machine translation (Machine Translation, abbreviated as: MT). However, this method will affect the overall translation effect due to the error propagation problem.

[0003] In recent years, end-to-end speech translation (End-to-End Speech Translation, abbreviated as: E2E ST) has gradually become a research hotspot. Its goal is to directly map the source language speech to the target language text through a unified model, thus avoiding the problem of error propagation. Although the existing end-to-end speech translation technology has made significant progress in some language pairs, it still faces huge challenges in dealing with resource-scarce languages (such as Mongolian to Chinese). Data scarcity, differences in cross-modal feature spaces, and the difficulty of model training are the main obstacles. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] Aiming at the deficiencies of the existing technology, the present invention provides an end-to-end Mongolian-Chinese speech translation method, system, storage medium and electronic device based on contrastive learning, and solves the technical problem of how to deal with resource-scarce languages in the existing end-to-end speech translation technology.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] An end-to-end Mongolian-Chinese speech translation method based on contrastive learning, based on an end-to-end Mongolian-Chinese speech translation baseline model, the model includes a speech encoder, a text embedding layer, a contrastive learning module, and a Transformer encoder-decoder; the method includes:

[0009] Construct a Mongolian-Chinese speech translation corpus composed of a series of speech transcription translation triples to obtain a parallel ASR task dataset, MT task dataset, and ST task dataset;

[0010] Introduce a pre-trained model with a pre-acquired MT parallel dataset, and based on the ASR task dataset, MT task dataset, and ST task dataset, alternately execute the following steps using a progressive multi-task strategy:

[0011] Adopt the pre-trained model with the MT parallel dataset;

[0012] Use the Mongolian speech in the ASR task data as the input of the speech encoder to obtain a first audio feature representation, and through the Transformer codec, generate a first word sequence, and construct an ASR loss to fine-tune the model;

[0013] Use the Mongolian text in the MT task dataset as the input of the text embedding layer respectively to obtain corresponding first text feature representations, and through the Transformer codec, generate a second word sequence, and construct an MT loss to fine-tune the model;

[0014] Use the Mongolian speech in the ST task dataset as the input of the speech encoder to obtain a second audio feature representation, and through the Transformer codec, generate a third word sequence, and construct an ST loss to fine-tune the model;

[0015] Design a total loss that combines contrastive learning loss, and optimize the model until convergence, including the following steps:

[0016] Average the Mongolian speech and its transcribed Mongolian text in the Mongolian-Chinese speech translation corpus according to the time dimension to obtain a first positive sample pair and a first negative sample pair, and construct a first contrastive learning loss based on the contrastive learning module;

[0017] Average the Mongolian speech and its translated Chinese text in the Mongolian-Chinese speech translation corpus according to the time dimension to obtain a second positive sample pair and a second negative sample pair, and construct a second contrastive learning loss based on the contrastive learning module;

[0018] Based on the ASR loss, the MT loss, the ST loss, the first contrastive learning loss, and the second contrastive learning loss, construct a total loss function and optimize the model until convergence;

[0019] Use the Mongolian speech to be translated as the input of the converged model to obtain a translated Chinese word sequence.

[0020] Preferably, before constructing the first contrastive learning loss and the second contrastive learning loss, it further includes:

[0021] Perform audio enhancement on the Mongolian speech in the Mongolian-Chinese speech translation corpus using a span mask augmentation method;

[0022] For the Mongolian text in the Mongolian-Chinese speech translation corpus, text enhancement is performed using the word repetition method;

[0023] And for the Mongolian speech in the Mongolian-Chinese speech translation corpus, sequence and feature dimension enhancement are respectively performed using the sequence cut-off and feature cut-off methods.

[0024] Preferably, the first contrastive learning loss is expressed as:

[0025]

[0026] Where i represents the index; N represents the batch size; log represents the contrastive function; exp represents the exponential function; sim(·) represents the similarity function;

[0027] s i represents the i-th Mongolian speech or the corresponding enhanced speech in any batch, and x i represents the transcribed Mongolian text or the corresponding enhanced text of s i , and (s i , x i ) is a first positive sample pair;

[0028] s j is the j-th Mongolian speech or the corresponding enhanced speech in the same batch except for s i , and x j represents the transcribed Mongolian text or the corresponding enhanced text of s j , and (s j , x j ) is a first negative sample pair;

[0029] f(·) represents the audio feature representation extraction function, and f′(·) represents the text feature representation extraction function; τ represents the temperature parameter.

[0030] Preferably, the second contrastive learning loss is expressed as:

[0031]

[0032] Where y i represents the translated Chinese text of s i , and (s i , y i ) is a second positive sample pair, and (s j , y j ) is a second negative sample pair.

[0033] Preferably, the total loss function is expressed as:

[0034] L = L ASR + L ST + L MT + λ1LCLL1 +λ2L CLL2

[0035] where λ1 and λ2 respectively represent the tuning hyperparameters of the weighted contrast loss term; L ASR 、L ST 、L MT respectively represent the ASR loss, MT loss, and ST loss; and

[0036]

[0037]

[0038]

[0039] where the conditional probability P(A|B) refers to the probability of event A occurring under the condition that another event B has already occurred; x n 、y n 、s n respectively represent the nth Mongolian speech and its transcribed Mongolian text and translated Chinese text in the current batch.

[0040] Preferably, the Adam optimizer is used to minimize the total loss function for model parameter update.

[0041] An end-to-end Mongolian-Chinese speech translation system based on contrastive learning, based on an end-to-end Mongolian-Chinese speech translation baseline model, the model includes a speech encoder, a text embedding layer, a contrastive learning module, and a Transformer encoder-decoder; the system includes:

[0042] A construction module for constructing a Mongolian-Chinese speech translation corpus consisting of a series of speech transcription and translation triples, and obtaining parallel ASR task datasets, MT task datasets, and ST task datasets;

[0043] A fine-tuning module for introducing a pre-trained model of the MT parallel dataset and, based on the ASR task dataset, MT task dataset, and ST task dataset, alternately performing the following steps using a progressive multi-task strategy:

[0044] Using the pre-trained model of the MT parallel dataset;

[0045] Taking the Mongolian speech in the ASR task data as the input of the speech encoder, obtaining a first audio feature representation, and generating a first word sequence through the Transformer encoder-decoder, and constructing an ASR loss to fine-tune the model;

[0046] Use the Mongolian texts in the MT task dataset as the inputs to the text embedding layer respectively, obtain the corresponding first text feature representations, and generate a second word sequence through the Transformer encoder-decoder, and construct an MT loss to fine-tune the model;

[0047] Use the Mongolian speech in the ST task dataset as the input to the speech encoder, obtain the second audio feature representation, and generate a third word sequence through the Transformer encoder-decoder, and construct an ST loss to fine-tune the model;

[0048] An optimization module, which is used to design a total loss that combines the contrastive learning loss, and optimize the model until convergence, including the following steps:

[0049] Average the Mongolian speech and its transcribed Mongolian text in the Mongolian-Chinese speech translation corpus according to the time dimension to obtain a first positive sample pair and a first negative sample pair, and construct a first contrastive learning loss based on the contrastive learning module;

[0050] Average the Mongolian speech and its translated Chinese text in the Mongolian-Chinese speech translation corpus according to the time dimension to obtain a second positive sample pair and a second negative sample pair, and construct a second contrastive learning loss based on the contrastive learning module;

[0051] Based on the ASR loss, the MT loss, the ST loss, the first contrastive learning loss, and the second contrastive learning loss, construct a total loss function and optimize the model until convergence;

[0052] A translation module, which is used to use the Mongolian speech to be translated as the input to the converged model, and obtain the translated Chinese word sequence.

[0053] A storage medium stores a computer program for end-to-end Mongolian-Chinese speech translation based on contrastive learning, wherein the computer program enables a computer to execute the end-to-end Mongolian-Chinese speech translation method described above.

[0054] An electronic device includes:

[0055] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the end-to-end Mongolian-Chinese speech translation method described above.

[0056] (III) Advantageous Effects

[0057] The present invention provides an end-to-end Mongolian-Chinese speech translation method, system, storage medium, and electronic device based on contrastive learning. Compared with the prior art, the following advantageous effects are achieved:

[0058] 1. The progressive multi-task strategy enables the model to simultaneously handle multiple tasks such as ASR tasks, MT tasks, and ST tasks. By jointly optimizing these tasks, the performance of the model in speech translation tasks is improved. During the multi-task learning process, the model can obtain richer semantic and syntactic information from multiple tasks, which helps to improve the model's understanding ability of the relationship between speech and text, thereby improving the translation quality. In addition, in multi-task learning, data and parameters can be shared between different tasks, enabling the model to better utilize the data, reducing the problems caused by data sparsity, and improving the generalization ability of the model.

[0059] 2. Using the contrastive learning method can reduce the gap between modalities, which helps the model to better capture speech features and can improve the translation effect compared to direct end-to-end training. Moreover, on this basis, jointly using an external MT parallel dataset for training enables the model to more comprehensively learn the mapping from speech to text. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0061] Figure 1 It is a block diagram of an end-to-end Mongolian-Chinese speech translation method based on contrastive learning provided by an embodiment of the present invention;

[0062] Figure 2 It is an architecture diagram of an end-to-end Mongolian-Chinese speech translation baseline model provided by an embodiment of the present invention;

[0063] Figure 3 It is a construction flow chart of a Mongolian-Chinese speech translation dataset obtained by assisting with machine translation provided by an embodiment of the present invention;

[0064] Figure 4 It is an example diagram of several speech transcription translation triples provided by an embodiment of the present invention.

[0065] Figure 5 It is a schematic diagram of span mask enhancement provided by an embodiment of the present invention;

[0066] Figure 6 It is a schematic diagram of a word repetition method provided by an embodiment of the present invention;

[0067] Figure 7 It is a schematic diagram of a cut-off strategy provided by an embodiment of the present invention. Detailed Implementation Manner

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] The embodiments of the present application provide an end-to-end Mongolian-Chinese speech translation method, system, storage medium, and electronic device based on contrastive learning, which solve the technical problem of how existing end-to-end speech translation technologies handle resource-scarce languages.

[0070] The technical solutions in the embodiments of the present application to solve the above technical problems are generally as follows;

[0071] The embodiments of the present invention propose a progressive multi-task training strategy to use the knowledge learned from other tasks for the speech translation task. Introducing tasks in stages can enable the model to gradually adapt to the data distributions of different tasks, reduce the risk of overfitting, and improve the generalization ability of the model. Moreover, the Mongolian-Chinese speech translation method based on contrastive learning narrows the semantic distance across modalities and languages, fills the application gap of contrastive learning in low-resource languages in the field of speech translation, and improves the translation quality and efficiency of Mongolian-Chinese speech translation.

[0072] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0073] Embodiment 1:

[0074] As Figure 1 shown, the embodiments of the present invention provide an end-to-end Mongolian-Chinese speech translation method based on contrastive learning. Based on an end-to-end Mongolian-Chinese speech translation baseline model, the model includes a speech encoder, a text embedding layer, a contrastive learning module, and a Transformer encoder-decoder; the method includes:

[0075] S1. Construct a Mongolian-Chinese speech translation corpus consisting of a series of speech transcription-translation triples to obtain parallel ASR task datasets, MT task datasets, and ST task datasets;

[0076] S2. Introduce a pre-trained model of the MT parallel dataset, and based on the ASR task dataset, MT task dataset, and ST task dataset, alternately execute the following steps using a progressive multi-task strategy:

[0077] S21. Use the pre-trained model of the MT parallel dataset;

[0078] S22. Use the Mongolian speech in the ASR task data as the input of the speech encoder to obtain the first audio feature representation, and through the Transformer codec, generate the first word sequence, and construct an ASR loss to fine-tune the model;

[0079] S23. Use the Mongolian texts in the MT task dataset as the input of the text embedding layer respectively to obtain the corresponding first text feature representations, and through the Transformer codec, generate the second word sequence, and construct an MT loss to fine-tune the model;

[0080] S24. Use the Mongolian speech in the ST task dataset as the input of the speech encoder to obtain the second audio feature representation, and through the Transformer codec, generate the third word sequence, and construct an ST loss to fine-tune the model;

[0081] S3. Design a total loss that combines the contrastive learning loss, and optimize the model until convergence, including the following steps:

[0082] S31. Average the Mongolian speech and its transcribed Mongolian text in the Mongolian-Chinese speech translation corpus along the time dimension to obtain the first positive sample pair and the first negative sample pair, and construct the first contrastive learning loss based on the contrastive learning module;

[0083] S32. Average the Mongolian speech and its translated Chinese text in the Mongolian-Chinese speech translation corpus along the time dimension to obtain the second positive sample pair and the second negative sample pair, and construct the second contrastive learning loss based on the contrastive learning module;

[0084] S33. Based on the ASR loss, the MT loss, the ST loss, the first contrastive learning loss, and the second contrastive learning loss, construct a total loss function and optimize the model until convergence;

[0085] S4. Use the Mongolian speech to be translated as the input of the converged model to obtain the translated Chinese word sequence.

[0086] The progressive multi-task training strategy proposed in the embodiments of the present invention uses the knowledge learned from other tasks for the speech translation task. Introducing tasks in stages can enable the model to gradually adapt to the data distributions of different tasks, reduce the risk of overfitting, and improve the generalization ability of the model. Moreover, the Mongolian-Chinese speech translation method based on contrastive learning narrows the semantic distance across modalities and languages, fills the application gap of contrastive learning in low-resource languages in the field of speech translation, and improves the translation quality and efficiency of Mongolian-Chinese speech translation.

[0087] First, it is necessary to introduce the relevant content of the end-to-end Mongolian-Chinese speech translation baseline model:

[0088] In an embodiment of the present invention, by analyzing the characteristics of the Transformer model for sequence task modeling and the good adaptability of contrastive learning for cross-modal speech and text, an end-to-end Mongolian-Chinese speech translation model based on contrastive learning is realized.

[0089] As Figure 2 shown, Figure 2 a schematic diagram of an end-to-end Mongolian-Chinese speech translation baseline model is given.

[0090] The model mainly consists of a speech encoder, a text embedding layer, and a Transformer encoder-decoder. The speech encoder is used to extract the audio features of the speech signal. The text embedding layer is the same as the word embedding in MT, and the two are connected to the Transformer encoder and then passed to the Transformer decoder. For explanation, the Transformer encoder further extracts the high-level semantic hidden representations of the two modalities. The Transformer decoder generates word sequences (transcription and translation) for the ST, MT, and ASR tasks. Since the model has a complete Transformer encoder-decoder as a sub-module, large-scale additional ASR and MT parallel data can be used for pre-training. Secondly, in the end-to-end Mongolian-Chinese speech translation experiment based on contrastive learning, different positive and negative sample construction strategies may have different impacts on the model performance. By analyzing the linguistic characteristics of Mongolian-Chinese bilingual speech, text, etc., a suitable construction strategy is selected to construct positive and negative samples to improve the performance of the translation model. Then the model consists of multiple network modules with different functions. In low-resource scenarios, such a network structure facilitates multi-task learning with joint ASR or MT models. In order to make full use of the existing relevant task data in stock to improve the performance of the model, the embodiment of the present invention proposes a training method combining multi-task learning.

[0091] The speech encoder is used to extract the low-level features of the speech signal. Inside the speech coding module, a multi-layer convolutional network structure with 2 channels of 1024 is adopted. The input is the waveform signal of the original Mongolian speech sampled at 16 kHz. The encoder first processes the original speech signal through a pre-trained model, hides the speech input in the latent space, and effectively captures the local time-frequency features through the convolutional layer, reducing the length of the sequence, thereby reducing the computational complexity of subsequent processing. The stride of each convolutional layer is set to 4. Through the convolutional operation, the time dimension of the input speech signal is reduced to one-fourth of the original length, thereby improving the computational efficiency while retaining enough information for subsequent processing. Therefore, the audio feature representation extracted from the original speech signal s can be represented by a = Speech-Encoder(s), where |a| << |s|.

[0092] The text embedding layer runs in parallel with the speech encoder. Its design and function are the same as those of the word embedding layer in traditional text translation models. It is responsible for converting each word in the input text into a dense vector representation, enabling the model to effectively process text information.

[0093] The features extracted by the convolutional layer of the speech encoder or the text embedding layer are fed into the Transformer encoder. The cross-modal attention mechanism is the core of this model, which is used to establish a direct information exchange channel between speech features and text features, enabling the model to process data inputs of different modalities. Moreover, through the dynamic adjustment of attention weights, the model can more accurately capture the parts in the speech signal that have a greater impact on the translation result. Specifically, the cross-modal attention mechanism uses the features generated by the speech encoder or the text embedding layer as queries, and the embedded representation of the text to be translated as keys and values, and calculates the weighted text features through the self-attention mechanism. This process enables the model to dynamically adjust the focus of attention on the text according to the speech input, thereby more accurately capturing the correspondence between speech and text in ST and ASR tasks or between text and text in MT tasks. Especially when dealing with complex sentence structures and technical terms, it improves the accuracy and naturalness of translation. The Transformer encoder consists of N = 6 identical layers, and each layer has two sub-layers. The first is the multi-head self-attention mechanism, and the second is a simple, position-wise fully connected feed-forward network. Residual connections are used around each of the two sub-layers, followed by layer normalization. That is, the output of each sub-layer is LayerNorm(a + Sublayer(a)), where Sublayer(a) is the function implemented by the sub-layer itself. Here, the Transformer encoder maps the input sequence of symbolic representations (a1,..., a n ) to a series of continuous representations z = (z1,..., z n ). At each step, the model is autoregressive, taking the previously generated symbol as an additional input when generating the next one. To facilitate these residual connections, all sub-layers in the model as well as the embedding layer produce outputs with a dimension of dmodel = 512. The Transformer encoder can capture global context information, making up for the deficiencies of the convolutional layer in dealing with long-range dependencies.

[0094] Given z, the Transformer decoder generates the output sequence y = (y1,..., y |y|) The fused features generated by the cross-modal attention mechanism are used to generate the text in the target language. The Transformer decoder also adopts a multi-layer self-attention and feed-forward network structure, but a masking mechanism is added to the self-attention layer to prevent the model from "peeking" at future words when generating the current word. In addition, the decoder also uses the speech features generated by the encoder for cross-modal attention calculation, which enables the decoding process to make full use of the information of the source language speech, efficiently generate accurate and fluent translation texts, ensure the end-to-end nature of the translation process, and reduce the loss and distortion of information. The Transformer decoder consists of a stack of N = 6 identical layers. In addition to the two sub-layers in each encoder layer, the decoder also inserts a third sub-layer, which performs multi-head attention on the output of the encoder stack. Similar to the encoder, residual connections are used around each sub-layer, followed by layer normalization, and attention is paid to the self-attention sub-layer in the decoder stack to prevent position attention to subsequent positions. This masking, combined with the fact that the output embedding is offset by one position, ensures that the prediction at position i depends only on the known outputs at positions less than i. And a beam search strategy of size 10 is adopted. In each step of expansion, several of the most likely solutions are selected from all the current candidate solutions to continue the expansion until the preset decoding step is reached.

[0095] Meanwhile, to adapt to multi-modal inputs - namely speech and text, the model introduces an additional lookup table for text tokens to achieve the mapping from tokens to embeddings, thus ensuring the flexibility and efficiency of the model when processing different types of inputs. To stabilize the model training process, a pre-layer normalization strategy is adopted. When dealing with the multi-task learning framework, different task indicator tokens (such as [src label]) are introduced to distinguish the three tasks of ST, ASR, and MT and to differentiate audio and text inputs. Whether it is audio or text, the input embedding e is fed into the Transformer encoder for processing. For audio inputs, a specific audio token is added at the front end of the embedding sequence and concatenated with the audio feature embedding e audio ∈R d to form the complete embedding representation of the audio input. This representation contains the acoustic features of the audio and also provides additional context information through the audio token, which helps the model better understand and process the audio input. For text inputs, the model places the language ID symbol at the front end of the sentence to indicate the target language, and the language ID symbol in the decoding stage serves as the initial token for output text generation, guiding the model to predict the text in the target language.

[0096] Since the Transformer-based codec is independent, the embodiments of the present invention can introduce a progressive multi-task speech translation training strategy that combines ASR, MT, and ST tasks, and improve the generalization ability and performance of the model by gradually introducing different tasks. This training strategy is fine-tuned on actual speech translation data, and needs to adapt to the task of translating from speech signals to text. By using the language patterns and translation knowledge learned in the pre-training stage, the model is optimized to better process speech inputs and generate accurate text translations.

[0097] Next, each step of the above solution will be introduced in detail in combination with Figure 2 :

[0098] In step S1, a Mongolian-Chinese speech translation corpus is constructed from a series of speech transcription translation triples, and parallel ASR task datasets, MT task datasets, and ST task datasets are obtained.

[0099] The Mongolian-Chinese speech translation dataset of the embodiments of the present invention is based on the currently known and publicly largest Mongolian speech recognition dataset M2ASR-Mongo. The construction process may include:

[0100] First, Mongolian texts are extracted from the Mongolian speech recognition dataset, and then, using a machine translation system, the Mongolian texts are translated into Chinese. Then, data category screening is performed for the news field, and the data is filtered and aligned and integrated. By means of random sampling, some data is randomly selected from the dataset and submitted to experts for manual review, correction, deletion, and update. Finally, a high-quality Mongolian-Chinese speech translation dataset is obtained. The construction process of obtaining the Mongolian-Chinese speech translation dataset with the assistance of machine translation is as Figure 3 shown.

[0101] Define the speech transcription translation triple D = {(s, x, y)}, where s = (s1,..., s |s| ) represents the input sequence, that is, the original Mongolian audio waveform or the extracted acoustic feature sequence; x = (x1,..., x |x| ) is the corresponding transcribed Mongolian text; y = (y1,..., y |y| ) represents the corresponding translated Chinese text. As Figure 4 shown, Figure 4 several examples of speech transcription translation triples are given.

[0102] It should be noted that studying how to effectively use this auxiliary transcription text for supervised learning and maximizing the use of the available triple-supervised dataset has become the key to improving the performance of the translation model.

[0103] To make full use of these triple-supervised datasets, the embodiments of the present invention further divide them into three parallel supervised sub-datasets, namely the ASR task dataset, the MT task dataset, and the ST task dataset, which are respectively used to solve automatic speech recognition, machine translation tasks, and end-to-end speech translation. This data organization method not only helps the model learn the mapping from audio to source language text, but also promotes the learning of direct translation from audio to target language, and at the same time provides additional training signals for translation from source language text to target language text.

[0104] In addition, considering that external machine translation datasets are usually much larger than the corpora specific to speech translation tasks, the embodiments of the present invention also introduce additional text translation pairs, namely the MT parallel dataset, which greatly enriches the training data and helps the model better understand and learn cross-language conversion rules. This method deepens the model's understanding of the semantic mapping between the source language and the target language, and further improves the model's ability to process speech and text information through cross-modal learning.

[0105] In step S2, a pre-trained model of the MT parallel dataset is introduced, and based on the ASR task dataset, the MT task dataset, and the ST task dataset, the following steps are alternately executed using a progressive multi-task strategy:

[0106] S21. Use the pre-trained model of the MT parallel dataset.

[0107] S22. Use the Mongolian speech in the ASR task data as the input of the speech encoder, obtain the first audio feature representation, and generate the first word sequence through the Transformer encoder-decoder, and construct an ASR loss to fine-tune the model.

[0108] S23. Use the Mongolian text in the MT task dataset as the input of the text embedding layer respectively, obtain the corresponding first text feature representation, and generate the second word sequence through the Transformer encoder-decoder, and construct an MT loss to fine-tune the model.

[0109] S24. Use the Mongolian speech in the ST task dataset as the input of the speech encoder, obtain the second audio feature representation, and generate the third word sequence through the Transformer encoder-decoder, and construct an ST loss to fine-tune the model.

[0110] Exemplarily, define L ASR 、L ST 、L MT to represent the ASR loss, the MT loss, and the ST loss respectively:

[0111]

[0112]

[0113]

[0114] Among them, the conditional probability P(A|B) refers to the probability of event A occurring under the condition that another event B has already occurred; x n , y n , s n respectively represent the nth Mongolian speech and its transcribed Mongolian text and translated Chinese text in the corresponding batch.

[0115] The progressive multi-task training strategy proposed in the embodiments of the present invention: In the initialization stage, the MT parallel dataset is used to pre-train the model to ensure that the model has a strong and effective starting point when processing the original speech signal. In the multi-task training loop stage, by randomly selecting different tasks for training, the model can not only establish knowledge connections between various tasks, but also avoid overfitting to any single task through this dynamic switching mechanism, thereby achieving better generalization performance. In addition, the optimization of the cross-entropy loss enables the model to progress in the direction of reducing prediction errors on each task, further refining and strengthening its capabilities in all aspects.

[0116] In addition, regarding the selection of the task order for alternating optimization, the following supplementary description is provided:

[0117] In the scenario of speech translation, progressive multi-task training usually starts with the ASR task because ASR involves extracting text information from speech signals and is the basis of the ST task. In the ASR task, the speech audio is processed by a speech encoder to generate an audio representation a', which is sent to the Transformer encoder to generate z', and finally processed by the Transformer decoder to generate a sequence x'. Subsequently, the model is guided to learn the MT task, that is, translating from the source language text to the target language text, and this step through the text embedding layer further improves the model's understanding of linguistic features. In the MT task, the source speech text is processed by the text embedding layer to generate a text representation b, which is sent to the Transformer encoder to generate b', and finally processed by the Transformer decoder to generate a sequence y'. Finally, the model will be trained on the ST task to directly translate the speech signal s into the target language text y. Through this phased training method, the model can gradually build the ability to translate from speech to text while maintaining the performance of each individual task.

[0118] In step S3, design the total loss that fuses the contrastive learning loss and optimize the model until convergence.

[0119] In the embodiments of the present invention, considering that external machine translation datasets are usually much larger than the corpora specific to speech translation tasks, introducing these additional text translation pairs can greatly enrich the training data and help the model better understand and learn cross - language conversion rules. This approach deepens the model's understanding of the semantic mapping between the source language and the target language, and through cross - modal learning, further improves the model's ability to process speech and text information.

[0120] By designing the construction method of positive and negative sample pairs, the discrimination ability of the model is improved, and the model's understanding of the complex mapping relationship between speech and text is deepened. By introducing a hard negative sample mining strategy, the training efficiency of the model and the generalization ability of the model can be further improved.

[0121] To improve the performance of the end - to - end Mongolian - Chinese speech translation model, performing data augmentation before the experiment can effectively expand the dataset size and provide more training samples for model training. By introducing data diversity, the model is helped to learn more generalized feature representations and reduce the risk of overfitting.

[0122] Therefore, in this step, before constructing the contrastive learning loss, a data augmentation operation is first performed to complete the construction of positive and negative sample pairs.

[0123] (1) For the Mongolian speech in the Mongolian - Chinese speech translation corpus, perform audio augmentation using the span mask augmentation method; specifically as follows:

[0124] The method of feature extraction and representation of the original audio signal is the key to achieving efficient and accurate translation. The audio features and their representation forms directly affect the model's ability to understand speech signals, and thus affect the translation quality. In speech translation experiments, the original audio waveform sequence s=(s1,..., s |s| ) undergoes pre - processing and feature extraction processes and is converted into a series of high - dimensional feature vectors to capture key acoustic information in the audio signal, such as pitch, rhythm, timbre, etc., while removing noise and irrelevant information. Commonly used audio feature representation methods include Mel - Frequency Cepstral Coefficients, Mel spectrograms, and directly using the converted original waveform, etc. At the same time, through self - supervised learning methods, rich feature representations can be automatically learned from unlabeled audio data, further enhancing the expressiveness of the features and the performance of the model.

[0125] Meanwhile, in order to improve the model's recognition ability for complex speech patterns and enhance its robustness, adopting effective audio data augmentation methods has become an important strategy to improve speech translation performance. The embodiment of the present invention adopts Span-Masked Augmentation, and its core idea is to randomly select consecutive time periods in the audio sequence, set them to silence or replace them with specific mask values, simulating the information loss situation in the audio signal. This enables the model to learn to pay more attention to the unmasked parts of the speech signal, thereby improving the model's sensitivity to key acoustic features and its processing ability for incomplete inputs. First, a part of all time steps in the original audio waveform is randomly selected as the starting index with probability p, and then a subsequence of M consecutive time steps starting from this index is set as the mask to generate a new modified audio sequence s′. The s′ is used as the input of the model, and the contrastive loss on its original corresponding transcript is calculated. By experimentally comparing different configurations, such as the mask probability p and the mask length M, the optimal parameter settings p = 0.25 and M = 3600 can be determined, as Figure 5 shown.

[0126] When the audio s after Span-Masked Augmentation is used as the model input, the goal of the model is to calculate the contrastive loss on the original corresponding transcript. This contrastive loss aims to minimize the difference between the masked audio s′ and its original transcript / target translation text, while maximizing the difference from other non-paired transcripts / non-paired translation texts in the batch. Since the masked audio segment is short, it is assumed that the masked audio and the original transcript / translation text form a positive pair, while the remaining transcripts / translation texts in the same batch are regarded as negative pairs. This method effectively utilizes the positive and negative relationships between the audio and the transcript and between the audio and the translation text, improving the model's ability to distinguish different speech patterns.

[0127] (2) For the Mongolian text in the Mongolian-Chinese speech translation corpus, the text augmentation is performed using the word repetition method; specifically as follows:

[0128] The text data is preprocessed and converted into a form suitable for model processing, which can enable the model to better understand and generate target language text. This includes splitting the text into smaller units (such as words, subwords, or characters) and converting these units into embedding representations in vector form. The text embeddings are obtained through pre-trained embedding models (such as Word2Vec, GloVe, or Transformer-based BERT), and these embeddings can capture semantic information, syntactic structures, and context relationships, providing rich language features for the model. The text data here includes translation texts and transcripts, and both are processed using the same data augmentation and positive and negative sample construction strategies.

[0129] The augmentation of text data can further improve the generalization ability. The text augmentation method of word repetition (WordRepetition) is adopted to augment the text data by randomly duplicating some words (or sub-words) in the original sentence. First, considering that the length of a text sentence is usually shorter than the corresponding audio representation, randomly repeating words in the sentence can increase the text length in a simple and effective way, making it better correspond to the duration of the audio data. At the same time, the repeated vocabulary does not change the original semantics of the sentence. This augmented text is suitable as an additional positive example for the corresponding speech, which helps the model learn a more robust text representation. For a given sentence x, each sub-word token x i can be repeated k times to generate the augmented sentence x'. Here, the selection of k follows the Poisson distribution k ∼ Poisson. This strategy not only increases the length of the sentence but also increases the variability of the text by introducing repeated occurrences of vocabulary, thereby improving the model's adaptability to different text forms. During the model training process, the augmented sentence x' is regarded as an additional positive example corresponding to the original audio s, while other samples in the same batch that have undergone the same word repetition operation are regarded as negative examples. The specific method is as Figure 6 shown.

[0130] (3) For the Mongolian speech in the Mongolian-Chinese speech translation corpus, sequence and feature dimension augmentation are respectively performed using the sequence cut-off and feature cut-off methods; specifically as follows:

[0131] The cut-off strategy is essentially a data augmentation technique. It increases the difficulty of model training by truncating a part of the information in a certain dimension of the input data, thereby improving the model's understanding and generalization ability of the data. It has shown remarkable effects in the field of text processing, especially in improving the model's ability to process context information. In the context of speech representation, the cut-off strategy is analogized to an operation performed on the speech feature matrix. Assuming that the speech feature is represented as a matrix of dimension T×d, where T represents the time dimension and d represents the feature dimension. This strategy operation involves selectively erasing a part of this matrix along each dimension, that is, setting the feature values to 0 in the selected area, aiming to simulate the information loss situation in the speech signal and forcing the model to learn to recover and predict the lost content from the remaining information.

[0132] Sequence cut-off and feature cut-off are adopted. Sequence cut-off focuses on the time dimension. By randomly selecting a continuous region on the time series and setting its values to 0, it simulates the discontinuity in the speech signal to help enhance the model's ability to process missing information in the time series. Feature cut-off is different from sequence cut-off. It targets the feature dimension and is achieved by erasing the selected feature channels in the feature matrix to improve the model's robustness in the face of incomplete feature information. For <original speech, transcribed text>, the cut-off audio representation and the corresponding original transcribed text are used as positive sample pairs, while other samples in the same batch are regarded as negative sample pairs. At the same time, for <original speech, translated text>, the cut-off audio representation and the corresponding target translation text are used as positive sample pairs, and other samples in the same batch are regarded as negative sample pairs to evaluate the effect of the cut-off strategy. This design allows the model to learn more robust representations in the face of partial information being cut off, thus showing better performance and generalization ability in practical applications. The specific method is as Figure 7 shown.

[0133] On this basis, this step designs the total loss that combines the contrastive learning loss and optimizes the model until convergence, which specifically includes the following steps

[0134] S31. Average the Mongolian speech and its transcribed Mongolian text in the Mongolian-Chinese speech translation corpus according to the time dimension to obtain the first positive sample pairs and the first negative sample pairs, and based on the contrastive learning module, construct the first contrastive learning loss using the multi-class N-pair contrastive learning Loss; expressed as:

[0135]

[0136] where, i represents the index; N represents the batch size; log represents the contrast function; exp represents the exponential function; sim(·) represents the similarity function;

[0137] s i represents the i-th Mongolian speech or the corresponding enhanced speech in any batch, x i represents the transcribed Mongolian text or the corresponding enhanced text of s i , and (s i , x i ) is a first positive sample pair;

[0138] s j is the j-th Mongolian speech or the corresponding enhanced speech in the same batch except s i , x j represents the transcribed Mongolian text or the corresponding enhanced text of s j and (sj , x j ) is a first negative sample pair;

[0139] f(·) represents the audio feature representation extraction function, and f′(·) represents the text feature representation extraction function; τ represents the temperature parameter, which is used to control the smoothness of the softmax function and affects the sensitivity of the model to the similarity difference between positive and negative samples.

[0140] S32. Along the time dimension, average the Mongolian speech and its translated Chinese text in the Mongolian-Chinese speech translation corpus to obtain a second positive sample pair and a second negative sample pair, and based on the contrast learning module, also use the multi-class N-pair contrast learning loss to construct a second contrast learning loss; expressed as:

[0141]

[0142] where y i represents the translated Chinese text of s i , and (s i , y i ) is a second positive sample pair, and (s j , y j ) is a second negative sample pair.

[0143] S33. Based on the ASR loss, the MT loss, the ST loss, the first contrast learning loss, and the second contrast learning loss, construct a total loss function and optimize the model until convergence.

[0144] where the total loss function is expressed as:

[0145] L = L ASR + L ST + L MT + λ1L CLL1 + λ2LC LL2

[0146] It is not difficult to find that in step S2, L ASR , L ST , and L MT all use the cross-entropy loss. These loss terms are based on the constructed triple-supervised dataset and aim to directly optimize the performance of the model on their respective tasks. For the end-to-end Mongolian-Chinese speech translation task, the two cross-modal contrast loss terms L CLL1 and L CLL2 specially introduced in this step accelerate the reduction of the distance between speech and text. Among them, L CLL1 is for the pair of Mongolian speech and its corresponding Mongolian transcribed text, while L CLL2Then, focus on the pairs of Mongolian speech and corresponding Chinese translation texts. The design of these two contrast loss terms is based on a core idea: by optimizing the model to make the representations between speech and text transcription and between speech and text translation closer, thereby improving the performance of the model in the cross-modal speech translation task.

[0147] In this loss function, λ1 and λ2 respectively represent the tuning hyperparameters of the weighted contrast loss terms, and their settings directly affect the balance and optimization direction of each task in the multi-task learning. By adjusting these hyperparameters, it can be ensured that while improving the accuracy of end-to-end speech translation, the model can also effectively perform speech recognition and text translation tasks.

[0148] Exemplarily, the embodiments of the present invention use the Adam optimizer to minimize the total loss function to update the model parameters until the model converges.

[0149] In step S4, the Mongolian speech to be translated is used as the input of the converged model, and the translated Chinese word sequence is obtained.

[0150] So far, the embodiments of the present invention have obtained an end-to-end Mongolian-Chinese speech translation baseline model after convergence. Based on this converged model, it can be used for the translation processing of the Mongolian speech to be translated to obtain the translated Chinese word sequence.

[0151] It can be understood that in fact, the end-to-end Mongolian-Chinese speech translation technology based on contrast learning and progressive multi-task training proposed in the embodiments of the present invention is not limited to the Mongolian-Chinese speech translation task. Its core idea and method have wide applicability and can be extended to other language translation tasks and even other technical fields.

[0152] Embodiment 2:

[0153] The embodiments of the present invention provide an end-to-end Mongolian-Chinese speech translation system based on contrast learning. Based on the end-to-end Mongolian-Chinese speech translation baseline model, the model includes a speech encoder, a text embedding layer, a contrast learning module, and a Transformer encoder-decoder; the system includes:

[0154] A construction module, configured to construct a Mongolian-Chinese speech translation corpus composed of a series of speech transcription and translation triples, and obtain parallel ASR task datasets, MT task datasets, and ST task datasets;

[0155] A fine-tuning module, configured to introduce a pre-trained model of the MT parallel dataset, and based on the ASR task dataset, MT task dataset, and ST task dataset, alternately execute the following steps using a progressive multi-task strategy:

[0156] Adopt the pre-trained model of the MT parallel dataset;

[0157] Use the Mongolian speech in the ASR task data as the input of the speech encoder to obtain the first audio feature representation, and generate the first word sequence through the Transformer codec, and construct an ASR loss to fine-tune the model;

[0158] Use the Mongolian text in the MT task dataset as the input of the text embedding layer respectively to obtain the corresponding first text feature representation, and generate the second word sequence through the Transformer codec, and construct an MT loss to fine-tune the model;

[0159] Use the Mongolian speech in the ST task dataset as the input of the speech encoder to obtain the second audio feature representation, and generate the third word sequence through the Transformer codec, and construct an ST loss to fine-tune the model;

[0160] An optimization module for designing a total loss that combines contrastive learning loss and optimizing the model until convergence, including the following steps:

[0161] Average the Mongolian speech and its transcribed Mongolian text in the Mongolian-Chinese speech translation corpus according to the time dimension to obtain the first positive sample pair and the first negative sample pair, and construct the first contrastive learning loss based on the contrastive learning module;

[0162] Average the Mongolian speech and its translated Chinese text in the Mongolian-Chinese speech translation corpus according to the time dimension to obtain the second positive sample pair and the second negative sample pair, and construct the second contrastive learning loss based on the contrastive learning module;

[0163] Based on the ASR loss, the MT loss, the ST loss, the first contrastive learning loss, and the second contrastive learning loss, construct a total loss function and optimize the model until convergence;

[0164] A translation module for using the Mongolian speech to be translated as the input of the converged model to obtain the translated Chinese word sequence.

[0165] Example 3:

[0166] An embodiment of the present invention provides a storage medium storing a computer program for end-to-end Mongolian-Chinese speech translation based on contrastive learning, wherein the computer program causes a computer to execute the end-to-end Mongolian-Chinese speech translation method as described in Example 1.

[0167] Example 4:

[0168] An embodiment of the present invention provides an electronic device, including:

[0169] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including those for executing the end-to-end Mongolian-Chinese speech translation method as described in Embodiment 1.

[0170] It is understandable that the end-to-end Mongolian-Chinese speech translation system, storage medium and electronic device provided by the embodiments of the present invention correspond to the end-to-end Mongolian-Chinese speech translation method provided by the embodiments of the present invention. For the explanations, examples, beneficial effects and other parts of the relevant content, reference can be made to the corresponding parts in the end-to-end Mongolian-Chinese speech translation method, which will not be elaborated here.

[0171] In summary, compared with the prior art, the following beneficial effects are achieved:

[0172] 1. The embodiments of the present invention refer to the research idea of the past speech translation data set, and convert the Mongolian speech recognition data set into a Mongolian-Chinese speech translation data set. After data processing, it is submitted to expert review and inspection. Through the correction and analysis of this data set, a high-quality Mongolian-Chinese speech translation data set is obtained.

[0173] 2. The embodiments of the present invention introduce progressive multi-task training. First, external Mongolian-Chinese machine translation data with a text data volume 10 times that of the M2CST-Mongo data volume is used for model pre-training to obtain rich language information in the target language domain. Subsequently, the model is fine-tuned on the M2CST-Mongo speech translation data set. Multiple tasks share parameters, and different tasks are randomly introduced in stages for training to further adapt to the speech translation task. The details of the progressive multi-task training algorithm are mainly concerned. This algorithm integrates the information of different tasks in the step-by-step training, promoting the overall learning and optimization process of the model.

[0174] 3. In the embodiments of the present invention, the cross-modal contrast learning method narrows the distance between the speech and text modalities. By optimizing the model to make the representations of positive samples (i.e., source language and target language sample pairs with similar semantics) closer, while the representations of negative samples (i.e., sample pairs with unrelated semantics) are farther away, guiding the model to pay more attention to the deep semantic connection between the source language and the target language during the learning process, so as to learn more robust feature representations. Contrast learning is used to capture and strengthen the semantic connection between the source language and the target language, and then reduce the difference between the two languages. The data augmentation of positive and negative sample pairs in the contrast learning method and its corresponding construction strategy are adopted. The formula and principle of the contrast loss function are mainly described, and the method principle of using large-scale external MT parallel data for pre-training to enhance its decoder is introduced.

[0175] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0176] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An end-to-end Mongolian-Chinese speech translation method based on contrastive learning, characterized in that: Based on the end-to-end Mongolian-Chinese speech translation baseline model, the model includes a speech encoder, a text embedding layer, a contrastive learning module, and a Transformer codec; the method includes: Construct a Mongolian-Chinese speech translation corpus consisting of a series of speech transcription and translation triplets, and obtain parallel ASR task datasets, MT task datasets, and ST task datasets; The pre-acquired MT parallel dataset pre-training model is introduced, and based on the ASR task dataset, the MT task dataset and the ST task dataset, the following steps are alternately performed using a progressive multi-task strategy: Using the MT parallel dataset to pre-train the model; Using the Mongolian speech in the ASR task data as input of the speech encoder, obtaining a first audio feature representation, and generating a first word sequence through the Transformer codec, and constructing an ASR loss to fine-tune the model; Using the Mongolian texts in the MT task dataset as inputs of the text embedding layer, respectively, obtaining corresponding first text feature representations, and generating second word sequences through the Transformer codec, and constructing MT loss to fine-tune the model; Using the Mongolian speech in the ST task dataset as the input of the speech encoder, obtaining a second audio feature representation, and generating a third word sequence through the Transformer codec, and constructing an ST loss to fine-tune the model; Design the total loss that integrates contrastive learning loss and optimize the model until convergence, including the following steps: According to the time dimension, the Mongolian speech and the transcribed Mongolian text of the Mongolian-Chinese speech translation corpus are averaged to obtain a first positive sample pair and a first negative sample pair, and a first contrastive learning loss is constructed based on the contrastive learning module; According to the time dimension, the Mongolian speech and the translated Chinese text of the Mongolian-Chinese speech translation corpus are averaged to obtain a second positive sample pair and a second negative sample pair, and a second contrastive learning loss is constructed based on the contrastive learning module; Based on the ASR loss, the MT loss, the ST loss, the first contrastive learning loss, and the second contrastive learning loss, construct a total loss function and optimize the model until convergence; The Mongolian speech to be translated is used as the input of the converged model to obtain the translated Chinese word sequence.

2. The end-to-end Mongolian-Chinese speech translation method as claimed in claim 1, characterized in that: Before constructing the first contrastive learning loss and the second contrastive learning loss, it also includes: For the Mongolian speech of the Mongolian-Chinese speech translation corpus, a span mask enhancement method is used to perform audio enhancement; For the Mongolian text of the Mongolian-Chinese speech translation corpus, a word repetition method is used to perform text enhancement; And for the Mongolian speech of the Mongolian-Chinese speech translation corpus, sequence cutoff and feature cutoff methods are used to perform sequence and feature dimension enhancement respectively.

3. The end-to-end Mongolian-Chinese speech translation method as claimed in claim 2, characterized in that: The first contrastive learning loss is expressed as: Where i represents the index; N represents the batch size; log represents the contrast function; exp represents the exponential function; sim(·) represents the similarity function; s i represents the i-th Mongolian speech or the corresponding enhanced speech in any batch, x i Indicates i The Mongolian text of the transcription or the corresponding enhanced text, (s i , x i ) is a first positive sample pair; s j For the same batch except s i The jth Mongolian speech or the corresponding enhanced speech, x j Indicates j The Mongolian text of the transcription or the corresponding enhanced text, (s j , x j ) is a first negative sample pair; f(·) represents the audio feature representation extraction function, f′(·) represents the text feature representation extraction function; τ represents the temperature parameter.

4. The end-to-end Mongolian-Chinese speech translation method as claimed in claim 3, characterized in that: The second contrastive learning loss is expressed as: Among them, y i Indicates i The translated Chinese text, (s i ,y i ) is a second positive sample pair, (s j ,y j ) is a second negative sample pair.

5. The end-to-end Mongolian-Chinese speech translation method as claimed in claim 4, characterized in that: The total loss function is expressed as: L=L ASR +L ST +L MT +λ1L CLL1 +λ2L CLL2 Among them, λ1 and λ2 represent the tuning hyperparameters of the weighted contrast loss term; L ASR , L ST , L MT denote ASR loss, MT loss and ST loss respectively; and Among them, the conditional probability P(A|B) refers to the probability of event A occurring under the condition that another event B has already occurred; n ,y n 、s n They respectively represent the nth Mongolian speech in the batch and its transcribed Mongolian text and translated Chinese text.

6. The end-to-end Mongolian-Chinese speech translation method as claimed in claim 1, characterized in that: The Adam optimizer is used to minimize the total loss function to update the model parameters.

7. An end-to-end Mongolian-Chinese speech translation system based on contrastive learning, characterized in that: Based on the end-to-end Mongolian-Chinese speech translation baseline model, the model includes a speech encoder, a text embedding layer, a contrastive learning module, and a Transformer codec; the system includes: The construction module is used to construct a Mongolian-Chinese speech translation corpus consisting of a series of speech transcription and translation triplets, and obtain parallel ASR task datasets, MT task datasets, and ST task datasets; The fine-tuning module is used to introduce the pre-acquired MT parallel data set pre-training model, and based on the ASR task data set, the MT task data set and the ST task data set, adopt a progressive multi-task strategy to alternately perform the following steps: Using the MT parallel dataset to pre-train the model; Using the Mongolian speech in the ASR task data as input of the speech encoder, obtaining a first audio feature representation, and generating a first word sequence through the Transformer codec, and constructing an ASR loss to fine-tune the model; Using the Mongolian texts in the MT task dataset as inputs of the text embedding layer, respectively, obtaining corresponding first text feature representations, and generating second word sequences through the Transformer codec, and constructing MT loss to fine-tune the model; Using the Mongolian speech in the ST task dataset as the input of the speech encoder, obtaining a second audio feature representation, and generating a third word sequence through the Transformer codec, and constructing an ST loss to fine-tune the model; The optimization module is used to design the total loss that integrates the contrastive learning loss and optimize the model until convergence, including the following steps: According to the time dimension, the Mongolian speech and the transcribed Mongolian text of the Mongolian-Chinese speech translation corpus are averaged to obtain a first positive sample pair and a first negative sample pair, and a first contrastive learning loss is constructed based on the contrastive learning module; According to the time dimension, the Mongolian speech and the translated Chinese text of the Mongolian-Chinese speech translation corpus are averaged to obtain a second positive sample pair and a second negative sample pair, and a second contrastive learning loss is constructed based on the contrastive learning module; Based on the ASR loss, the MT loss, the ST loss, the first contrastive learning loss, and the second contrastive learning loss, construct a total loss function and optimize the model until convergence; The translation module is used to use the Mongolian speech to be translated as the input of the converged model to obtain the translated Chinese word sequence.

8. A storage medium, characterized in that: It stores a computer program for end-to-end Mongolian-Chinese speech translation based on contrastive learning, wherein the computer program enables a computer to execute the end-to-end Mongolian-Chinese speech translation method as described in any one of claims 1 to 6.

9. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include programs for executing the end-to-end Mongolian-Chinese speech translation method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Cross-type intermodal confusion method and device based on adaptive optimal transmission

    CN121768398A