Simultaneous translation from a source language to a target language
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- LEMON INC(GB)
- Filing Date
- 2024-07-18
- Publication Date
- 2026-04-15
AI Technical Summary
Traditional simultaneous speech translation systems suffer from error propagation and latency due to cascaded architectures, failing to deliver high-quality translation that matches human interpreters' effectiveness.
An end-to-end model-based approach using a machine learning model that converts audio segments into embeddings, constructs model inputs with previous segments and translations, and generates translations iteratively, incorporating a knowledge base and multi-modal retrieval to enhance semantic integrity and accuracy.
Improves translation accuracy by considering semantic integrity and domain-specific knowledge, reducing latency and error propagation, and approaching human-level translation quality.
Smart Images

Figure CN2024106277_22012026_PF_FP_ABST
Abstract
Description
SIMULTANEOUS TRANSLATION FROM A SOURCE LANGUAGE TO A TARGET LANGUAGEFIELD
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular to a method, apparatus, device, and computer readable storage medium for speech translation.BACKGROUND
[0002] Oral translation of conversation, statements, questions, etc. involves the translation of words spoken in a source language to words spoken in a target language. Simultaneous speech translation (SiST) is recognized as one of the most challenging tasks in the translation domain. Machine-assisted automatic interpretation has been receiving much attention in the natural language processing (NLP) community. Traditional simultaneous translation approaches usually employ a cascaded system, involving a streaming Automatic Speech Recognition (ASR) model, a punctuation model and a Machine Translation (MT) model. However, such cascaded systems often suffer error propagation and latency from the ASR module. Despite these advancements in both academic SiST models and commercial SiST engines, the translation quality is still far from satisfactory.SUMMARY
[0003] In a first aspect of the present disclosure, a method for speech translation is provided. The method comprises: in response to obtaining a first audio segment in a source language, converting the first audio segment into a first audio embedding; in response to determining that at least one audio segment is obtained before the first audio segment and / or at least one translation segment in a target language corresponding to the at least one audio segment, constructing a first model input based on the first audio feature embedding, at least one audio embedding corresponding to the at least one audio segment, and at least one translation embedding corresponding to the at least one translation segment; and generating a first translation segment in the target language corresponding to the first audio segment based on the first model input using a trained machine learning model.
[0004] In a second aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform acts comprising: in response to obtaining a first audio segment in a source language, converting the first audio segment into a first audio embedding; in response to determining that at least one audio segment is obtained before the first audio segment and / or at least one translation segment in a target language corresponding to the at least one audio segment, constructing a first model input based on the first audio feature embedding, at least one audio embedding corresponding to the at least one audio segment, and at least one translation embedding corresponding to the at least one translation segment; and generating a first translation segment in the target language corresponding to the first audio segment based on the first model input using a trained machine learning model.
[0005] In a third aspect of the present disclosure, a computer readable storage medium is provided. The computer readable storage medium has a computer program stored thereon, when executed by a processor, implementing the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, an apparatus for speech translation is provided. The apparatus comprises: an embedding converting module configured to in response to obtaining a first audio segment in a source language, convert the first audio segment into a first audio embedding; an input constructing module configured to in response to determining that at least one audio segment is obtained before the first audio segment and / or at least one translation segment in a target language corresponding to the at least one audio segment, constructing a first model input based on the first audio feature embedding, at least one audio embedding corresponding to the at least one audio segment, and at least one translation embedding corresponding to the at least one translation segment; and a translation generating module configured to generate a first translation segment in the target language corresponding to the first audio segment based on the first model input using a trained machine learning model.
[0007] It should be understood that what is described in this Summary is not intended to define key features or important features of the embodiments of the disclosure, nor is it intended to limit the scope of the disclosure. Other features of the present disclosure will become readily apparent from the description below.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other features, advantages and aspects of various embodiments of the present disclosure will become more apparent with reference to the following detailed description taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numeral represents the same or similar elements, where:
[0009] FIG. 1 illustrates a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0010] FIG. 2 illustrates a schematic diagram of architecture of model-based speech translation according to some embodiments of the present disclosure;
[0011] FIGS. 3A-3C illustrate examples of the machine learning models according to some embodiments of the present disclosure;
[0012] FIGS. 3D-3I illustrate some cases for comparison the translation results between relevant product and some embodiments of the present disclosure;
[0013] FIG. 4 illustrates a flowchart of a process for speech translation according to some embodiments of the present disclosure;
[0014] FIG. 5 illustrates a schematic structural block diagram of an apparatus for speech translation according to some embodiments of the present disclosure; and
[0015] FIG. 6 illustrates a block diagram of an electronic device that may implement one or more embodiments of the present disclosure.DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it would be appreciated that the present disclosure can be implemented in various forms and should not be interpreted as limited to the embodiments described herein. On the contrary, these embodiments are provided for a more thorough and complete understanding of the present disclosure. It would be appreciated that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0017] In the description of the embodiments of the present disclosure, the term “including” and similar terms should be understood as open-ended inclusion, that is, “including but not limited to” . The term “based on” should be understood as “at least partially based on” . The terms “one embodiment” or “the embodiment” should be understood as “at least one embodiment” . The term “some embodiments” should be understood as “at least some embodiments” . The following may also include other explicit and implicit definitions. As used herein, the term “model” may represent an association between various data. For example, the above correlation relationship can be obtained based on various technical solutions that are currently known and / or will be developed in the future.
[0018] It is to be understood that, before applying the technical solutions disclosed in various implementations of the present disclosure, the user should be informed of the type, scope of use, and use scenario of the personal information involved in the subject matter described herein in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0019] It is to be understood that, before applying the technical solutions disclosed in various implementations of the present disclosure, the user should be informed of the type, scope of use, and use scenario of the personal information involved in the subject matter described herein in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0020] For example, in response to receiving an active request from the user, prompt information is sent to the user to explicitly inform the user that the requested operation would acquire and use the user’s personal information. Therefore, according to the prompt information, the user may decide on his / her own whether to provide the personal information to the software or hardware, such as electronic devices, applications, servers, or storage media that execute operations of the technical solutions of the subject matter described herein.
[0021] As an optional but non-limiting implementation, in response to receiving an active request from the user, the way of sending the prompt information to the user may, for example, include a pop-up window, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry a select control for the user to choose to “agree” or “disagree” to provide the personal information to the electronic device.
[0022] It is to be understood that the above process of notifying and obtaining the user authorization is only illustrative and does not limit the implementations of the present disclosure. Other methods that satisfy relevant laws and regulations are also applicable to the implementations of the present disclosure.
[0023] As used herein, the term “model” may learn the correlation relationship between corresponding inputs and outputs from training data, so that corresponding outputs may be generated for given inputs after training. The generation of the model may be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using a plurality of layers of processing units. Neural networks models are an example of deep learning-based models. Herein, “model” may also be referred to as “machine learning model” , “learning model” , “machine learning network” , or “learning network” , and these terms are used interchangeably herein.
[0024] A “neural network” is a machine learning network based on deep learning. Neural networks are capable of processing inputs and providing corresponding outputs, and typically include an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications often include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence such that the output of the previous layer is provided as the input of the subsequent layer, where the input layer receives the input of the neural network, and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons) , each of which processes input from the previous layer.
[0025] Generally, machine learning may roughly include three stages, namely a training stage, a testing stage and an application stage (also referred to as an inference stage) . In the training stage, a given model may be trained using a large amount of training data, and parameter values are continuously updated iteratively until the model may obtain consistent inferences from the training data that meet the expected goals. Through training, the model may be thought of as being able to learn associations from inputs to outputs (also referred to as input-to-output mappings) from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to test whether the model may provide the correct output, thereby determining the performance of the model. In the application stage, the model may be used to process the actual input and determine the corresponding output based on the parameter values obtained through training.
[0026] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure may be implemented. As shown in FIG. 1, the example environment 100 may include a speech translation system 110 which apply a machine learning model 112 to implement speech translation.
[0027] The speech translation system 110 receive an audio 101 captured in a physical environment, such as in a meeting or from any other data source. The audio 101 contains content in a certain language (referred to as a source language) . The speech translation system 110 translate the audio 101 in the source language by using the machine learning model 112, to obtain a translation result 102 in a target language which may be different from the source language. The translation result 102 may then be presented to the user (s) . The source language and the target language may be any languages, e.g., natural languages.
[0028] The machine learning model 112 may be a model that is suitable for processing audio modality of data. The translation result 102 may be in any suitable modality, e.g., may be a translation text sequence, or a translation audio. The translation text sequence and the translation audio are corresponding with each other in the target language. For example, the machine learning model 112 may output the translation text sequence which is then converted into a translation audio via text-to-speech (TTS) technology, to be playback to the user (s) . The machine learning model 112 may output the translation audio which is then converted to a translation text sequence to be presented to the user (s) via audio speech recognition (ASR) technology.
[0029] In the use case of simultaneous speech translation, the speech translation system 110 is configured to provide a translation result for an audio segment in an audio stream, instead of waiting for all the audio are captured.
[0030] The machine learning model 112 may be, for example, any neural network that may perform feature extraction and feature aggregation, including but not limited to Fully Convolutional Network (FCN) , Convolutional Neural Network (CNN) , Recurrent Neural Network (RNN) , etc., and the embodiments of the present disclosure are not restricted in this regard. In some embodiments, the machine learning model 112 may be stored locally in the speech translation system 110. The speech translation system 110 may directly utilize the local machine learning model 112 to implement feature extraction and feature aggregation when it needs to perform tasks associated with feature extraction and feature aggregation. In some embodiments, the machine learning model 112 may also be a model stored in the cloud. The speech translation system 110 may send the input 101 to the machine learning model 112 in the cloud and obtain the translation result 102 from the machine learning model 112 in the cloud.
[0031] The speech translation system 110 may include or may be implemented in any type of computing-capable device, including a terminal device or a server device. The terminal device may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs) , audio / video player, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, electronic book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices, or any combination thereof. Server devices may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0032] It should be understood that the structure and functionality of environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0033] Traditional simultaneous translation approaches usually employ a cascaded system, involving a streaming Automatic Speech Recognition (ASR) model, a punctuation model and a Machine Translation (MT) model. However, such cascaded systems often suffer error propagation and latency from the ASR module. Despite these advancements in both academic SiST models and commercial SiST engines, the translation quality is still far from satisfactory. However, from the user-centered perspective, these systems cannot deliver enough valid information to listeners, which heavily affects communication effectiveness. In contrast, professional human interpreters usually deliver much more the necessary information. Such a discrepancy implies further improvements for machine-assisted SiST systems to reach human capabilities.
[0034] Motivated by the huge success of large scale language model in machine translation and speech translation, it is proposed to apply language model to accomplish the SiST task. A challenge for incorporating the language model into the SiST task is the read-write policy, where the language model needs to provide partial translation for input speech.
[0035] According to example embodiments of the present disclosure, an improved solution for speech translation is provided. In this solution, an end-to-end model-based approach is proposed to accomplish simultaneous interpretation by iteratively performing multiple actions for audio segments. For a first audio segment in a source language, it is converted into a first audio embedding. A model input is constructed based on the first audio feature embedding, at least one audio embedding corresponding to at least one previous audio segment (if available) , and at least one translation embedding corresponding to the at least one translation segment. The model input is provided to a trained machine learning model, to generate a translation segment in the target language corresponding to the first audio segment. The translation segment and / or the first audio segment may be stored and provided for use in translation of the following audio segments.
[0036] In this way, by taking the previous audio segments and previous translation segments into account, the translation accuracy of the audio segment can be improved, to provide translation with semantic integrity.
[0037] Some example embodiments of the present disclosure will continue to be described below with reference to the accompanying drawings.
[0038] FIG. 2 illustrates a schematic diagram of an architecture 200 for speech translation according to some embodiments of the present disclosure. The architecture 200 may be implemented at the speech translation system 110 in FIG. 1, and various units and / or encoders in the architecture 200 may be implemented by or utilizing the machine learning model 112 of FIG. 1 in the speech translation system 110. For ease of discussion, the architecture 200 will be described with reference to the environment 100 in FIG. 1.
[0039] As shown in FIG. 2, at step 201, the speech translation system 110 receives an audio segment 212 in the environment 210. This step is marked as an <Input> operation, to obtain input to the machine learning model 112.
[0040] In some embodiments, the machine learning model 112 may be constructed based on an encoder-conditioned language model architecture. The audio segment 212 is converted by the audio encoder into an audio embedding (also referred to as audio representation) for processing by the conditioned language model.
[0041] FIG. 3A shows example architecture of the machine learning model 112, including an audio encoder 310 and a conditioned language model 330. The conditioned language model 330 may be based on a large language model. The audio encoder 310 transforms an audio segment to an audio embedding to input to the conditioned language model 330.
[0042] In some embodiments, the audio encoder 310 may contain a large-scale speech conformer pretrained on a large of speech data and an audio adapter to connect to the conditioned language model 330. The audio adapter may down-sample the audio representations and the resulting representations are linearly projected to match the dimension of the embedding layer of the conditioned language model 330. In some examples, the projected audio representations may lower the computational latency for SiST.
[0043] In some embodiments, the language model is a medium size decoder-only transformer to balance performance and computation efficiency. The conditioned language model 330 may be pretrained on a large amount of text data and fine-tuned with instructions. The conditioned language model 330 may directly take the continuous embedding from both the audio encoder and text embedder as input. The conditioned language model 330 may autoregressively generates the translation response of the provided audio segment (and optionally, provide the transcription of the audio segment) .
[0044] It is assumed that the audio segment 212 is obtained for translation in the example of FIG. 3A. If there is no previous audio segment obtained or stored for the current audio segment 212 to be processed, a model input to the conditioned language model 330 may include the audio embedding of the audio segment 212 and an instruction 324, which contains prompt information to indicate a translation from the source language to a target language. As shown in FIG. 3A, the instruction 324 may indicate streaming translation into “English” (which is the target language) . The conditioned language model 330 may then perform the translation of the audio segment 212 based on such instruction. In some examples, the instruction 324 may be configured to further indicate the conditioned language model 330 to perform transcription on the audio segment. For example, the instruction 324 may indicate the conditioned language model 330 to transcribe to Chinese and then translate to English. In this case, in addition to provide a translation result of the input audio in English, the conditioned language model 330 may also provide a Chinese transcription text of the input audio segment in Chinese.
[0045] In some embodiments, the machine learning model 112 is configured to perform streamlining generation. The machine learning model 112 may generate the translation result by detecting whether a currently generated text sequence has a semantic integrity. In oral, people may have modal particles, stammering, repetitions, the like, and sometimes background speech may be inserted into the user who is speaking. In some cases, to reduce the complexity in capturing and uploading decision, audio segments with relatively fixed durations are input for simultaneous translation. In these cases, an audio segment may be start in any time point of an audio stream and end in any time point of the audio stream. These make it inaccurate to directly translate an ASR result of the audio segment in the source language into the target language.
[0046] Therefore, by considering the semantic integrity in speech translation, instead of translating all the content contained in an audio segment, some of the content in the audio segment may not be translated into the target language as they may compromise the semantic integrity of the translation text sequence.
[0047] In some embodiments, as the machine learning model 112 may output the translation result (e.g., the translation text sequence) in a streaming manner, it may sequentially detect a semantic integrity of one or more translation text units in the target language that are translated from the audio segment. If the machine learning model 112 determines a first part of the current audio segment is translated with a semantic integrity, and a second part following the first part may compromise the semantic integrity, the machine learning model 112 may generate only the translation text sequence corresponding to the first part. For example, as shown in FIG. 3B, the part 302 in the audio segment 212 may not help complete the semantic integrity, and thus the conditioned language model 330 may not output the translation text sequence 226 to include translation of the part 302.
[0048] In some embodiments, the machine learning model 112 may detect a semantic integrity of a translation text sequence in the target language that is translated from the audio segment based on the context information, e.g., the one or more previous audio segments and / or the translation result corresponding to the one or more previous audio segments which may be stored in the memory 240.
[0049] A challenge for the machine learning model is to require understanding and translation of terms and uncommon phrases that language models cannot learn from training data. For this challenge, in some embodiments, as shown in FIG. 2, one or more of the following modules are included to augment the machine learning model 112: a knowledge base 230 that stores terms and paired translations, and a memory 240 that stores the context of speech. The knowledge base 230 and the memory 240 may be embedded in the speech translation system 110 or may be external modules.
[0050] The memory 240 stores previous transcriptions (optional) and translations during an interaction session for an audio stream in the environment 210. At round r, it first reads previous audio e.g., one or more audio segments obtained before the current audio segment 212, where tr-1 is the predicted cut-off timestamp of round r-1 and Tr is the end time for the audio stream at round r. In some embodiments, the speech translation system 110 may retrieve relevant information kr from the knowledge base 230 at step 202 (marked as <Retrieve>) . The speech translation system 110 further loads historical context information y1: r-1 from the last round memory at step 203 (marked as <LOAD_MEM>) . Then the model input constructed for the audio segment 212 may be represented as kr, y1: r-1. In some examples, as shown in FIG. 3A, the model input to the conditioned language 330 of the machine learning model 112 may further include context information 326.
[0051] In some embodiments, the cut-off stamp (s) predicted for one or more previous audio segments may also be included in the model input if such cut-off stamp (s) are output by the machine learning model 112 and stored in the memory 240.
[0052] The speech translation system 110 may provide the model input to the machine learning model 112 to generate a translation result, e.g., a translation text sequence, in the target language for the audio segment 212 in the source language. In some embodiments, the machine learning model 112 may determine a cut-off timestamp in the audio segment 212 at round r, which identifies a time in the audio segment 212 before which a complete semantic is detected. Then the machine learning model 112 may generate the translation result corresponding to an audio portion of the audio segment 212 before the cut-off timestamp.
[0053] The output of the machine learning model 112 at round r may be represented as follows:
[0054] where tr is the predicted cut-off timestamp indicating the end time for the current translation round r. yr is then forwarded to update the memory 240. In some embodiments, when instructed to output the transcription, the machine learning model 112 may optionally engage chain of think (CoT) to generate the transcription first and then the translation of the audio segment at round r. For the following round r + 1, the audio stream begins with the predicted cut-off timestamp.
[0055] In some embodiments, the model output yr at each round may at least include a translation segment (e.g., a translation text sequence) in the target language. The translation segment (e.g., a translation text sequence) may then be presented to the user. For example, in FIG. 2, a translation text sequence 226 in English corresponding to the audio segment 212 may be presented to the user in the environment 210 at step 204 (marked as <Output>) . In some embodiments, the model output yr at each round may further include the transcription result of the audio segment in the source language. In this case, the transcription result may be presented to the user. For example, in FIG. 2, a transcription text sequence 224 in Chinese corresponding to the audio segment 212 may be presented to the user in the environment 210.
[0056] In the architecture 200, a new-driven read-write strategy is applied. Without the requirement of complicated human pre-design, the strategy could balance translation quality and latency effortlessly. Unlike most systems where the outputs are frequently rewritten during the translation process for better quality, the strategy proposed herein guarantees all the outputs are deterministic while maintaining high quality.
[0057] In real-world scenarios, the accurate speech transcription or translation of professional and domain-specific terms is challenging. Even human interpreters require prior domain knowledge to understand those terms, including names of people, locations, jargon, or special in-domain terms. For example, an interpreter unfamiliar with the machine learning theory may not recognize the word “Rademacher complexity” when hearing it. Therefore, in various scenarios, human interpreters often prepare in advance to get familiar with the corresponding domain knowledge. Motivated by the preparatory trajectory of human interpreters, in some embodiments, it is proposed to integrate an external database to empower the machine learning model with necessary domain-specific knowledge. Each item in the database contains a key and the corresponding value in text modality. The key, which may appear in the speech, is used as the input for the retriever. The value of the item may be itself, a paired translation of the target language, or even an explanation of the key.
[0058] Theoretically, all items in the external database might be added into the prompt to provide information for the translation. Considering that the knowledge base may contain tremendous terms that not only increase the inference time but may also lower the model performance because of noisy intervention. Simply prompting the machine learning model with all the terms not only increases the inference time but may also hurt the performance of the machine learning model 112 because of noisy intervention. Therefore, in some embodiments, a Multi-Modal Retrieval Augmented Generation (MM-RAG) process may be incorporated in the architecture 200. The speech translation system 110 may apply a multi-modal retriever to extract knowledge from the knowledge base 230 based on the audio segment input. The multi-modal retriever first retrieves the relevant terms from the database based on the input speech. A small number of filtered items are incorporated into the prompt of the machine learning model 112 for in-context learning as shown in FIG. 2.
[0059] When constructing the model input for the audio segment 212, the speech translation system 110 may determine, from the knowledge base 230 (also referred to as a knowledge base) , at least one term matching the audio segment 212. In some embodiments, the speech translation system 110 may apply the multi-modal retriever to determine the matching between the term and the audio segment.
[0060] With the retrieved knowledge and previous context from the memory 240, the machine learning model 112 has the in-context learning ability to better utilize the provided contextual information. To achieve this, a series of in-context learning data may be collected to train the model. Compared with the previous approaches for intervention, such as shallow-fusion and traditional substitution-based methods, which generates fixed translation for given translation pair, the incorporation of the knowledge base can achieve better results and generates more coherent text. For example, in some internet companies, the Chinese characters “大盘” means “overall performance” in English, while in most cases, it should be translated into “stock market” in English. In some embodiments of the present disclosure, the correct translation may be selected by the speech translation system given different context. In addition, in some embodiments, the monolingual text from both source and target languages may be used to help the translation.
[0061] In some embodiments, the knowledge base 230 may include a number of pairs of terms in the source language and its corresponding in the target language, or vice versa. For example, as shown in FIG. 2, a CN-EN pair for “Ising model” , an EN-CN pair for “Ising” may be contained in the knowledge base 230. In some embodiments, the knowledge base 230 may additionally or alternatively contain terms in one of the source language or the target language. For example, in FIG. 2, the knowledge base 230 may contain a Chinese and / or English expression of “Ferromagnetism” . It would be appreciated that the knowledge base 230 may contain more other pairs of terms in two or more different languages, or terms in one language. The scope of the present disclosure is not limited in this regard.
[0062] In some embodiments, the multi-modal retriever may employ audio and text encoders to independently encode the audio segment and text key of the terms in the knowledge base 230. To enhance the alignment between audio embeddings and text embeddings, the multi-modal retriever may include an embedding fusion layer, which includes a multi-head attention module followed by a pooling layer. The resulting pooled representation is subsequently fed into a linear projection layer to produce the final scores, indicating the probability of the text key’s presence in the audio segment. Terms with top scores are forwarded to the machine learning model 112 to enhance the translation quality.
[0063] For example, for the audio segment 212, the speech translation system 110 may employ the multi-modal retriever to retrieve, at step 202, the CN-EN pair for “Ising model” , and / or the EN-CN pair for “Ising” , as the two pairs are determined to match with the audio segment 212. Then the model input to the machine learning model 112 for translation of the audio segment 212 may further include the retrieved terms, e.g., include the text embeddings of the retrieved terms. In some examples, as shown in FIG. 3A and FIG. 3C, the model input to the conditioned language 330 of the machine learning model 112 may further include the context information 326. In some embodiments, in the translation result for the current audio segment, the matched terms in the target language may be included in the translation text sequence or translation audio corresponding to the input audio segment. For example, in FIG. 2, the translation text sequence 226 in English may include the matched terms “Ising model” , “Ising” .
[0064] In some embodiments, the multi-modal retriever may be independently trained with a substantial dataset of speech recognition data. During training, words are randomly selected from speech transcription to serve as the positive sample, indicating their appearance in the speech. In some embodiments, negative words are selected from different sentences, indicating the speech does not mention these words. A label of 1 may be assigned to positive samples and a label of 10 may be assigned to negative samples, aiming to minimize the Binary Cross Entropy (BCE) loss. This approach helps refine the model’s ability to distinguish relevant from irrelevant information, enhancing its overall performance and accuracy. In some embodiments, the positive sample may be labeled as 1 and the negative samples may be labeled as 0, minimizing the Binary Cross Entropy (BCE) loss. To evaluate the effectiveness of the multi-modal retriever, an in-house retrieve development set may be constructed. Each sample in the development set includes a short audio chunk and the mentioned terms in the audio. Note that the term here may be defined as special keywords, such as name, location, abbreviation, and domain-specific word.
[0065] According to the embodiments of the present disclosure, the speech translation system 110 may iteratively translate respective audio segments obtained from the environment 210. For example, after the audio segment 212 is translated, in FIG. 3C, a following audio segment is further obtained, which may be processed in a similar way as the audio segment 212. The following audio segment may be considered as connected to the previous audio segment 212 as an audio segment 314 because the previous audio segment 212 is obtained as context to facilitate the translation. The connected audio segment may be converted by the audio encoder 310 to obtain the corresponding audio embeddings. The model input constructed for this audio segment may further the audio embeddings, and embeddings corresponding to the context information 326, the instruction 324, and the retrieved knowledge 322 from the knowledge base 250 (if available) . The context information, e.g., the translation segment (s) corresponding to the previous audio segment (s) may also be converted into the embedding space as the translation embedding, for processing by the conditioned language model. Based on the model input, the conditioned language model 330 may output a translation segment, e.g., a translation text sequence for the new audio segment. In some embodiments, by considering the semantic integrity of the translation result, the conditioned language model 330 may output partial translation that is not included in the translation text sequence because this partial translation may compromise the semantic integrity for the previous translation result but benefit the semantic integrity of the current audio segment. Similarly, a part of the new audio segment that compromises the semantic integrity of the current audio segment (e.g., the part 304 in FIG. 3C) may not be directly translated into the target language in the translation text sequence.
[0066] In the architecture 200, the memory 240 stores translations and transcriptions in previous rounds y1: r-1. It has at least two functions. Firstly, it works with the input speech to determine which part of the speech has been translated and which part has not, helping the machine learning model 112 make the read-write decisions and outputs the translation of the unfinished parts. Secondly, understanding human speech often requires context. For example, when a speaker talks about “barrel bridge” , it often refers to the bridges built upon rivers that are supported by barrels. However, in the context of “watch” , it refers to a mechanical structure in the watch. The phenomenon of polysemy in different contexts can lead to vastly different translation outcomes. Therefore, the machine learning model 112 may be able to retrieve the context of the long speech for translating some keywords, and make appropriate translations under different contexts.
[0067] At each round, the translation result for the previous round is loaded as context information to facilitate the machine learning model 112 to detect the semantic integrity of the translation text sequence generated for the audio segment to be translated in the current round. FIG. 3B illustrates a schematic diagram for such iterative translation. As shown in FIG. 3B, at round r for translation of the audio segment 212, the step 205 marked as <LOAD_MEM> may obtain relevant translations y1: r-1 350 at round r-1 from the memory 210, to the machine learning model 112 as a prompt (apart of the model input) . After the machine learning model 112 generates the translation output yr 352 at round r, the step marked as <UPDATE_MEM>may store the translation output 352 to the memory 240 and obtains y1: r. That is, the translation output generated from consecutive rounds may be aggregated and stored in the memory 240. In some embodiments, the translation output or a part of it stored in the memory 240 may be flushed, depending on the upper limit of size configured for the memory 240.
[0068] For example, if the upper limit of translation results corresponding to a total of 30 seconds of audio is configured, the stored information in the memory 240 may be deleted following the first-in-first-out (FIFO) principle. In other example, all the stored information in the memory 240 before the first 30 seconds may be deleted to wait for further context information in a next of 30 seconds. It would be appreciated that there may be various other ways to control the time window of the context information stored in the memory.
[0069] To enable the machine learning model 112 to provide partial translation with semantic integrity for input partial audio, during the model training phase, professional human interpreters are imitated to learn their policy of segmenting a complete sentence into several semantic segments through syntactic boundaries (pauses, commas, conjunctions, etc. ) and contextual meaning. To enable the machine learning model 112 to learn such a policy, a data-driven policy learning process is applied. The training dataset may include partial training samples that are annotated by human interpreters, which includes the read-write timing for segmentation. From the training data, the machine learning model 112 can obtain the robust read-write policy for SiST.
[0070] Unlike predetermined read-write probabilities and heuristic waiting policies detailed in prior research, interpreters engage in a dynamic process of listening (read) and translating (write) . They attentively listen to the speaker’s speech and segment lengthy sentences into semantic chunks, representing the smallest linguistic units capable of conveying a complete thought independently. Upon identifying a chunk that encapsulates sufficient information, they proceed to translate this segment into the target language, thereby providing an accurate and contextually appropriate translation.
[0071] Emulating the strategies of human interpreters, in some embodiments of the present disclosure, the read-write policy is not explicitly defined for the machine learning model 112. The machine learning model 112 may determine the policy by waiting for complete semantic segments. Specifically, given partial speech, the machine learning model 112 may generate the translation for the complete segments (with a complete semantic) of the input speech. The machine learning model 112 may be trained with segmented speech data to learn such ability.
[0072] In the training phase, given source audio x1: M, it may be segment into a series of n segments y1: n. Then n training samples may be obtained, represented as where and yj represents the j-th segment of the audio and the corresponding translation. For training, the objective is to output all complete segmented translation and the cut-off time (optional) given random partial input audio x1: t. The training objective may be defined to minimize the predicted output translation and cut-off time for a segment in the training samples, which may be represented as follows:
[0073] where indicates uniform distribution over time of speech. Trained with Equation (2) , the machine learning model 112 is capable of generating the cut-off timestamp for the input speech. Additionally, the objective function makes the machine learning model 112 wait for appropriate time for starting translation as the machine learning model 112 will output nothing when it determines that the current audio segment does not contain a complete semantic.
[0074] In some embodiments, for the model-based scheme, the scarcity of training data continues to hinder the performance on the SiST task. Addressing the data scarcity of the SiST task, in some embodiments, a three-stage training methodology: pretraining, continual training, and fine-tuning. In the pretraining stage, in some embodiments, the language model and the audio encoder are independently pretrained on large-size datasets. Then, the machine learning model 112 may be continually trained with billions of tokens of mediocre-quality synthesized speech translation data, aiming to align the speech and text modalities. In some embodiments, multiple training tasks may be included to enhance the in-context learning ability of the machine learning model 112 to better utilize the contextual information from the retriever and prior translation. In the last stage, the machine learning model 112 is fine-tuned with a small amount of human-annotated data, to improve the robustness and translation quality.
[0075] Specifically, for streaming and higher-quality translation, three types of tasks may be applied for training the machine learning model 112: Automatic Speech Recognition (ASR) , Speech Translation (ST) , and Text Translation (MT) . To align the modalities of the pretrained language model (e.g., LLM) and audio encoder, the machine learning model 112 is continually trained on various tasks with a substantial volume of paired data. It may further strengthen the in-context learning ability of the approach herein by incorporating translation in the memory and knowledge from external databases. As a result, the ST tasks may be expanded to different configurations as shown in Table 1. A ST translation can either be streaming or offline, direct or COT, with or without context, which leads to 8 different tasks.
[0076] In some embodiments, the major challenge of developing an end-to-end SiST model is the data scarcity of simultaneous ST. To this end, it is proposed a synthetic data construction pipeline. With a strong language model, two types of speech translation data, the ASR training data and the MT training data may be synthesized for continual training: offline ST data and context-aware segmented streaming ST data. In some embodiments, in a first training stage of multi-task continual training, the machine learning model 112 may be trained using a first training dataset, the first training dataset comprises a first number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language. In some embodiments, the first training dataset may include the offline ST data and / or context-aware segmented streaming ST data.
[0077] In some embodiments, ASR data may be used to construct the offline ST data. Given the ground-truth transcription of speech, a trained language model may be utilized to translate the source language to target languages. To ensure the readability and conciseness of the target language, the trained language model is prompted to conduct Inverse Text Normalization (ITN) , filler word smoothing, etc. In some embodiments, the streaming ST data consists of fine-grained audiotext alignments and translation pairs for segmented semantic chunks. Compared to offline ST data, streaming ST data is even more challenging to collect. It is found that some human interpreter often segments long speech into a few semantic chunks, each of which can be translated independently to ensure an effective and smooth translation. Motivated by such findings, a trained language model is utilized to construct streaming ST data by imitating the chunking process. Long speech data are used to construct the streaming ST data, as the additional history can provide better contextual information. First, prompt information to the language model is to instruct the model to break down the ASR transcription into multiple independent semantic chunks, which are then translated into the target language. Subsequently, the semantic chunks are aligned with the corresponding audio chunks, obtaining the streaming ST data. Such training data enable the machine learning model 112 to handle incomplete speech inputs and generate partial translation in coherent semantics.
[0078] In some embodiments, the first training stage may be performed based on at least one of the following: a first training task configured to cause the machine learning model to translate a first sample audio in the source language to a first translation result in the target language; a second training task configured to cause the machine learning to translate a transcription text corresponding to a second sample audio in the source language into a second translation result in the target language; a third training task configured to cause the machine learning model to translate a sample audio segment in the source language into a third translation segment with a complete semantic in the target language; a fourth training task configured to cause the machine learning model to translate a third sample audio with a complete semantic into a fourth translation result in the target language. Each of the above training tasks may be performed with or without historical translation as context information. The multiple training tasks may be defined as in below Table 1.
[0079] Table 1
[0080] Even though machine learning model 112 may possess a good translation quality on the SiST tasks after the previous multi-task continual training stage, in some embodiments, it is proposed to further boost the performance by fine-tuning on human-annotated streaming ST data with diverse tasks. In some embodiments, for a second training stage for multi-task Supervised Fine-tuning, the machine learning model 112 may be trained using a second training dataset, the second training dataset comprises a second number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language. In some embodiments, the translation quality level of the second training dataset may be higher than a translation quality level of the first training dataset. Such high-quality data enables the model to better align with the segmentation methodologies of professional human interpreters. Furthermore, this process enhances the model’s robustness to speech disfluencies such as stuttering, ensuring smoother communication in real-world scenarios.
[0081] The source of human-annotated streaming ST data originates from real world scenarios that contain various speech characteristics, such as disfluencies, stuttering, code-mixing, and specialized terminologies. Such features ensure the robustness of the machine learning model 112 in diverse conditions. In some embodiments, professional human interpreters may be requested to provide high-quality annotations for simultaneous segmentation and interpretation of the speech data. Additionally, terminologies and jargon are identified and translated within the context, further strengthening the context-aware capabilities of the machine learning model 112.
[0082] In some examples, case studies to show the ability of the machine learning model 112 in translating complicated speech for CN-EN and EN-CN in some example tables in FIGS. 3D-3F and some example tables in FIGS. 3G-3I. One of the most-performed cascaded systems Product-X is chosen for comparison with the translation result according to the example embodiment of the present disclosure (marked as CLASI Translation) . The Product-X adopted a cascaded approach for SiST. It is one of the best SiST systems in the market. The Golden Transcription indicates the ground-truth translation in the test, and “CLASI ASR” indicates that ASR result of the test audio segment according to the example embodiment of the present disclosure, and “Product X ASR” indicates the ASR result of the test audio segment from Product-X.
[0083] Detailed explanations are described in the tables. For CN-EN direction, cases regarding robustness to recognition errors, reasoning ability, and trending words translation are presented in FIGS. 3D-3F. For EN-CN direction, the cases regarding native, expressive, and accurate terminology translations are presented in FIGS. 3G-3I.
[0084] FIG. 4 illustrates a flowchart of a process 400 for speech translation according to some embodiments of the present disclosure. The process 400 may be implemented at the speech translation system 110, for example. For ease of discussion, the process 400 will be described with reference to the environment 100 of FIG. 1.
[0085] At block 410, in response to obtaining a first audio segment in a source language, the speech translation system 110 converts the first audio segment into a first audio embedding.
[0086] At block 420, in response to determining that at least one audio segment is obtained before the first audio segment and / or at least one translation segment in a target language corresponding to the at least one audio segment, the speech translation system 110 construct a first model input based on the first audio feature embedding, at least one audio embedding corresponding to the at least one audio segment, and at least one translation embedding corresponding to the at least one translation segment.
[0087] At block 430, the speech translation system 110 generates a first translation segment in the target language corresponding to the first audio segment based on the first model input using a trained machine learning model.
[0088] In some embodiments, generating the first translation segment comprises: determining a cut-off timestamp in the first audio segment, the cut-off timestamp identifying a time in the first audio segment before which a complete semantic is detected; and generating the first translation segment corresponding to an audio portion of the first audio segment before the cut-off timestamp.
[0089] In some embodiments, constructing the first model input further comprises: constructing the first model input further based on at least one cut-off stamp in the at least one audio segment, respectively.
[0090] In some embodiments, constructing the first model input further comprises: determining, from a knowledge base, at least one term matching the first audio segment; and constructing the first model input further based on the at least one term. In some embodiments, generating the first translation segment comprises: generating the first translation segment to include the at least one term in the target language.
[0091] In some embodiments, constructing the first model input further based on the at least one term comprises: constructing the first model input further based on: at least one pair of the at least one term in the source language and a translation of the at least one term in the target language; at least one pair of the at least one term in the target language and a translation of the at least one term in the source language; the at least one term in the source language, or the at least one term in the target language.
[0092] In some embodiments, the first model input further comprises prompt information to indicate a translation from the source language to the target language.
[0093] In some embodiments, the first audio segment is obtained in an interaction session, and constructing the first model input further comprises: constructing the first model input further based on historical context information in the interaction session.
[0094] In some embodiments, the process 400 further comprises: in response to obtaining a second audio segment subsequent to the first audio segment in the source language, converting the second audio segment into a second audio embedding; constructing a second model input based on the second audio feature embedding and at least the first audio embedding and a first translation embedding corresponding to the first translation segment; and providing the second model input to the trained machine learning model to obtain a second model output, the second model output indicating a second translation segment in the target language corresponding to the second audio segment.
[0095] In some embodiments, a training process of the machine learning model comprises: a first training stage configured to train the machine learning model using a first training dataset, the first training dataset comprises a first number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language; and a second training stage configured to train the machine learning model using a second training dataset, the second training dataset comprises a second number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language, a translation quality level of the second training dataset being higher than a translation quality level of the first training dataset.
[0096] In some embodiments, the first training stage is performed based on at least one of the following: a first training task configured to cause the machine learning model to translate a first sample audio in the source language to a first translation result in the target language; a second training task configured to cause the machine learning to translate a transcription text corresponding to a second sample audio in the source language into a second translation result in the target language; a third training task configured to cause the machine learning model to translate a sample audio segment in the source language into a third translation segment with a complete semantic in the target language; a fourth training task configured to cause the machine learning model to translate a third sample audio with a complete semantic into a fourth translation result in the target language.
[0097] FIG. 5 illustrates a schematic structural block diagram of an apparatus 500 for speech translation according to some embodiments of the present disclosure. The apparatus 500 may be implemented as or included in the speech translation system 110. Each unit / component in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0098] As shown, the apparatus 500 includes an embedding converting module 510 configured to in response to obtaining a first audio segment in a source language, convert the first audio segment into a first audio embedding; an input constructing module 520 configured to in response to determining that at least one audio segment is obtained before the first audio segment and / or at least one translation segment in a target language corresponding to the at least one audio segment, constructing a first model input based on the first audio feature embedding, at least one audio embedding corresponding to the at least one audio segment, and at least one translation embedding corresponding to the at least one translation segment; and a translation generating module 530 configured to generate a first translation segment in the target language corresponding to the first audio segment based on the first model input using a trained machine learning model.
[0099] In some embodiments, the translation generating module 530 is further configured to:determine a cut-off timestamp in the first audio segment, the cut-off timestamp identifying a time in the first audio segment before which a complete semantic is detected; and generate the first translation segment corresponding to an audio portion of the first audio segment before the cut-off timestamp.
[0100] In some embodiments, the input constructing module 520 is further configured to: construct the first model input further based on at least one cut-off stamp in the at least one audio segment, respectively.
[0101] In some embodiments, the input constructing module 520 is further configured to: determine, from a knowledge base, at least one term matching the first audio segment; and construct the first model input further based on the at least one term. In some embodiments, the translation generating module 530 is further configured to: generating the first translation segment to include the at least one term in the target language.
[0102] In some embodiments, the input constructing module 520 is further configured to: construct the first model input further based on: at least one pair of the at least one term in the source language and a translation of the at least one term in the target language; at least one pair of the at least one term in the target language and a translation of the at least one term in the source language; the at least one term in the source language, or the at least one term in the target language.
[0103] In some embodiments, the first model input further comprises prompt information to indicate a translation from the source language to the target language.
[0104] In some embodiments, the first audio segment is obtained in an interaction session, and constructing the first model input further comprises: constructing the first model input further based on historical context information in the interaction session.
[0105] In some embodiments, the apparatus 500 further comprises: a second embedding converting module configured to in response to obtaining a second audio segment subsequent to the first audio segment in the source language, converting the second audio segment into a second audio embedding; a second input constructing module configured to construct a second model input based on the second audio feature embedding and at least the first audio embedding and a first translation embedding corresponding to the first translation segment; and an output obtaining module configured to provide the second model input to the trained machine learning model to obtain a second model output, the second model output indicating a second translation segment in the target language corresponding to the second audio segment.
[0106] In some embodiments, a training process of the machine learning model comprises: a first training stage configured to train the machine learning model using a first training dataset, the first training dataset comprises a first number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language; and a second training stage configured to train the machine learning model using a second training dataset, the second training dataset comprises a second number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language, a translation quality level of the second training dataset being higher than a translation quality level of the first training dataset.
[0107] In some embodiments, the first training stage is performed based on at least one of the following: a first training task configured to cause the machine learning model to translate a first sample audio in the source language to a first translation result in the target language; a second training task configured to cause the machine learning to translate a transcription text corresponding to a second sample audio in the source language into a second translation result in the target language; a third training task configured to cause the machine learning model to translate a sample audio segment in the source language into a third translation segment with a complete semantic in the target language; a fourth training task configured to cause the machine learning model to translate a third sample audio with a complete semantic into a fourth translation result in the target language.
[0108] FIG. 6 illustrates a block diagram of an electronic device 600 capable of implementing multiple implementations of the disclosure. It should be understood that the electronic device 600 shown in FIG. 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The electronic device 600 shown in FIG. 6 may be used to implement the speech translation system 110 in FIG. 1 or the apparatus 500 in FIG. 6.
[0109] As shown in FIG. 6, the electronic device 600 is in the form of a general electronic device. The components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and can execute various processes based on the programs stored in the memory 620. In a multiprocessor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0110] The electronic device 600 typically includes multiple computer storage medium. Such medium may be any available medium that is accessible to the electronic device 600, including but not limited to volatile and non-volatile medium, removable and non-removable medium. The memory 620 may be volatile memory (for example, a register, cache, a random access memory (RAM) ) , a non-volatile memory (for example, a read-only memory (ROM) , an electrically erasable programmable read-only memory (EEPROM) , a flash memory) , or any combination thereof. The storage device 630 may be any removable or non-removable medium, and may include a machine readable medium such as a flash drive, a disk, or any other medium, which may be used to store information and / or data (such as training data for training) and may be accessed within the electronic device 600.
[0111] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage medium. Although not shown in FIG. 6, a disk driver for reading from or writing to a removable, non-volatile disk (such as a "floppy disk" ) , and an optical disk driver for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each driver may be connected to the bus (not shown) by one or more data medium interfaces. The memory 620 may include a computer program product 625, which has one or more program units configured to execute various methods or acts of various implementations of the present disclosure.
[0112] The communication unit 640 communicates with a further electronic device through the communication medium. In addition, functionality of components in the electronic device 600 may be implemented by a single computing cluster or multiple computing machines, which can communicate through a communication connection. Therefore, the electronic device 600 may be operated in a networking environment using a logical connection with one or more other servers, a network personal computer (PC) , or another network node.
[0113] The input device 650 may be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as required. The external device, such as a storage device, a display device, etc., communicate with one or more devices that enable users to interact with the electronic device 600, or communicate with any device (for example, a network card, a modem, etc. ) that makes the electronic device 600 communicate with one or more other electronic devices. Such communication may be executed via an input / output (I / O) interface (not shown) .
[0114] According to the example implementations of the present disclosure, a computer-readable storage medium is provided, on which a computer-executable instruction or computer program is stored, wherein the computer-executable instructions is executed by the processor to implement the method described above. According to the example implementations of the present disclosure, a computer program product is also provided. The computer program product is physically stored on a non-transient computer-readable medium and includes computer-executable instructions, which are executed by the processor to implement the method described above. According to the exemplary implementations of the present disclosure, a computer program product is provided having stored thereon a computer program, and when the program is executed by a processor, the method described above is implemented.
[0115] Various aspects of the present disclosure are described herein with reference to the flow chart and / or the block diagram of the method, the apparatus, the device and the computer program product implemented in accordance with the present disclosure. It would be appreciated that each block of the flowchart and / or the block diagram and the combination of each block in the flowchart and / or the block diagram may be implemented by computer-readable program instructions.
[0116] These computer-readable program instructions may be provided to the processing units of general-purpose computers, specialized computers, or other programmable data processing devices to produce a machine that generates an apparatus to implement the functions / actions specified in one or more blocks in the flow chart and / or the block diagram when these instructions are executed through the computer or other programmable data processing apparatuses. These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions enable a computer, a programmable data processing apparatus and / or other devices to work in a specific way. Therefore, the computer-readable medium containing the instructions includes a product, which includes instructions to implement various aspects of the functions / actions specified in one or more blocks in the flowchart and / or the block diagram.
[0117] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operational steps may be executed on a computer, other programmable data processing apparatus, or other devices, to generate a computer-implemented process, such that the instructions which execute on a computer, other programmable data processing apparatuses, or other devices implement the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.
[0118] The flowchart and the block diagram in the drawings show the possible architecture, functions and operations of the system, the method and the computer program product implemented in accordance with the present disclosure. In this regard, each block in the flowchart or the block diagram may represent a part of a unit, a program segment or instructions, which contains one or more executable instructions for implementing the specified logic function. In some alternative implementations, the functions labeled in the block may also occur in a different order from those labeled in the drawings. For example, two consecutive blocks may actually be executed in parallel, and sometimes can also be executed in a reverse order, depending on the functionality involved. It should also be noted that each block in the block diagram and / or the flowchart, and combinations of blocks in the block diagram and / or the flowchart, may be implemented by a dedicated hardware-based system that executes the specified functions or acts, or by the combination of dedicated hardware and computer instructions.
[0119] Each implementation of the present disclosure has been described above. The above description is an example, not exhaustive, and is not limited to the disclosed implementations. Without departing from the scope and spirit of the described implementations, many modifications and changes are obvious to ordinary skill in the art. The selection of terms used in the present disclosure aims to best explain the principles, practical application or improvement of technology in the market of each implementation, or to enable other ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1.A method for speech translation, comprising:in response to obtaining a first audio segment in a source language, converting the first audio segment into a first audio embedding;in response to determining that at least one audio segment is obtained before the first audio segment and / or at least one translation segment in a target language corresponding to the at least one audio segment, constructing a first model input based on the first audio feature embedding, at least one audio embedding corresponding to the at least one audio segment, and at least one translation embedding corresponding to the at least one translation segment; andgenerating a first translation segment in the target language corresponding to the first audio segment based on the first model input using a trained machine learning model.2.The method of claim 1, wherein generating the first translation segment comprises:determining a cut-off timestamp in the first audio segment, the cut-off timestamp identifying a time in the first audio segment before which a complete semantic is detected; andgenerating the first translation segment corresponding to an audio portion of the first audio segment before the cut-off timestamp.3.The method of claim 1, wherein constructing the first model input further comprises:constructing the first model input further based on at least one cut-off stamp in the at least one audio segment, respectively.4.The method of claim 1, wherein constructing the first model input further comprises:determining, from a knowledge base, at least one term matching the first audio segment; andconstructing the first model input further based on the at least one term; andwherein generating the first translation segment comprises:generating the first translation segment to include the at least one term in the target language.5.The method of claim 4, wherein constructing the first model input further based on the at least one term comprises:constructing the first model input further based on:at least one pair of the at least one term in the source language and a translation of the at least one term in the target language;at least one pair of the at least one term in the target language and a translation of the at least one term in the source language;the at least one term in the source language, orthe at least one term in the target language.6.The method of claim 1, wherein the first model input further comprises prompt information to indicate a translation from the source language to the target language.7.The method of claim 1, wherein the first audio segment is obtained in an interaction session, and constructing the first model input further comprises:constructing the first model input further based on historical context information in the interaction session.8.The method of claim 1, further comprising:in response to obtaining a second audio segment subsequent to the first audio segment in the source language, converting the second audio segment into a second audio embedding;constructing a second model input based on the second audio feature embedding and at least the first audio embedding and a first translation embedding corresponding to the first translation segment; andproviding the second model input to the trained machine learning model to obtain a second model output, the second model output indicating a second translation segment in the target language corresponding to the second audio segment.9.The method of claim 1, wherein a training process of the machine learning model comprises:a first training stage configured to train the machine learning model using a first training dataset, the first training dataset comprises a first number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language; anda second training stage configured to train the machine learning model using a second training dataset, the second training dataset comprises a second number of samples, a sample comprising a sample audio or sample audio segments in the source language, and a corresponding translation result in the target language, a translation quality level of the second training dataset being higher than a translation quality level of the first training dataset.10.The method of claim 9, wherein the first training stage is performed based on at least one of the following:a first training task configured to cause the machine learning model to translate a first sample audio in the source language to a first translation result in the target language;a second training task configured to cause the machine learning to translate a transcription text corresponding to a second sample audio in the source language into a second translation result in the target language;a third training task configured to cause the machine learning model to translate a sample audio segment in the source language into a third translation segment with a complete semantic in the target language;a fourth training task configured to cause the machine learning model to translate a third sample audio with a complete semantic into a fourth translation result in the target language.11.An electronic device comprising:at least one processing unit; andat least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the electronic device to perform the method of any of claims 1 to 10.12.An apparatus for speech translation, comprising:an embedding converting module configured to in response to obtaining a first audio segment in a source language, convert the first audio segment into a first audio embedding;an input constructing module configured to in response to determining that at least one audio segment is obtained before the first audio segment and / or at least one translation segment in a target language corresponding to the at least one audio segment, constructing a first model input based on the first audio feature embedding, at least one audio embedding corresponding to the at least one audio segment, and at least one translation embedding corresponding to the at least one translation segment; anda translation generating module configured to generate a first translation segment in the target language corresponding to the first audio segment based on the first model input using a trained machine learning model.13.A computer readable storage medium having stored thereon a computer program, when executed by a processor, implementing the method of any of claims 1 to 10.
Citation Information
Patent Citations
Robust direct speech-to-speech translation
US11960852B2
Enhanced speech-to-speech translation system and methods
US20110307241A1