Training method and translation method of multi-language instant speech translation model
By extracting features from audio and text corpora and fusing them into a unified semantic space, building a multilingual instant speech translation model, and using a large language model to expand the knowledge base, the accuracy problem of traditional translation methods in multilingual mixed scenarios is solved, and efficient and smooth multilingual translation is achieved.
Patent Information
- Application Number
- CN202510811452.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional translation methods have poor translation effects in high-real-time and multi-language support scenarios, and are difficult to handle situations where the same audio contains a mixture of multiple languages, resulting in a poor user experience.
By extracting features from audio and text corpora, fusing the features and mapping them to a unified semantic space, a multilingual instant speech translation model is built, and the knowledge base is expanded using a large language model to achieve multilingual translation.
It improves the accuracy and fluency of multilingual translation, can handle single-language and mixed-language input, is suitable for real-time communication, and reduces training costs and time expenditure.
Smart Images

Figure CN120673749A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of multimodal learning technology, and in particular to a training method and a translation method for a multilingual instant speech translation model. Background Art
[0002] Language translation refers to the process of converting spoken or written language from one language into another. Traditional translation methods typically rely on a three-stage process: first, automatic speech recognition (ASR) converts source language speech into text, then machine translation (MT) technology translates the source text into the target language, and finally, text-to-speech (TTS) synthesizes the target text into speech output. While this approach can meet the needs of some scenarios, it can be relatively ineffective in complex scenarios.
[0003] For example, in scenarios requiring high real-time performance and multilingual support, the limitations of traditional methods are particularly evident. Because traditional translation methods typically require a single, fixed input language and a fixed target language, they struggle to handle situations where multiple languages are mixed within the same audio clip, resulting in a poor user experience. Summary of the Invention
[0004] The embodiments of the present application provide a training method and a translation method for a multilingual instant speech translation model to solve the problem of inaccurate multilingual instant speech translation by traditional translation methods.
[0005] In a first aspect, an embodiment of the present application provides a training method for a multilingual instant speech translation model, the method comprising: extracting speech features from an audio corpus, and extracting text features from a text corpus corresponding to the audio corpus; wherein the audio corpus and its corresponding text corpus are divided into semantic categories, and the same category contains corpora corresponding to multiple different languages; the speech features are vector representations of the audio corpus, and the text features are vector representations of the text corpus; the speech features and text features representing the same semantics in the same language are fused to obtain fused features corresponding to each semantics; the fused features are respectively mapped to the same preset semantic space to obtain multimodal features corresponding to each fused feature; and a preset multimodal basic model is trained using the multimodal features as training data to obtain a multilingual instant speech translation model.
[0006] In this way, multiple input forms such as voice and text can be integrated into a unified model architecture, so that the model can handle translation tasks between multiple languages.
[0007] In one possible implementation, after fusing speech features and text features representing the same semantics in the same language to obtain the fused features, the method further includes aligning multimodal features representing the same semantics in different languages using a contrastive loss function.
[0008] Based on this, the distance between multimodal features representing the same semantics in different languages can be minimized in the semantic space. Multimodal features corresponding to the same semantics in different languages are tightly aggregated in the semantic space, which can effectively improve the multilingual instant speech translation model's accurate understanding of semantics and cross-language conversion capabilities, and reduce translation errors.
[0009] In one possible implementation, the method further includes: obtaining additional text corpora, where the additional text corpora correspond to different languages and / or domains than the text corpora; obtaining additional audio corpora corresponding to the additional text corpora; and incrementally training a multilingual instant speech translation model based on the additional text corpora and the additional audio corpora to update certain parameters of the multilingual instant speech translation model. Incremental training allows updating only certain model parameters, avoiding full model retraining and maintaining the stability of the model's learned knowledge while rapidly adapting to new data, ensuring stable translation performance and reducing training costs and time.
[0010] In the second aspect, the embodiment of the present application provides a knowledge base construction method, which is applied to the multilingual instant speech translation model obtained by training the training method provided by the first aspect and its various implementation methods. The method includes: constructing a first language knowledge base; wherein the first language knowledge base includes a vocabulary consisting of multiple first entries, each first entry includes a text unit and its corresponding text vector and audio vector, the text vector has the same dimension as the text feature, the audio vector has the same dimension as the speech feature, and the text vector and text feature expressing the same semantics are the same, the audio vector and speech feature expressing the same semantics are the same, and the text unit is the granular text corresponding to the first language; using the large language model LLM to expand the first entry, to obtain the text expansion vector and audio expansion vector corresponding to each first entry; the text expansion vector and audio expansion vector are the vectors corresponding to the text unit in at least one second language; the text expansion vector and audio expansion vector are added to the corresponding first entry to form a multilingual knowledge base. In this way, a multilingual knowledge base can be constructed to provide knowledge support for the model to optimize the language coverage and translation quality of the model.
[0011] In one possible implementation, expanding the first entry using a large language model (LLM) includes: generating an extended text corresponding to the text unit using the LLM, the extended text including corresponding text content in at least one second language; and / or obtaining target text content corresponding to each text unit in at least one second language from a public data source; wherein the target text content is located within paragraph content; and extracting the target text content from the paragraph content using the LLM to obtain the extended text. In this manner, the extended text can be obtained.
[0012] In one possible implementation, after generating the extended text corresponding to a text unit using the LLM and / or obtaining the extended text, the method further includes: generating extended audio corresponding to the extended text using an audio generation model; converting the extended text into an extended text vector; and converting the extended audio into an extended audio vector. This allows for automated expansion of the multilingual knowledge base, improving its coverage and timeliness.
[0013] In one possible implementation, before expanding the first entry using the Large Language Model (LLM), the method further includes: expanding the vocabulary using the LLM to form first entries corresponding to other text units in the first language, where the other text units belong to the same or different domains as the text unit. In this way, the multilingual knowledge base can cover more domain knowledge.
[0014] In a third aspect, an embodiment of the present application provides a translation method based on a multilingual instant speech translation model, which is trained based on the training method in the aforementioned first aspect and its various implementations. The translation method includes: obtaining audio to be translated, which is formed by a single language or a mixture of multiple languages; converting the audio to be translated into source language text; converting the source language text into a source language text vector; searching the source language text vector in a multilingual knowledge base to obtain a mapping relationship; wherein the multilingual knowledge base is constructed by the knowledge base construction method in the aforementioned first aspect and its various implementations, and the mapping relationship includes the source language text vector and its corresponding source language audio vector, as well as the text vector corresponding to the source language text in the target language; forming a prompt word based on the mapping relationship; inputting the audio to be translated and the prompt word into the multilingual instant speech translation model, so that the multilingual instant speech translation model translates the audio to be translated into the target language audio based on the prompt word.
[0015] Based on this, it is possible to process single-language and mixed-language inputs and generate translation output in the target language. Furthermore, the translation method provided in the embodiments of this application does not include the steps of converting audio to text, translating into target language text, and converting target language text to audio. The multilingual instant speech translation model can directly convert source language speech to target language speech, making the translation process smoother and more natural, suitable for real-time communication.
[0016] In one possible implementation, the method further includes: obtaining interaction logs at a preset frequency, the interaction logs including user feedback on the target language audio, and also including the audio to be translated and / or source language text involved in the interaction process; determining whether the audio to be translated includes language and / or specific domain vocabulary not covered by the multilingual knowledge base based on the feedback results, the audio to be translated and / or the source language text; if the audio to be translated includes language and / or specific domain vocabulary not covered by the multilingual knowledge base, obtaining extended text of the uncovered language and / or extended text of the specific domain from a public data source; determining a text extension vector and an audio extension vector based on the extended text; and adding the text extension vector and the audio extension vector to the vocabulary of the multilingual knowledge base based on the semantics of the extended text.
[0017] This allows for alignment between new languages or domains and existing languages, enabling dynamic updates to the multilingual knowledge base, thereby maintaining its timeliness and comprehensiveness. By combining user feedback and real-time interactive content, the multilingual knowledge base's coverage can be optimized, particularly by supplementing and improving low-resource languages, achieving continuous iterative updates.
[0018] In a fourth aspect, an embodiment of the present application also provides a training device for a multilingual instant speech translation model, the device comprising: a corpus conversion module, configured to extract speech features from an audio corpus, and to extract text features from a text corpus corresponding to the audio corpus; wherein the audio corpus and its corresponding text corpus are divided according to semantic categories, and the same category contains corpora corresponding to multiple different languages; speech features are vector representations of the audio corpus, and text features are vector representations of the text corpus; a cross-modal fusion module, configured to fuse speech features and text features representing the same semantics in the same language to obtain fused features corresponding to each semantics. A vector mapping module, configured to map the fused features to the same preset semantic space respectively to obtain multimodal features corresponding to each fused feature. An iterative training module, configured to train a preset multimodal basic model using multimodal features as training data to obtain a multilingual instant speech translation model.
[0019] In a fifth aspect, an embodiment of the present application further provides a knowledge base construction device, which includes: a first construction module, configured to: construct a first language knowledge base; wherein the first language knowledge base includes a vocabulary consisting of multiple first entries, each first entry includes a text unit and its corresponding text vector and audio vector, the text vector has the same dimension as the text feature, the audio vector has the same dimension as the voice feature, and the text vector and text feature expressing the same semantics are the same, the audio vector and voice feature expressing the same semantics are the same, and the text unit is a granular text corresponding to the first language; an expansion module, configured to: expand the first entry using a large language model LLM to obtain a text expansion vector and an audio expansion vector corresponding to each first entry; the text expansion vector and the audio expansion vector are vectors corresponding to the text unit in at least one second language; and the text expansion vector and the audio expansion vector are added to the corresponding first entry to form a multilingual knowledge base.
[0020] In a sixth aspect, an embodiment of the present application further provides a translation device based on a multilingual instant speech translation model, the device comprising: an acquisition module configured to: acquire audio to be translated, the audio to be translated being formed by a single language or a mixture of multiple languages; a translation module configured to: convert the audio to be translated into source language text; convert the source language text into a source language text vector; search the source language text vector in a multilingual knowledge base to obtain a mapping relationship; wherein the multilingual knowledge base is constructed by the knowledge base construction method in the aforementioned first aspect and its various implementation methods, the mapping relationship includes the source language text vector and its corresponding source language audio vector, and also includes the text vector corresponding to the source language text in the target language; a prompt word is formed based on the mapping relationship; the audio to be translated and the prompt word are input into the multilingual instant speech translation model, so that the multilingual instant speech translation model translates the audio to be translated into the target language audio based on the prompt word.
[0021] In the seventh aspect, an embodiment of the present application also provides a computing device, comprising: one or more processors; and a memory configured to: store one or more programs; wherein, when the one or more programs are executed by one or more processors, the one or more processors implement a method as described in any one of the first, second or third aspects above.
[0022] In an eighth aspect, an embodiment of the present application provides a chip, which is used to execute the method of any one of the first, second or third aspects above.
[0023] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a computer, a method as described in any one of the first, second, or third aspects is implemented.
[0024] In a tenth aspect, an embodiment of the present application provides a program product, comprising a computer program, which, when executed by a processor, implements a method as described in any one of the first, second, or third aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram of the architecture of the translation system provided in an embodiment of the present application;
[0026] Figure 2 A flowchart of a method for training a multilingual instant speech translation model provided in an embodiment of the present application.
[0027] Figure 3 A flowchart of a method for training a multilingual instant speech translation model and a method for constructing a knowledge base provided in an embodiment of the present application;
[0028] Figure 4 A flowchart of a method for constructing a knowledge base provided in an embodiment of the present application;
[0029] Figure 5 A flowchart of a translation method based on a multilingual instant speech translation model provided in an embodiment of the present application;
[0030] Figure 6 A schematic diagram of a process for dynamically updating a knowledge base provided in an embodiment of the present application;
[0031] Figure 7 A schematic diagram of the structure of a training device for a multilingual instant speech translation model provided in an embodiment of the present application;
[0032] Figure 8 A schematic diagram of the structure of the knowledge base construction device provided in an embodiment of the present application.
[0033] Figure 9 A schematic diagram of the structure of a translation device based on a multilingual instant speech translation model provided in an embodiment of the present application;
[0034] Figure 10 A schematic diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] The following will describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. To facilitate the clear description of the technical solutions in the embodiments of the present application, the first, second, etc. descriptions in the embodiments of the present application are only used for illustration and to distinguish the described objects. There is no order, nor does it represent a special limitation on the number of devices in the embodiments of the present application, and it does not constitute any limitation on the embodiments of the present application.
[0036] Before introducing the technical solutions of the embodiments of the present application, an exemplary introduction to the terms involved in the embodiments of the present application is first given.
[0037] 1. Speech-to-speech model: This is an artificial intelligence model that can output voice answers based on input voice. It is widely used in scenarios such as real-time translation, voice interaction, and voice conversion.
[0038] 2. Knowledge base (enterprise knowledge base): It is a systematic collection of knowledge that can store, organize and manage knowledge in some form so that knowledge can be retrieved, shared and utilized. The knowledge in the knowledge base can be applied to solve specific problems, support decision making, automate processes, etc.
[0039] 3. Prompt: refers to the input text used to guide or stimulate the model to perform a specific task. The purpose is to help the model understand the type of task the user wants to perform or the required output format.
[0040] 4. Multimodality: This refers to the simultaneous use of multiple data types or perceptual methods (called modalities) to process information. In the field of artificial intelligence, especially machine learning and deep learning, multimodal methods combine multiple data types (such as text, images, audio, video, etc.) to more comprehensively understand and process complex information.
[0041] Figure 1 A schematic diagram of the architecture of the translation system provided in an embodiment of the present application.
[0042] like Figure 1 As shown, in order to solve the problem of inaccurate translation of multi-language mixed speech, an embodiment of the present application provides a translation system, which can be built based on a computing device such as a server. In terms of form, the server can be a rack server or a whole cabinet server; in terms of performance, it can be a general server, a GPU (graphics processing unit) server, etc. or an artificial intelligence (AI) server.
[0043] The translation system architecture at least includes a processing module 1001, a multilingual instant speech translation model 1002 and a multilingual knowledge base 1003. Wherein, the multilingual instant speech translation model 1002 can be obtained by the training method of the multilingual instant speech translation model provided by the embodiment of the present application. After training, the model can process multilingual mixed speech input and translate it into the audio of the target language. And, the multilingual knowledge base 1003 can be constructed by the knowledge base construction method provided by the embodiment of the present application, for providing prior knowledge. When building the multilingual knowledge base, the embodiment of the present application can make the text vector and audio vector in the knowledge base consistent with the feature dimension of the multilingual instant speech translation model 1002 and semantically aligned. In this way, the text vector and audio vector in the knowledge base can be utilized to provide rich semantic information and contextual support for the multilingual instant speech translation model 1002, thereby improving the accuracy and robustness of translation.
[0044] Furthermore, processing module 1001 can be used to execute the multilingual real-time model-based translation method provided in the embodiment of the present application to call multilingual real-time speech translation model 1002 and multilingual knowledge base 1003 to implement translation. Through the collaborative work of processing module 1001, multilingual real-time speech translation model 1002, and multilingual knowledge base 1003, the translation system of the embodiment of the present application can efficiently and accurately handle multilingual mixed speech translation tasks.
[0045] The following is a detailed introduction to the training method of the multilingual instant speech translation model, the knowledge base construction method, and the translation method based on the multilingual instant model with reference to the accompanying drawings.
[0046] First, the training method for the instant speech translation model provided in the embodiments of this application can be applied to a computing device, which can specifically be a server or a terminal device. For example, the server can be a rack server or a whole cabinet server in terms of form; and can be a general-purpose server, a GPU server, or an artificial intelligence server in terms of performance.
[0047] Figure 2 A flowchart of a method for training a multilingual instant speech translation model provided in an embodiment of the present application.
[0048] Figure 3 A flowchart of the training method and knowledge base construction method for the multilingual instant speech translation model provided in the embodiments of the present application.
[0049] like Figure 2 and Figure 3 As shown, the training method provided in the embodiment of the present application may include the following steps S101-S104.
[0050] S101: extracting speech features from an audio corpus, and extracting text features from a text corpus corresponding to the audio corpus.
[0051] First, the embodiment of the present application can construct a training data set, which includes multilingual audio corpus and multilingual text corpus for training. In addition, in order to achieve the purpose of real-time multilingual translation, the embodiment of the present application can use corpus composed of multiple different languages. The multiple different languages can include official languages such as Arabic, Chinese, English, French, Russian and Spanish, as well as common languages such as German, Japanese, Korean, and minority languages or dialects to meet the multilingual translation needs in different scenarios. The embodiment of the present application does not make specific limitations on this.
[0052] In the training data set, audio corpus and text corpus can be granularized and split into basic language units such as words and phrases. In addition, audio corpus and text corpus can be divided according to semantic categories. The same category contains corpus corresponding to multiple languages, forming multilingual synonymous aligned audio corpus and multilingual synonymous aligned text corpus. For example, in the semantic category of "greeting", "hello" in Chinese, "hello" in English, "hello" in Arabic, "hello" in English, "hello" in Arabic, "hello" in Chinese, "hello" in English, "hello" in Arabic, "hello" in English ... Audio and text expressions in different languages, such as the French word "Bonjour"; in this way, different language corpora with the same or similar semantics are aggregated, so that subsequent model training can effectively capture the similarities and differences between different languages when expressing the same semantics.
[0053] In the embodiment of the present application, the text corpus and the audio corpus can be obtained from a public multilingual corpus to enrich the diversity of the corpus. In addition, the embodiment of the present application can also obtain the corpus based on the following steps of generating the corpus. Specifically, step S101 may include the following steps S1011-S1012.
[0054] S1011: Input prompt words to a large language model (LLM) so that the LLM generates training text corpus based on the prompt words.
[0055] S1012: Input the training text corpus into the TTS model, so that the TTS model generates training audio corpus corresponding to the training text corpus.
[0056] In practical applications, both LLM generation and public data source acquisition can be used simultaneously, and this embodiment of the present application does not specifically limit this.
[0057] In some implementations, before feature extraction, audio and text corpora can be tokenized. For audio corpora, tokenization involves segmenting continuous speech signals into acoustic units (e.g., phonemes) to form a sequence of discrete speech units. For text corpora, tokenization involves breaking text into its smallest processing units, such as morphemes.
[0058] Furthermore, during the feature extraction step of the audio corpus, acoustic feature extraction techniques such as Mel-Frequency Cepstral Coefficients (MFCC) and Linear Prediction Cepstral Coefficients (LPCC) can be used to convert the speech information contained in the audio corpus into computer-processable numerical features, ultimately obtaining speech features in vector form. Furthermore, text feature embedding can be used to map words, subwords, or characters in the text corpus into text vectors, i.e., text features.
[0059] It should also be noted that in the process of feature extraction, the consistency of the dimensions of the speech feature and text feature vectors should be maintained. For example, the extracted speech features and text features are both 256-dimensional vectors or both 512-dimensional vectors, so as to achieve the purpose of unified feature representation. In some implementations, speech features and text features of the same dimension can be extracted. In another implementation, speech features and text features of different dimensions can be extracted, and then the feature alignment module maps the speech features and text features of different dimensions to the same dimension.
[0060] In some implementations, before feature extraction, the audio corpus and text corpus may be cleaned to remove noise and redundant information.
[0061] S102: Fusing speech features and text features that represent the same semantics in the same language to obtain fused features corresponding to each semantics.
[0062] The purpose of feature fusion is to obtain a more comprehensive and accurate feature representation by combining the multimodal information of speech and text. This can effectively capture the relationship between speech and text, ensure that the acoustic information of speech and the semantic information of text can complement each other, ensure that the trained model is robust to complex semantics, dialects and different accents, and improve the model's adaptability to complex semantics and accent changes.
[0063] In the embodiment of the present application, a deep learning framework (such as the Transformer architecture) can be used to construct a multimodal feature representation. Through the cross-modal attention mechanism, the speech branch and the text branch can perceive each other's information, and realize the interaction and fusion of the two modal features. Finally, after multiple layers of feature interaction and fusion operations, the output integrates the fusion features of speech acoustic information and text semantic information, that is, multimodal features, to provide a high-quality multimodal data foundation for subsequent model training. The embodiment of the present application does not make specific limitations on this.
[0064] S103: Mapping the fused features to the same preset semantic space respectively to obtain multimodal features corresponding to each fused feature.
[0065] The semantic space is a vector space shared by fused features corresponding to different languages (i.e., the aforementioned multimodal features). By mapping fused features into the same semantic space, we can break down language barriers and achieve semantic alignment across languages, making the representation of the same meaning closer in different languages. Training models based on these multimodal features can improve the semantic consistency and fluency of translations.
[0066] In practical applications, the mapping process can be implemented using mapping functions or neural network architectures, such as the Transformer network based on the attention mechanism, the generative adversarial network (GAN), etc., to transform the fusion features (multimodal features) of each language from the original feature space to a unified semantic space coordinate system.
[0067] S104: Using the multimodal features as training data, a preset multimodal basic model is trained to obtain a multilingual instant speech translation model.
[0068] Among them, the preset multimodal basic model can be an open source Transformer architecture model, such as mBART, mT5, etc., and the embodiments of this application do not make specific limitations on this.
[0069] Furthermore, the parameters of the preset multimodal basic model can be adjusted through supervised learning, and multiple rounds of iterations can be performed until the model translation accuracy reaches the expected standard.
[0070] From the above content, it can be seen that the embodiment of the present application provides a training method for a multilingual instant speech translation model, which includes: extracting speech features from an audio corpus, and extracting text features from a text corpus corresponding to the audio corpus; wherein the audio corpus and its corresponding text corpus are divided into semantic categories, and the same category contains corpora corresponding to multiple different languages; the speech features are vector representations of the audio corpus, and the text features are vector representations of the text corpus; the speech features and text features representing the same semantics in the same language are fused to obtain fused features corresponding to each semantics; the fused features are mapped to the same preset semantic space respectively to obtain multimodal features corresponding to each fused feature; the multimodal features are used as training data to train a preset multimodal basic model to obtain a multilingual instant speech translation model. In this way, multiple input forms such as speech and text can be integrated into a unified model architecture so that the model can handle translation tasks between multiple languages.
[0071] Furthermore, after mapping the fused features to the same preset semantic space in step S103, the multimodal features in different languages expressing the same semantics have a certain degree of discreteness in the same semantic space. To achieve semantic alignment, embodiments of the present application can further minimize the distance in the semantic space between the multimodal features in different languages expressing the same semantics. Specifically, step S103 can be followed by the following step S105.
[0072] S105: Using a contrast loss function, align the multimodal features representing the same semantics in different languages to minimize the distance between the multimodal features representing the same semantics in different languages in the semantic space.
[0073] Among them, the contrastive loss function is a function used to optimize feature representation, which can make the distance between positive sample pairs (i.e., multimodal features representing the same semantics in different languages) in the semantic space as small as possible, while making the distance between negative sample pairs (i.e., multimodal features representing different semantics) as large as possible. Specifically, the contrastive loss function can be expressed as:
[0074]
[0075] in: represents the distance between the i-th positive sample pair in the semantic space, represents the distance between the i-th negative sample pair in the semantic space, and m is the preset margin, which is used to control the minimum distance between positive and negative sample pairs. If The value is Otherwise the value is 0.
[0076] It should be noted that the contrast loss function here is only an example introduction. In actual application, the contrast loss function can be adjusted based on the training effect.
[0077] Based on this, multimodal features corresponding to the same semantics in different languages are tightly aggregated in the semantic space, which can effectively improve the multilingual instant speech translation model's accurate understanding of semantics and cross-language conversion capabilities, and reduce translation errors.
[0078] Furthermore, the embodiments of the present application can also provide an incremental update mechanism, which only updates part of the model parameters through the incremental fine-tuning module, avoiding retraining of the entire model.
[0079] Specifically, the embodiment of the present application can judge the accuracy of the model translation during the training process and decide whether to trigger incremental training based on the judgment result. For example, the translation accuracy can be evaluated by comparing the translation result of the model on the translation sample audio with the annotated correct translation, and the translation audio sample can come from the training data set. When the translation results of the model in multiple consecutive batches of training do not meet the requirements, for example, the translation accuracy is lower than the preset threshold, or the error rate fluctuates significantly, the embodiment of the present application can automatically trigger the incremental update mechanism.
[0080] Furthermore, the incremental training step may include the following steps S106 - S108 , and steps S106 - S108 may be performed after step S104 .
[0081] S106: Acquire newly added text corpus, where the newly added text corpus corresponds to a different language and / or a different field than the text corpus.
[0082] The acquisition of the newly added text corpus may refer to the step of generating the training corpus in the aforementioned step S101, which will not be described in detail here.
[0083] S107: Obtain the newly added audio corpus corresponding to the newly added text corpus.
[0084] The newly added audio corpus may be obtained by performing TTS conversion on the newly added text corpus.
[0085] S108: Incrementally train the multilingual instant speech translation model based on the newly added text corpus and the newly added audio corpus to update some parameters of the multilingual instant speech translation model.
[0086] Specifically, the newly added text corpus and the newly added audio corpus can be processed based on the steps of the aforementioned steps S102 and S103, and then fine-tuning can be performed. During the fine-tuning process, low-rank adaptation (LoRA) technology can be used to update only part of the model parameters, such as the weights of the last few layers or the parameters of a specific module, rather than updating all the parameters of the entire model. In this way, retraining the entire model can be avoided, maintaining the stability of the model's learned knowledge, while quickly adapting to new data, ensuring stable translation performance, and reducing training costs and time overhead.
[0087] Through continuous iterative training, the embodiment of the present application can obtain an efficient and adaptable multilingual instant speech translation model, and the model can ensure the translation quality and accuracy in different scenarios.
[0088] To further enhance the model's translation capabilities and accuracy, embodiments of the present application also provide a knowledge base construction method, which builds a multilingual knowledge base to provide knowledge support for the model, particularly providing knowledge support in specialized fields, thereby optimizing the model's language coverage and translation quality. This knowledge base construction method can be applied to a computing device, specifically a server or terminal device.
[0089] The following is a detailed introduction to the steps of building a multilingual knowledge base with reference to the accompanying drawings.
[0090] Figure 4 A flowchart of the knowledge base construction method provided in an embodiment of the present application.
[0091] like Figure 3 and Figure 4 As shown, the training method provided in the embodiment of the present application may further include the following steps S201-S204.
[0092] S201: Construct a first language knowledge base; wherein the first language knowledge base is a vocabulary consisting of multiple first entries, each first entry includes a text unit and its corresponding text vector and audio vector, and the text unit is a granular text corresponding to the first language.
[0093] Among them, the first language knowledge base is a single language knowledge base, that is, it is composed of knowledge of a single language (first language). The first language is, for example, Chinese or English. The text unit can be a vocabulary, a phrase, or a sentence. Taking English as an example, for the sentence "I love playing football", it can be granularized into the words "I", "love", "playing", and "football" as different text units. The text units in the vocabulary can be obtained from an existing corpus, an industry terminology set, or a public language resource, and the embodiments of the present application do not specifically limit this.
[0094] In some implementations, the first entry may also include context information, syntax, and grammar corresponding to the text unit, which is not specifically limited in the embodiments of the present application.
[0095] Afterwards, the text unit can be mapped into a text vector through text feature embedding technology, and the audio corresponding to the text unit can be mapped into an audio vector through acoustic feature extraction methods such as MFCC and LPCC.
[0096] It is worth noting that text vectors and text features have the same dimensions, audio vectors and speech features have the same dimensions, and text vectors and text features that express the same semantics are the same, as are audio vectors and speech features that express the same semantics. This way, when applying the knowledge in the knowledge base (including but not limited to text vectors and audio vectors) to the multilingual instant speech translation model, the model does not need to perform complex dimensionality conversion operations due to the consistent dimensions. Instead, the vector information in the knowledge base can be directly integrated into the calculation process, reducing data processing time and computing resource consumption. Furthermore, the same semantics can ensure translation accuracy.
[0097] In some implementations, the embodiments of the present application may adopt semantic verification rules to verify the accuracy and consistency of the data in the first language knowledge base, ensuring that each text unit is semantically consistent with its corresponding text vector and audio vector. For example, for a text unit and its corresponding audio vector, consistency can be verified by manually checking whether the audio accurately reflects the text content. In addition, automated tools or pre-trained models can also be used to verify whether the text unit and its vector representation match to ensure the reliability of data quality. The embodiments of the present application do not limit the specific semantic verification method, and a suitable method can be selected according to actual needs and resources.
[0098] Understandably, the scope of vocabulary coverage has certain field limitations. For example, the first language knowledge base only covers the medical, computer, and communications fields, but not finance. To adapt the multilingual instant speech translation model to more practical application needs, when building the knowledge base, we can, on the one hand, expand the vocabulary coverage by collecting professional corpora from more fields; on the other hand, for existing fields, we can continuously update and improve the vocabulary content to include newly emerging professional terms and expressions.
[0099] In the embodiment of the present application, the knowledge base can cover more fields by expanding the entries, as shown in the following step S202.
[0100] S202: Expand the vocabulary using the LLM to form first entries corresponding to other text units in the first language, where the other text units belong to the same or different fields as the text unit.
[0101] It is understood that step S202 is a step of expanding the vocabulary, which can vertically expand the vocabulary to increase the number of words covered by the vocabulary. Specifically, new domain corpus can be collected, and then the new domain corpus can be tokenized and vectorized to form new text units, i.e., other text units, and corresponding text vectors and audio vectors. Finally, the new text unit and its corresponding vector are added as the first new entry in the vocabulary to achieve the purpose of item expansion, thereby improving the breadth of knowledge base coverage.
[0102] In some implementations, data can be obtained through Internet searches, such as obtaining text materials from professional websites, academic databases, industry reports and other channels through web crawler technology, and combining it with TTS to form text-to-audio. For example, when adding new knowledge in the tourism field, you can crawl text such as attraction introductions and hotel reservation guides on travel guide websites, and then generate audio files through TTS. In addition, for covered fields such as medicine and computers, new professional vocabulary will continue to emerge with the development of the industry and technological updates, such as the names of newly discovered diseases in the medical field and newly proposed algorithm concepts in the computer field. The embodiments of this application can also regularly review and supplement the vocabulary in these fields to ensure the timeliness and accuracy of the knowledge base.
[0103] In this way, we can make associative expansions based on existing text units and further explore semantic associations to make the content of the first language knowledge base richer and more diverse.
[0104] It should also be noted that the expansion step of step S202 may not be limited to the model training phase, and the model reasoning phase may also be performed in real time to achieve dynamic expansion of the knowledge base.
[0105] Furthermore, the embodiment of the present application can make the knowledge base cover more language knowledge by means of vector expansion. Specifically, the following step S203 can be further included after step S100 or step S200.
[0106] S203: Expand the first entries using the large language model (LLM) to obtain a text expansion vector and an audio expansion vector corresponding to each first entry; the text expansion vector and the audio expansion vector are vectors corresponding to the text unit in at least one second language.
[0107] Step S203 is the step of vector expansion of the first entry, which can horizontally expand the vocabulary. This horizontal expansion can make the vocabulary cover more languages, thereby forming a multilingual knowledge base. In this way, when performing multilingual instant speech translation, the model can obtain richer cross-language semantic information, significantly improving the accuracy and fluency of translation.
[0108] Specifically, the vector expansion step includes the following steps S2031-S2036.
[0109] S2031: Generate an extended text corresponding to the text unit using the LLM, where the extended text includes corresponding text content in at least one second language.
[0110] The second language is a language different from the first language, and the second language is not limited to one. For example, the first language may be Chinese, and the second language may be English, Korean, Japanese, Italian, French, etc. The embodiments of this application do not limit the specific number of second languages. In actual applications, the number of second languages can be expanded based on actual conditions.
[0111] It is understood that the LLM has semantic understanding and generation capabilities. In embodiments of the present application, specific prompt words can be input into the LLM to guide the LLM to output the text content corresponding to the text unit in at least one second language. For example, if the English word "I" is input into the LLM, the LLM can generate its corresponding French translation "Je" or Spanish translation "Yo".
[0112] Using LLM to generate extended text is one way of text expansion. The embodiment of the present application can also obtain public text from a public data source, which can specifically include the following steps S2032-S2033.
[0113] S2032: Obtain target text content corresponding to each text unit in at least one second language from a public data source.
[0114] In an embodiment of the present application, target text content can be obtained from data sources such as public databases and forums using web crawler technology. The target text content obtained using this method can be located within paragraph content. For example, the target text content may be nested within a long paragraph on a web page or within a section of a document. In an embodiment of the present application, further extraction of the target text content is possible.
[0115] S2033: Input the paragraph content into the LLM, and use the LLM to extract the target text content from the paragraph content to obtain the extended text.
[0116] Specifically, the present embodiment can leverage the text understanding and information extraction capabilities of LLM to locate and extract the second language target text content corresponding to the first language text unit from a paragraph. For example, from a webpage paragraph containing translations in multiple languages, the French translation "Je" corresponding to the English word "I" can be extracted.
[0117] In this way, the extended text can be obtained. In order to add the extended text to the entry, the following steps S2034-S2035 can be included after steps S2031 and S2033.
[0118] S2034: Generate extended audio corresponding to the extended text using the audio generation model.
[0119] It is understandable that when the extended text is input into the audio generation model, the audio generation model can simulate natural and fluent speech based on the input extended text, and the speech is a second language.
[0120] S2035: Convert the extended text into a text extension vector.
[0121] S2036: Convert the extended audio into an audio extension vector.
[0122] In the embodiment of the present application, the extended text and extended audio can be tokenized, and then the tokenized extended text can be mapped to a text extension vector through text feature embedding technology, and the tokenized extended audio can be mapped to an audio extension vector through acoustic feature extraction methods such as MFCC and LPCC. The embodiment of the present application does not make specific limitations on this.
[0123] Furthermore, step S204 may be included after step S203.
[0124] S204: Add the text extension vector and the audio extension vector to the corresponding first entry to form a multilingual knowledge base.
[0125] In summary, the vocabulary can be expanded vertically and horizontally so that the first entry includes text vectors and audio vectors corresponding to multiple languages and different fields, thereby forming a multilingual knowledge base.
[0126] It should also be noted that in the embodiment of the present application, the number of first language knowledge bases (single language knowledge bases) is at least one. In actual applications, single knowledge bases can be constructed for several key languages, that is, multiple first language knowledge bases are formed. Each first language knowledge base is then expanded in terms of entries and vectors to cover a richer range of knowledge. Finally, the vocabularies of multiple first language knowledge bases can be aggregated to form a multilingual knowledge base including multiple vocabularies, thereby improving the coverage of the knowledge base.
[0127] In some implementations, a single language knowledge base may include one or more vocabularies, and different vocabularies may be divided according to fields, which is not specifically limited in the embodiments of the present application.
[0128] In some implementations, after entry expansion and / or vector expansion, that is, after a new language or field is added, an automated semantic verification module can be triggered to perform verification. Specifically, the automated semantic verification module can ensure semantic consistency between different languages through methods such as two-way translation verification and multilingual semantic comparison. Furthermore, the automated semantic verification module can also use context similarity matching technology to check whether the semantics of each first entry in different languages are accurately aligned to avoid deviations and mistranslations. In addition, the automated semantic verification module can also generate a verification report to record inconsistent entries, thereby triggering manual review or algorithm optimization to ensure the reliability of the knowledge base.
[0129] S205: Connect the multilingual knowledge base to the multilingual instant speech translation model.
[0130] In this way, the multilingual instant speech translation model can perform reasoning based on a multilingual knowledge base, which can improve translation accuracy.
[0131] In practical applications, a real-time knowledge base query interface can be constructed. By calling this interface, relevant information can be retrieved from the knowledge base. In this way, the multilingual instant speech translation model can perform inference based on the retrieval results, thereby achieving the purpose of connecting the multilingual knowledge base to the multilingual instant speech translation model.
[0132] In some implementations, the multilingual knowledge base may also be expanded based on the model training results, such as adding new languages and adding knowledge in new fields, etc., which is not specifically limited in the embodiments of the present application.
[0133] It should be noted that the multilingual knowledge base is not limited to being used in the inference phase of the multilingual instant speech translation model to achieve multilingual translation. The multilingual knowledge base can also be applied to the training phase of the multilingual instant speech translation model. During the training phase, the text vectors and audio vectors in the multilingual knowledge base can serve as additional supervisory signals to help the model better learn the semantic mapping relationship between different languages.
[0134] Further, the embodiment of the present application also provides a translation method based on a multilingual instant speech translation model to realize multilingual translation. The translation method can be run on a computing device, and the computing device is, for example, a translation terminal, and specifically can be a smart phone, a tablet computer, a special translation device, etc., which have voice processing and computing power. These devices can be equipped with a voice input interface, an audio processing chip, and a processor and memory for running the multilingual instant speech translation model, can receive the user's voice input in real time, and translate based on the translation method provided in the embodiment of the present application, and finally output the target language audio.
[0135] Figure 5A flowchart of a translation method based on a multilingual instant speech translation model provided in an embodiment of the present application.
[0136] like Figure 5 As shown, an embodiment of the present application provides a translation method based on a multilingual instant speech translation model. The multilingual instant speech translation model of the translation method can be trained based on the training method of the aforementioned multilingual instant speech translation model. The translation method can include the following steps S301-S306.
[0137] S301: Acquire audio to be translated, where the audio to be translated is in a single language or a mixture of multiple languages.
[0138] In practice, the audio to be translated may come from a variety of scenarios, including but not limited to recorded speeches in conferences, real-time voice conversations in international exchanges, and audio content from multimedia materials. Monolingual audio refers to audio information in only one language, such as an academic presentation in English or a news broadcast in Chinese. Multilingual audio can be audio that includes multiple languages, such as English statements interspersed with German technical terms. Multilingual audio includes but is not limited to Chinese, English, Malay, Thai, and Vietnamese.
[0139] For example, at the Global AI Technology Seminar, an engineer proposed: "When training this Transformer model, batch normalization can accelerate convergence, and using the PyTorch framework to build the network structure will be more efficient. At the same time, attention should be paid to distributed training on a GPU cluster. A / B testing should also be used to optimize hyperparameters to ensure the model's performance in NLP (natural language processing) tasks." This is clearly a multilingual audio track that mixes Chinese and English.
[0140] It is understandable that the computing device can be deployed with a front-end application, which is the interface for the computing device to interact with the user. The front-end application can be a translation software, voice assistant, etc. When the user has a translation demand, the audio acquisition can be triggered by the response operation of the front-end application. Furthermore, the front-end application can call the built-in microphone of the computing device to obtain audio. The user only needs to turn on the recording function of the front-end application and then clearly speak the content to be translated. After obtaining the audio to be translated, the front-end application can perform preliminary pre-processing on it. For example, the audio format is first checked and converted. If the obtained audio format does not meet the input requirements of the multilingual instant speech translation model, the front-end application can convert it into a suitable format, such as the common WAV format. At the same time, the front-end application can also perform noise reduction processing on the audio, remove interference factors such as environmental noise and current noise, improve the quality of the audio, and provide purer audio data for subsequent translation work.
[0141] S302: Convert the audio to be translated into source language text.
[0142] It is understandable that the step of converting the audio to be translated into the source language text is a speech-to-text step, which can be specifically implemented based on ASR technology.
[0143] S303: Convert the source language text into a source language text vector.
[0144] After converting the audio to be translated into source language text, embodiments of the present application can perform vector conversion on the source language text to obtain a source language text vector. This can be achieved through word embedding technology. In a specific implementation, to ensure the accuracy of knowledge retrieval, the source language text vector should maintain the same dimensionality as the text vector in the multilingual knowledge base.
[0145] S304: Search the source language text vector in the multilingual knowledge base to obtain a mapping relationship. The mapping relationship includes the source language text vector and its corresponding source language audio vector, as well as the text vector corresponding to the source language text in the target language.
[0146] It can be understood that a multilingual knowledge base is a database that stores text and audio information in multiple languages. The multilingual knowledge base can be constructed based on the aforementioned knowledge base construction method. The embodiment of the present application can perform vector retrieval in the multilingual knowledge base to obtain the text vector with the highest similarity to the source language text vector. This can be achieved through vector similarity calculation, for example, using cosine similarity to measure the similarity between the source language text vector and the text vector corresponding to the text unit stored in the multilingual knowledge base.
[0147] When a text vector with the highest similarity to the source language text vector is found in the multilingual knowledge base, related vectors in the entry where the text vector is located can be further extracted to obtain the source language audio vector (i.e., the audio vector corresponding to the source language text stored in the first entry), and the text vector corresponding to the source language text in the target language (i.e., the text vector in the target language corresponding to the source language text stored in the first entry).
[0148] In this way, the text information and audio information of the audio to be translated can be obtained, as well as the text information in the target language corresponding to the audio to be translated. In particular, it can be determined whether the audio to be translated contains professional domain knowledge, thereby ensuring the accuracy of the translation results in terms of language and professional content.
[0149] S305: forming prompt words based on the mapping relationship.
[0150] Specifically, the prompt word can include a source language text vector, a source language audio vector, and a target language text vector. These vector information can be organized in a specific format to form a multimodal prompt word, which can be constructed through preset rules or algorithms. The embodiments of this application do not make specific limitations on this.
[0151] S306: Input the audio to be translated and the prompt word into the multi-language instant speech translation model, so that the multi-language instant speech translation model translates the audio to be translated into the target language audio based on the prompt word. It is understandable that the target language is specifically specified by the user, for example, Chinese.
[0152] Furthermore, after receiving the audio to be translated and its corresponding prompt word, the multilingual instant speech translation model can parse the vector information in the prompt word, use the source language text vector and audio vector in the prompt word to enhance the understanding of the input audio, and refer to the target language text vector to guide the translation generation process. Ultimately, the target language audio can be obtained, realizing language-to-language (speech2speech) output.
[0153] It can be seen that the translation method provided in the embodiment of the present application does not include the steps of converting audio to text, translating into target language text, and converting target language text to audio. The multilingual instant voice translation model can directly realize the conversion of source language voice directly to target language voice, making the translation process smoother and more natural, suitable for real-time communication. In addition, users do not need to manually input text or perform complex operations, and can interact directly through voice, which enhances the naturalness and intuitiveness of language communication and achieves the effect of natural conversation. In addition, the multilingual instant voice translation model has strong multilingual adaptability and is suitable for different languages and dialects, especially multilingual mixed conference scenarios. It can realize real-time translation of the same audio containing different languages, which can enhance the user experience.
[0154] Furthermore, during the translation process, embodiments of the present application can also dynamically update the model's language knowledge and synchronize the new knowledge into the multilingual knowledge base, improving the coverage and accuracy of future translations. During this process, embodiments of the present application can optimize the knowledge base through a feedback mechanism combined with real-time interactive content.
[0155] Figure 6 A schematic diagram of the process of dynamically updating the knowledge base provided in an embodiment of the present application.
[0156] like Figure 6 As shown, the translation method provided in the embodiment of the present application may further include the following steps S307-S311.
[0157] S307: Obtaining an interaction log at a preset frequency. The interaction log includes the user's feedback on the target language audio, as well as the audio to be translated and / or the source language text involved in the interaction process.
[0158] Among them, it can be a back-end module of a computing device, and the back-end module can run on a server or a cloud platform. The embodiment of the present application does not make specific limitations on this.
[0159] Furthermore, the preset frequency is, for example, once a day or once a week. The interaction log can be generated in real time by the computing device based on the user's usage process, and can include structured stored data. Furthermore, the feedback results can be user reviews, error annotations, modification suggestions and other information collected through the front-end application. Among them, user reviews can be in the form of a combination of quantitative ratings (such as 1-5 stars) and text descriptions. For example, a user gives a 3-star rating and leaves a message "some technical terms are translated awkwardly"; the error annotation function can allow users to directly circle the problematic segments in the target language audio; modification suggestions can guide users to propose more accurate translation expressions, such as guiding users to propose "here 'elastic expansion of cloud computing' should be translated as 'Elastic Expansion of Cloud Computing'".
[0160] S308: Based on the feedback result, the audio to be translated and / or the source language text, determine whether the audio to be translated includes language and / or specific domain vocabulary not covered by the multilingual knowledge base.
[0161] After obtaining the interaction log, it can be analyzed. For example, natural language processing technology is used to perform semantic analysis on user feedback results to extract key descriptions of translation issues. For example, from the user's feedback that "a certain professional vocabulary was translated incorrectly", the specific content in the audio to be translated can be accurately located. Then, the audio to be translated, the source language text and the content in the multilingual knowledge base are compared. For example, it can be identified whether the audio to be translated contains small languages or dialects that are not covered by the knowledge base. At the vocabulary level, a vocabulary matching algorithm can be used to determine whether there are professional terms in specific fields, emerging vocabulary, and other content that is missing in the multilingual knowledge base. For example, in translation interactions in the medical field, if the audio to be translated contains the name of a new drug or cutting-edge treatment technology term, and there is no relevant record in the multilingual knowledge base, it can be identified that this is a specific field vocabulary that is not covered.
[0162] S309: When the audio to be translated includes languages and / or domain-specific vocabulary not covered by the multilingual knowledge base, obtaining extended texts of the uncovered languages and / or domain-specific vocabulary from a public data source.
[0163] During the retrieval process, web crawler technology can be used to automatically capture relevant text content, and through natural language processing text screening algorithms, extended text related to uncovered languages or specific field vocabulary can be accurately extracted from massive data to ensure that the information obtained is accurate and authoritative.
[0164] In some implementations, the extended text may also be obtained through LLM generation, which is not specifically limited in the embodiments of the present application.
[0165] S310: Determine a text extension vector and an audio extension vector based on the extended text.
[0166] The step of determining the text extension vector and the audio extension vector based on the text may refer to the aforementioned step S203 and will not be described in detail here.
[0167] S311: Adding a text extension vector and an audio extension vector to a vocabulary of a multilingual knowledge base based on the semantics of the extended text.
[0168] This allows for alignment between new languages or domains and existing languages, enabling dynamic updates to the multilingual knowledge base, thereby maintaining its timeliness and comprehensiveness. By combining user feedback and real-time interactive content, the multilingual knowledge base's coverage can be optimized, particularly by supplementing and improving low-resource languages, achieving continuous iterative updates.
[0169] In some implementations, the newly added knowledge may also be verified by calling an automated semantic verification module, which will not be described in detail here.
[0170] In some implementations, existing language knowledge can also be optimized, such as removing or correcting infrequently used invalid entries to ensure the timeliness of the multilingual knowledge base.
[0171] In some implementations, the embodiments of the present application may also establish a knowledge base version management mechanism to record changes in each update to facilitate backtracking and adjustment, which is not specifically limited in the embodiments of the present application.
[0172] From the above content, it can be seen that the embodiments of the present application provide a training method for a multilingual instant speech translation model, a knowledge base construction method, and a translation method based on a multilingual instant speech translation model. The training method can be customized for specific fields or application scenarios by fine-tuning the pre-trained language model (multimodal basic model), that is, performing incremental training to optimize the performance of the model in a specific context. In addition, the method can improve the model's ability to understand specific contexts, industry terms, and special expressions through careful adjustment of training data, thereby improving translation accuracy and naturalness, helping to process complex conversation content or proper nouns in specific fields, and ensuring that the translation results are more accurate and in line with the context.
[0173] In addition, the knowledge base construction method can build a professional knowledge base that includes multiple languages, integrating professional terms and expressions in various languages in different fields (such as medicine, law, technology, etc.). Through this multilingual knowledge base, the translation method can provide customized translation services based on the linguistic characteristics and industry needs of different languages, enhancing the depth and accuracy of cross-language communication. It not only optimizes the quality of machine translation, but also enhances the application effect in multilingual and multi-field environments, and can more accurately adapt to global communication needs. It improves the intelligence level of the cross-language real-time translation system, and can achieve more efficient and more accurate translation services, especially in cross-border conferences and multilingual business processing, which has significant application value.
[0174] Figure 7 A schematic diagram of the structure of a training device for a multilingual instant speech translation model provided in an embodiment of the present application.
[0175] like Figure 7 As shown, an embodiment of the present application provides a training device for a multilingual instant speech translation model, the device comprising:
[0176] The corpus conversion module 1101 is configured to: extract speech features from the audio corpus, and extract text features from the text corpus corresponding to the audio corpus; wherein the audio corpus and its corresponding text corpus are divided into semantic categories, and the same category contains corpora corresponding to multiple different languages; the speech features are vector representations of the audio corpus, and the text features are vector representations of the text corpus.
[0177] The cross-modal fusion module 1102 is configured to: fuse the speech features and text features that represent the same semantics in the same language to obtain fused features corresponding to each semantics.
[0178] The vector mapping module 1103 is configured to: map the fused features to the same preset semantic space respectively to obtain the multimodal features corresponding to each fused feature.
[0179] The iterative training module 1104 is configured to train a preset multimodal basic model using multimodal features as training data to obtain a multilingual instant speech translation model.
[0180] In a possible implementation, the vector mapping module 1103 is further configured to align multimodal features representing the same semantics in different languages using a contrastive loss function.
[0181] In one possible implementation, the iterative training module 1104 is further configured to: obtain newly added text corpus, where the newly added text corpus corresponds to a different language and / or corresponds to a different field than the text corpus; obtain newly added audio corpus corresponding to the newly added text corpus; and perform incremental training on the multilingual instant speech translation model based on the newly added text corpus and the newly added audio corpus to update some parameters of the multilingual instant speech translation model.
[0182] Figure 8 A schematic diagram of the structure of the knowledge base construction device provided in an embodiment of the present application.
[0183] like Figure 8 As shown, the embodiment of the present application also provides a knowledge base construction device, which includes:
[0184] The first construction module 2101 is configured to: construct a first language knowledge base; wherein the first language knowledge base includes a vocabulary consisting of a plurality of first entries, each first entry includes a text unit and its corresponding text vector and audio vector, the text vector has the same dimension as the text feature, the audio vector has the same dimension as the speech feature, and the text vector and text feature that express the same semantics are the same, and the audio vector and speech feature that express the same semantics are the same, and the text unit is the granular text corresponding to the first language;
[0185] The expansion module 2102 is configured to: expand the first entries using the large language model (LLM) to obtain a text expansion vector and an audio expansion vector corresponding to each first entry; the text expansion vector and the audio expansion vector are vectors corresponding to the text unit in at least one second language; and add the text expansion vector and the audio expansion vector to the corresponding first entry to form a multilingual knowledge base.
[0186] In one possible implementation, the expansion module 2102 is specifically configured to: generate an extended text corresponding to a text unit using LLM, the extended text including corresponding text content in at least one second language; and / or obtain target text content corresponding to each text unit in at least one second language from a public data source; wherein the target text content is located in paragraph content; and extract the target text content from the paragraph content using LLM to obtain the extended text.
[0187] In a possible implementation, the expansion module 2102 is further configured to: generate extended audio corresponding to the extended text using an audio generation model; convert the extended text into a text extension vector; and convert the extended audio into an audio extension vector.
[0188] In a possible implementation, the expansion module 2102 is further configured to: utilize the LLM to expand the vocabulary to form first entries corresponding to other text units in the first language, where the other text units belong to the same or different fields as the text unit.
[0189] Figure 9 A schematic diagram of the structure of a translation device based on a multilingual instant speech translation model provided in an embodiment of the present application.
[0190] like Figure 9 As shown, an embodiment of the present application provides a translation device based on a multilingual instant speech translation model. The multilingual instant speech translation model can be trained based on the aforementioned training method. The device includes:
[0191] The acquisition module 3101 is configured to: acquire audio to be translated, where the audio to be translated is in a single language or a mixture of multiple languages;
[0192] The translation module 3102 is configured to: convert the audio to be translated into source language text; convert the source language text into a source language text vector; search the source language text vector in a multilingual knowledge base to obtain a mapping relationship; wherein the multilingual knowledge base is constructed based on the knowledge base construction method in the aforementioned embodiment, and the mapping relationship includes the source language text vector and its corresponding source language audio vector, as well as the text vector corresponding to the source language text in the target language; form a prompt word based on the mapping relationship; input the audio to be translated and the prompt word into the multilingual instant speech translation model, so that the multilingual instant speech translation model translates the audio to be translated into the target language audio based on the prompt word.
[0193] In one possible implementation, the device also includes a knowledge base update module 3103, which is specifically configured to: obtain interaction logs at a preset frequency, wherein the interaction logs include user feedback results on the target language audio, and also include the audio to be translated and / or source language text involved in the interaction process; based on the feedback results, the audio to be translated and / or the source language text, determine whether the audio to be translated includes languages and / or specific domain vocabulary not covered by the multilingual knowledge base; if the audio to be translated includes languages and / or specific domain vocabulary not covered by the multilingual knowledge base, obtain extended text of the uncovered languages and / or extended text of the specific domain from a public data source; determine a text extension vector and an audio extension vector based on the extended text; and add the text extension vector and the audio extension vector to the vocabulary of the multilingual knowledge base based on the semantics of the extended text.
[0194] Figure 10 A schematic diagram of a computing device provided in an embodiment of the present application.
[0195] like Figure 10 As shown, the computing device may include a server, a terminal, or other devices; the computing device includes one or more processors 1201 and a memory 1202. The memory 1202 is configured to store one or more programs. When the one or more programs are executed by the one or more processors 1201, the one or more processors 1201 implement the training method, knowledge base construction method, or translation method based on the multilingual instant speech translation model in the above-mentioned embodiments.
[0196] Continue to see Figure 10 The computing device 1200 may further include: a communication interface 1203 and a communication bus 1204 .
[0197] The processor 1201, the memory 1202 and the communication interface 1203 communicate with each other via a communication bus 1204. The communication interface 1203 is used to communicate with other devices such as a client or other server network elements.
[0198] In some embodiments, one or more processors 1201 are configured to execute one or more programs 1205, which may specifically execute the relevant steps in the aforementioned multilingual instant speech translation model training method, knowledge base construction method, or translation method based on the multilingual instant speech translation model. Specifically, program 1205 may include program code, which may include computer-executable instructions.
[0199] For example, the processor 1201 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. The computing device 1200 may include one or more processors of the same type, such as one or more CPUs, or different types of processors, such as one or more CPUs and one or more ASICs.
[0200] In some embodiments, the memory 1202 is used to store one or more programs 1205. The memory 1202 may include a high-speed RAM memory, and may also include a non-volatile memory (NVM), such as at least one disk storage.
[0201] The program 1205 can be specifically called by the processor 1201 to enable the computing device 1200 to execute a training method for a multilingual instant speech translation model, a knowledge base construction method, or a translation method operation based on a multilingual instant speech translation model.
[0202] Some embodiments of the present application provide a computer-readable storage medium storing at least one executable instruction. When the executable instruction is executed on a computing device 1200, the computing device 1200 executes the training method for a multilingual instant speech translation model, the knowledge base construction method, or the translation method based on the multilingual instant speech translation model in the above-mentioned embodiments.
[0203] The executable instructions can be specifically used to enable the computing device 1200 to execute a training method for a multilingual instant speech translation model, a knowledge base construction method, or a translation method operation based on a multilingual instant speech translation model.
[0204] For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0205] The beneficial effects that can be achieved by the readable storage medium provided in some embodiments of the present application can refer to the beneficial effects of the corresponding multilingual instant speech translation model training method, knowledge base construction method or translation method based on the multilingual instant speech translation model provided above, and will not be repeated here.
[0206] It should be noted that, in the application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0207] Each embodiment in this specification is described in a related manner. Similar portions between the embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so their description is relatively simple. For related portions, refer to the description of the method embodiments.
[0208] The logic and / or steps represented in the flowchart or otherwise described herein may be considered, for example, as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device).
[0209] For the purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, apparatus, or device.
[0210] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic device, and a portable compact disc read-only memory (CDROM).
[0211] In addition, the computer readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in the computer memory. It should be understood that various parts of the present application can be implemented in hardware, software, firmware or a combination thereof.
[0212] In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0213] The above implementation methods are only specific implementation methods of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present application should be included in the scope of protection of the present application.
Claims
1. A training method for a multilingual instant speech translation model, characterized in that: The method comprises: Extracting speech features from an audio corpus, and extracting text features from a text corpus corresponding to the audio corpus; wherein the audio corpus and the corresponding text corpus are divided into semantic categories, and the same category includes corpora corresponding to multiple different languages; the speech features are vector representations of the audio corpus, and the text features are vector representations of the text corpus; Fusing the speech features and text features representing the same semantics in the same language to obtain fused features corresponding to the respective semantics; Mapping the fused features to the same preset semantic space respectively to obtain multimodal features corresponding to each fused feature; The multimodal features are used as training data to train a preset multimodal basic model to obtain the multilingual instant speech translation model.
2. The training method of the multilingual instant speech translation model according to claim 1, characterized in that: After fusing the speech features and the text features that represent the same semantics in the same language to obtain fused features, the method further includes: The multimodal features representing the same semantics in different languages are aligned using a contrast loss function.
3. The training method of the multilingual instant speech translation model according to claim 1, wherein: The method further comprises: Acquire a new text corpus, where the new text corpus corresponds to a different language and / or a different field than the text corpus; Obtaining the newly added audio corpus corresponding to the newly added text corpus; Incremental training is performed on the multilingual instant speech translation model based on the newly added text corpus and the newly added audio corpus to update some parameters of the multilingual instant speech translation model.
4. A knowledge base construction method, characterized in that: A multilingual instant speech translation model trained by the training method according to any one of claims 1 to 3, the method comprising: Constructing a first language knowledge base; wherein the first language knowledge base includes a vocabulary consisting of a plurality of first entries, each of the first entries includes a text unit and its corresponding text vector and audio vector, the text vector has the same dimension as the text feature, the audio vector has the same dimension as the speech feature, and the text vector and the text feature that express the same semantics are the same, and the audio vector and the speech feature that express the same semantics are the same, and the text unit is a granular text corresponding to the first language; Expanding the first entries using a large language model (LLM) to obtain a text expansion vector and an audio expansion vector corresponding to each of the first entries; the text expansion vector and the audio expansion vector are vectors corresponding to the text unit in at least one second language; The text extension vector and the audio extension vector are added to the corresponding first entry to form a multilingual knowledge base.
5. The knowledge base construction method according to claim 4, characterized in that: The expanding the first entry by using a large language model (LLM) includes: generating an extended text corresponding to the text unit using the LLM, the extended text including at least one corresponding text content in the second language; and / or, Obtaining target text content corresponding to each of the text units in at least one of the second languages from a public data source; wherein the target text content is located in a paragraph content; The target text content is extracted from the paragraph content using the LLM to obtain an extended text.
6. The knowledge base construction method according to claim 5, characterized in that: After generating the extended text corresponding to the text unit by using the LLM, and / or after obtaining the extended text, the method further includes: Generate extended audio corresponding to the extended text using an audio generation model; The extended text is converted into the text extension vector, and the extended audio is converted into the audio extension vector.
7. The knowledge base construction method according to claim 4, characterized in that: Before expanding the first entry using the large language model LLM, the method further includes: The vocabulary is expanded using the LLM to form the first entries corresponding to other text units in the first language, where the other text units belong to the same or different fields as the text unit.
8. A translation method based on a multilingual instant speech translation model, characterized in that: The multilingual instant speech translation model is obtained by training based on the training method according to any one of claims 1 to 3, and the translation method includes: Obtaining audio to be translated, where the audio to be translated is in a single language or a mixture of multiple languages; Converting the audio to be translated into source language text; Converting the source language text into a source language text vector; Searching the source language text vector in a multilingual knowledge base to obtain a mapping relationship; wherein the multilingual knowledge base is constructed based on the knowledge base construction method according to any one of claims 4 to 7, and the mapping relationship includes the source language text vector and its corresponding source language audio vector, and also includes a text vector corresponding to the source language text in the target language; forming a prompt word based on the mapping relationship; The audio to be translated and the prompt word are input into the multilingual instant speech translation model, so that the multilingual instant speech translation model translates the audio to be translated into a target language audio based on the prompt word.
9. The translation method based on the multilingual instant speech translation model according to claim 8, characterized in that: The method further comprises: Obtaining an interaction log at a preset frequency, wherein the interaction log includes a user's feedback on the target language audio, and also includes the audio to be translated and / or the source language text involved in the interaction process; Based on the feedback result, the audio to be translated and / or the source language text, determining whether the audio to be translated includes language and / or domain-specific vocabulary not covered by the multilingual knowledge base; When the audio to be translated includes a language and / or a vocabulary in a specific field that is not covered by the multilingual knowledge base, obtaining extended text in the uncovered language and / or the extended text in the specific field from a public data source; Determining a text extension vector and an audio extension vector based on the extended text; The text extension vector and the audio extension vector are added to the vocabulary of the multilingual knowledge base based on the semantics of the extended text.
10. A computing device, characterized in that include: one or more processors; as well as, a memory configured to: store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the multilingual instant speech translation model according to any one of claims 1-3, or implement the knowledge base construction method according to any one of claims 4-7, or implement the translation method based on the multilingual instant speech translation model according to any one of claims 8-9.