Fault-tolerant translation method, method and device for training fault-tolerant translation model

By identifying and fusing multiple candidate statements and using the fault-tolerant translation model of the self-attention mechanism for translation, the problem of error propagation in the cascaded translation system is solved, and the error tolerance and real-time nature of translation is improved.

CN114492467BActive Publication Date: 2025-05-13UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111675437.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-13
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

In machine translation scenarios, especially in cascading translation systems, there are error propagation problems, which affect the accuracy of translation and require more robust fault-tolerant translation.

Method used

By identifying multiple candidate statements of the source language statement, fusion candidate statements are determined, and the fusion candidate statements are encoded, pre-selected and decoded using a fault-tolerant translation model with a self-attention mechanism to obtain the target language statement.

Benefits of technology

Improves the error tolerance and real-time nature of translation. The fault-tolerant translation model can learn the differential parts of multiple candidate statements in a word-level granularity, reducing the loss of computing resources and memory resources required for semantic representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492467B_ABST
    Figure CN114492467B_ABST
Patent Text Reader

Abstract

The present application provides a method for error-tolerant translation, a method for training an error-tolerant translation model, and an apparatus. The error-tolerant translation method comprises: identifying multiple candidate sentences corresponding to a source language sentence; determining a fused candidate sentence based on the multiple candidate sentences, the fused candidate sentence comprising multiple candidate words corresponding to the first word in the source language sentence and label information for indicating the multiple candidate words; encoding, pre-selecting, and decoding the fused candidate sentence using an error-tolerant translation model to obtain a target language sentence. The present application can improve the error tolerance and real-time performance of translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing, and in particular to a method and device for training an error-tolerant translation model. Background Art

[0002] At present, the widespread application of deep learning technology and the large amount of data generated by the development of the Internet have brought great opportunities for the development of natural language processing. With the update of machine hardware and the introduction of relevant optimization algorithms, the model can use a large amount of corpus for learning, making many tasks in the field of natural language processing such as text recognition and machine translation have achieved breakthrough results.

[0003] However, in machine translation scenarios, especially in cascade translation systems, there is often the problem of error propagation, which directly affects the accuracy of the translation system and requires the translation system to be able to achieve more robust and fault-tolerant translation. Summary of the invention

[0004] In view of this, the embodiments of the present application provide a method for error-tolerant translation, a method for training an error-tolerant translation model, and a device, which can improve the error tolerance and real-time performance of translation.

[0005] In a first aspect, an embodiment of the present application provides a fault-tolerant translation method, comprising: identifying multiple candidate sentences corresponding to a source language sentence; determining a fused candidate sentence based on the multiple candidate sentences, the fused candidate sentence comprising multiple candidate words corresponding to the first word in the source language sentence and label information for indicating the multiple candidate words; encoding, pre-selecting and decoding the fused candidate sentences using a fault-tolerant translation model to obtain a target language sentence.

[0006] In certain embodiments of the present application, a fused candidate sentence is determined based on multiple candidate sentences, including: sorting the multiple candidate words according to their recognition score priorities, embedding the multiple candidate words as a sequence in the fused candidate sentence, wherein the position of the sequence in the fused candidate sentence is the same as the position of the first word in the source language sentence.

[0007] In certain embodiments of the present application, the error-tolerant translation method is a simultaneous interpretation method, and the recognition score includes an acoustic recognition score.

[0008] In certain embodiments of the present application, the fault-tolerant translation model includes an encoder with a self-attention mechanism and a decoder with a self-attention mechanism, and the fault-tolerant translation model is used to encode and decode the fused candidate sentences to obtain the target language sentences, including: using the encoder to encode and pre-select the fused candidate sentences to obtain the semantic representation of the fused candidate sentences; using the decoder to decode the semantic representation to obtain the target language sentence.

[0009] In certain embodiments of the present application, the encoder and the decoder are jointly trained based on a first loss function and a second loss function, wherein the first loss function is set for the task of semantically matching multiple candidate words based on the context information of the fused candidate sentences, and the second loss function is set for the machine translation task.

[0010] In some embodiments of the present application, the first loss function is expressed by the following formula:

[0011]

[0012] Among them, T represents the number of multiple candidate intervals in the sentence, K j represents the number of multiple candidate words in the jth candidate interval, and They represent the correct and i-th multiple candidate words or phrases in the j-th candidate interval respectively.

[0013] In certain embodiments of the present application, the error-tolerant translation method of the first aspect further includes: randomly selecting a second word in a sample sentence, and replacing the second word in the sample sentence position with a neighbor candidate of the second word in a candidate dictionary to generate a candidate training corpus for an error-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by frequency of occurrence in speech recognition.

[0014] In certain embodiments of the present application, the label information includes at least one of start information, end information and intermediate information, wherein the start information is used to indicate the starting position of multiple candidate words in the fused candidate sentence, the end information is used to indicate the ending position of multiple candidate words in the fused candidate sentence, and the intermediate information is used to indicate the intermediate position of multiple candidate words in the fused candidate sentence.

[0015] In some embodiments of the present application, determining a fused candidate sentence based on multiple candidate sentences includes: using a longest common substring algorithm to determine a fused candidate sentence based on the multiple candidate sentences.

[0016] In some embodiments of the present application, identifying multiple candidate sentences of a source language sentence includes: performing speech recognition on a speech signal of the source language sentence to obtain multiple candidate sentences of the source language sentence.

[0017] In certain embodiments of the present application, the fused candidate sentence also includes: at least one third word, each of the multiple candidate sentences includes at least one third word, and the position of the at least one third word in the fused candidate sentence is the same as the position of the at least one third word in the source language sentence.

[0018] In certain embodiments of the present application, the error-tolerant translation model includes a neural network model with a self-attention mechanism.

[0019] In some embodiments of the present application, the neural network model includes a transformer model.

[0020] In a second aspect, a method for training a fault-tolerant translation model is provided, comprising: identifying multiple candidate sentences corresponding to a source language sentence sample; determining a fused candidate sentence based on the multiple candidate sentences, the fused candidate sentence comprising multiple candidate words corresponding to a first word in the source language sentence sample and label information for indicating the multiple candidate words; encoding and pre-selecting the fused candidate sentence using an encoder to obtain a semantic representation of the fused candidate sentence, and determining a first loss function based on the semantic representation; decoding the semantic representation using a decoder to obtain a target language sentence, and determining a second loss function based on the target language sentence; jointly training the encoder and the decoder based on the first loss function and the second loss function to obtain a fault-tolerant translation model.

[0021] In some embodiments of the present application, the first loss function is expressed by the following formula:

[0022]

[0023] Among them, T represents the number of multiple candidate intervals in the sentence, K j represents the number of multiple candidate words in the jth candidate interval, and They represent the correct and i-th multiple candidate words or phrases in the j-th candidate interval respectively.

[0024] In certain embodiments of the present application, the method of the second aspect also includes: randomly selecting a second word in the sample sentence, and replacing the second word in the sample sentence position with a neighbor candidate of the second word in a candidate dictionary to generate a candidate training corpus for a fault-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by frequency of occurrence in speech recognition.

[0025] In a third aspect, a fault-tolerant translation device is provided, comprising: an identification module for identifying multiple candidate sentences corresponding to a source language sentence; a determination module for determining a fused candidate sentence based on the multiple candidate sentences, the fused candidate sentence comprising multiple candidate words corresponding to the first word in the source language sentence and label information for indicating the multiple candidate words; and a coding and decoding module for encoding, pre-selecting and decoding the fused candidate sentence using a fault-tolerant translation model to obtain a target language sentence.

[0026] In a fourth aspect, a device for training a fault-tolerant translation model is provided, comprising: an identification module for identifying multiple candidate sentences corresponding to a source language sentence sample; a determination module for determining a fused candidate sentence based on the multiple candidate sentences, the fused candidate sentence comprising multiple candidate words corresponding to the first word in the source language sentence sample and label information for indicating the multiple candidate words; an encoding and decoding module for encoding and pre-selecting the fused candidate sentence using an encoder to obtain a semantic representation of the fused candidate sentence, determining a first loss function based on the semantic representation, decoding the semantic representation using a decoder to obtain a target language sentence, and determining a second loss function based on the target language sentence; a training module for jointly training the encoder and the decoder based on the first loss function and the second loss function to obtain a fault-tolerant translation model.

[0027] In a fifth aspect, an electronic device is provided, comprising: a processor; and a memory for storing processor executable instructions, wherein the processor is used to execute the error-tolerant translation method described in the method of the first aspect and the method for training an error-tolerant translation model described in the method of the second aspect.

[0028] In a sixth aspect, a storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the fault-tolerant translation method described in the method of the first aspect and the method of the second aspect.

[0029] According to the technical solution of the present application, by identifying multiple candidate sentences corresponding to the source language sentence, determining the fused candidate sentence according to the multiple candidate sentences, and using the fault-tolerant translation model to encode, pre-select and decode the fused candidate sentence, the target language sentence is obtained, which can improve the fault tolerance and real-time performance of the translation. Since the fault-tolerant translation model can learn the difference parts of the identified multiple candidate sentences in a targeted manner at the word-level granularity, it is helpful to improve the accuracy of fault-tolerant translation. In addition, since the fusion of multiple candidate sentences can reduce the loss of computing resources and memory resources required for the subsequent semantic representation of the sentence, the real-time performance of the translation is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Shown is a schematic diagram of the system architecture of a simultaneous interpretation system provided by an exemplary embodiment of the present application.

[0031] Figure 2 Shown is a flow chart of a fault-tolerant translation method provided by an exemplary embodiment of the present application.

[0032] Figure 3 Shown is a flow chart of a fault-tolerant translation method provided by another exemplary embodiment of the present application.

[0033] Figure 4AShown is a schematic diagram of a process of fusing multiple candidate sentences provided by an exemplary embodiment of the present application.

[0034] Figure 4B Shown is a schematic diagram of a process of fusing multiple candidate sentences provided by another exemplary embodiment of the present application.

[0035] Figure 5 Shown is a flowchart of a method for training an error-tolerant translation model provided by an exemplary embodiment of the present application.

[0036] Figure 6 Shown is a schematic diagram of the structure of a fault-tolerant translation system provided by another exemplary embodiment of the present application.

[0037] Figure 7 Shown is a schematic diagram of the structure of a fault-tolerant translation system provided by another exemplary embodiment of the present application.

[0038] Figure 8 Shown is a schematic diagram of the structure of a fault-tolerant translation device provided by an exemplary embodiment of the present application.

[0039] Fig. 9 Shown is a schematic diagram of the structure of an apparatus for training a fault-tolerant translation model provided by an exemplary embodiment of the present application.

[0040] Fig.10 Shown is a block diagram of an electronic device for executing an error-tolerant translation method or a method for training an error-tolerant translation model provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0042] Application Overview

[0043] Simultaneous interpretation, also known as simultaneous voice interpretation or simultaneous interpretation for short, is a translation method in which the translator continuously interprets the speech content from one language into another language for the audience without interrupting the speaker. In order to ensure the effectiveness of simultaneous interpretation, it is necessary to balance real-time performance and accuracy.

[0044] Due to the increasing development of recognition and translation technology, simultaneous interpreters are gradually being replaced by machines, and the cascaded voice simultaneous interpretation system that has emerged has gradually been applied in real scenarios. In the cascaded voice simultaneous interpretation system, the speaker's voice is first converted into text through a voice recognition device, and then the text is converted into the corresponding translation through a machine translation device, and finally displayed on the screen or converted into voice and played through an audio output device.

[0045] However, in actual application scenarios, there are many problems such as loud noise and accents. Speech recognition devices are bound to have recognition errors, and the cascaded speech interpretation system has the problem of error propagation, which will directly affect the accuracy of the machine translation device. Therefore, in the cascaded speech interpretation system, it is necessary to be able to achieve more robust error-tolerant translation in the case of recognition errors, which requires not only high translation quality of the machine translation device, but also fast response. Therefore, it is urgent to provide a lightweight recognition and error-tolerant translation technology.

[0046] In order to achieve error-tolerant translation of simultaneous speech, an additional error correction module can be introduced between the speech recognition device and the machine translation device, or more robust training can be performed on the input text at the translation end.

[0047] The former solution first corrects the speech recognition results by introducing an additional error correction module between the recognition module and the translation module, and then sends them to the machine translation device for translation. This often requires a separate recognition error correction model to be trained or relevant rules to be formulated manually. However, since it is difficult for artificially formulated rules to cover all aspects, they may have certain utility in certain fields or scenarios, but the effect is poor after being transferred to other fields and scenarios. Therefore, they lack accuracy and generalization in simultaneous interpretation. In addition, when the recognition system has certain changes and updates, the relevant rules need to be re-formulated, and the update cost is relatively high. However, there is a problem of error propagation in the cascaded speech simultaneous interpretation system, which may face the risk that the original recognition result is correct, but then it will be wrong after passing through the error correction module. In addition, the additional introduction of the error correction module will lead to an increase in computing and storage resources.

[0048] The latter solution improves the robustness of the recognition input text by constructing a method of constructing pairs of erroneous source text and correct translation sentences and adding mask information to the translation model. The purpose is that even when there is a recognition error in the speaker's speech, the machine translation device can still give a translation result that conforms to the speaker's original intention. In this solution, since there is no clear text information that needs to be corrected in the input text, it is difficult for the translation model to strike a balance between translation fidelity and error correction strength; and this solution lacks the use of the output information of the upstream recognition system. Since it only models the text channel, it is easy to make the error correction more biased towards the semantic level, while ignoring the use of the speaker's acoustic information.

[0049] In order to solve the above problems, the embodiments of the present application provide a method and device for fault-tolerant translation. The technical solution of the present application is described in detail below in conjunction with the accompanying drawings.

[0050] Exemplary Systems

[0051] Figure 1 The system architecture diagram of a simultaneous interpretation system 100 provided by an exemplary embodiment of the present application is shown, which shows an application scenario of using a machine translation device to implement simultaneous interpretation. The simultaneous interpretation system 100 may include a machine translation device 110, a speech recognition device 120, a speech collection device 130, and an output device 140. The speech collection device 130 is connected to the speech recognition device 120, and the machine translation device 110 is connected between the speech recognition device 120 and the output device 140.

[0052] In one embodiment, the voice collection device 130 can be used to collect the user's voice data. The voice recognition device 120 is used to perform voice recognition on the voice data and convert it into text data. The machine translation device 110 is used to convert the text data into a translated text in another language using machine translation technology. The output device 140 can be a display device or an audio output device for presenting the translated text on the screen or converting it into audio and playing it by the audio output device.

[0053] Specifically, in the simultaneous interpretation system 100, the speaker's voice is first collected by the voice collection device 140, then recognized and converted into text by the voice recognition device 120, and then the text is converted into a corresponding translation by the machine translation device 110, and finally output by the output device 140, for example, displayed on the screen or converted into voice and played through an audio output device.

[0054] It should be understood that the above-mentioned devices can be integrated on the same computing device. For example, the computing device can be a user terminal such as a PC, a mobile terminal, a translation pen, a tablet computer, or a server. The present application does not limit the implementation form of the above-mentioned devices. The above-mentioned devices can also be implemented as independent devices. For example, the voice acquisition device 130 can be a voice input device such as a microphone. The output device 140 can be a display screen or an audio output device (for example, a speaker), and the machine translation device 110 or the voice recognition device 120 can be a server.

[0055] It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principle of the present application, and the embodiments of the present application are not limited thereto. On the contrary, the embodiments of the present application can be applied to any scenario that may be applicable.

[0056] Exemplary Methods

[0057] Figure 2The figure shows a schematic flowchart of a fault-tolerant translation method provided by an exemplary embodiment of the present application. Figure 2 The method can be executed by a computing device (e.g., a server). As Figure 2 shown, the fault-tolerant translation method includes the following.

[0058] 210: Identify multiple candidate sentences corresponding to the source language sentence.

[0059] Specifically, identifying the source language sentence can be (e.g., in the simultaneous interpretation scenario) from speech data, or can be (e.g., in the text translation scenario) from a text file. The source language can be English, or can be Chinese or other languages, and the embodiments of the present application do not limit this.

[0060] The so-called candidate sentences can refer to multiple possible recognition results that may occur after the source language is recognized. For example, in the simultaneous interpretation scenario, if the speech content is: "5G has brought business growth to China Mobile", since in Chinese, the pronunciation of 5G, "wuji", is similar to the pronunciation of weapon, "wuqi", confusion may occur during speech recognition and 5G may be misrecognized as weapon. Then the candidate sentences may include "Weapon has brought business growth to China Mobile" and "5G has brought business growth to China Mobile".

[0061] It should be noted that although the present application is described by taking the simultaneous interpretation scenario or the text translation scenario as examples, the embodiments of the present application can be applied to other similar scenarios.

[0062] 220: Determine a fused candidate sentence according to the multiple candidate sentences. The fused candidate sentence includes multiple candidate words corresponding to the first word in the source language sentence and label information for indicating the multiple candidate words.

[0063] Specifically, after identifying multiple candidate sentences, these candidate sentences can be fused to obtain a fused candidate sentence. The so-called fusion means fusing multiple candidate sentences into one candidate sentence, and at the same time retaining at least part or all of the information of the multiple candidate sentences in this candidate sentence. Since the multiple candidate sentences are for the same source language sentence, there may be sentence components in these candidate sentences that are the same as those in the source language sentence, or there may be different sentence components. For the same words in the same position in these candidate sentences, only one needs to be retained. For the different words in these candidate sentences, as candidate words, all can be retained. It should be understood that only a part can also be retained. For example, n candidate words (n < m) are retained from m candidate sentences. Still taking the above speech content as an example, the fused candidate sentence can be "‘5G’‘Weapon’ has brought business growth to China Mobile", where 5G and weapon are candidate words.

[0064] It should be understood that a source language sentence may have candidate words at one or more positions, that is, multiple candidate sentences may have candidate words at one position or at multiple positions.

[0065] The so-called word can refer to the smallest language unit that can be used independently, or it can refer to a fragment in a sentence. For example, in English, it can refer to a word (word) or a phrase (phase), and in Chinese, it can refer to a word or a phrase. The source language can be Chinese, English or other languages, and the embodiments of this application are not limited to this. Taking the above-mentioned voice content as an example, the word can be "5G", "weapons", "China Mobile" or "brought", etc.

[0066] The tag information can not only be used as an indicator to indicate that there are candidate words in the fused candidate sentence, but also as a delimiter to indicate the position of the candidate words. Specifically, the tag information can be inserted before or after the candidate word, and the tag information can be a word or a letter, or a numerical value. When the tag information is a numerical value, it can be used to indicate that the number of words corresponding to the numerical value before or after the information is the candidate word.

[0067] 230: Encode, pre-select and decode the fusion candidate sentences using the error-tolerant translation model to obtain the target language sentences.

[0068] Specifically, the fault-tolerant translation model includes a neural network model with a self-attention mechanism, for example, the neural network model includes a transformer model. The fault-tolerant translation model is used to perform translation tasks. When a fused candidate sentence with label information and candidate words is input into the fault-tolerant translation model, the fault-tolerant translation model can implicitly realize selective translation of multiple candidate words based on information in two channels, the source language and the target language. For example, the fault-tolerant translation model can utilize the cross-attention mechanism and historical decoding information, improve the fault-tolerant ability of the fault-tolerant translation model to identify multiple candidate words by introducing target-side information, and preliminarily screen the information of multiple candidate words based on the fault-tolerant encoding module, and then further use the fault-tolerant decoding module to implicitly select based on the preliminary screening information.

[0069] The embodiment of the present application first merges multiple identified candidate sentences into one sentence based on word granularity, seeks common ground while reserving differences, retains the same sentence components in multiple candidate sentences, converts multiple candidate words into sequences according to the priority of acoustic recognition scores for different multiple candidate words in the sentence, and uses special symbols to separate and distinguish different multiple candidate words; for example, for multiple candidate sentences: "5g has brought business growth to China Mobile" and "weapons have brought business growth to China Mobile", the fused sentence is "[start]5g[mid]weapons[end] has brought business growth to China Mobile". Secondly, the encoder is used to score the multiple candidate words in the fused candidate sentence, thereby achieving the purpose of pre-selecting multiple candidate words on the encoder side. Finally, the decoding module can also learn the attention weights of different candidate words based on the information in the parallel sentence pairs, and combine the pre-selection information of multiple candidate words in the encoder to achieve implicit selection and translation of multiple candidate words on the decoder side.

[0070] According to an embodiment of the present application, by identifying multiple candidate sentences corresponding to a source language sentence, determining a fused candidate sentence according to multiple candidate sentences, and utilizing a fault-tolerant translation model to encode, pre-select and decode the fused candidate sentence, a target language sentence is obtained, and the fault tolerance rate and real-time performance of the translation can be improved. Since the fault-tolerant translation model can learn the difference parts of the identified multiple candidate sentences in a targeted manner with word-level granularity, it is helpful to improve the accuracy of fault-tolerant translation. In addition, since the fusion of multiple candidate sentences can reduce the subsequent loss of computing resources and memory resources required for semantic representation of sentences, the real-time performance of the translation is improved.

[0071] In some embodiments, multiple candidate words are arranged in a preset order in the fused candidate sentence, for example, they may be arranged in order of priority, or in other embodiments, multiple candidate words may be randomly distributed.

[0072] Compared with the data enhancement scheme using the translation model, due to the lack of acoustic information, error correction often tends to be biased towards the semantic level, while ignoring the use of the speaker's acoustic information. For example, for the two sentences "5g is developing very fast" and "weapons are developing very fast", it is impossible to distinguish them only from the perspective of the language model. At this time, it is necessary to rely on the recognition system to identify the speaker's acoustic score (or rating). For example, the candidate word with a high acoustic score is set at the center position or close to the center position, that is, the higher the score, the closer it is to the middle position, that is, the use of acoustic information.

[0073] In certain embodiments of the present application, a fused candidate sentence is determined based on multiple candidate sentences, including: sorting the multiple candidate words according to their recognition score priorities, embedding the multiple candidate words as a sequence in the fused candidate sentence, wherein the position of the sequence in the fused candidate sentence is the same as the position of the corresponding word in the source language sentence.

[0074] In certain embodiments of the present application, the error-tolerant translation method is a simultaneous interpretation method, and the recognition score includes an acoustic recognition score.

[0075] For example, in a simultaneous cascade translation system, the accuracy of the translation system directly depends on the accuracy of the acoustic recognition system. The speech recognition system can generate multiple candidate sentences based on the acoustic score and linguistic score of the speaker's statement, and sort the multiple candidate sentences. In order to fully tap the semantic differences between the multiple candidate sentences of the speech recognition system, the embodiment of the present application uses a fusion module of word granularity to fuse multiple candidate sentences in a way that seeks common ground while reserving differences. This is in that: first, by converting different words in multiple candidate sentences into sequences, the integrity of the multiple candidate sentence spaces can be retained and the scoring information of the speech recognition system for multiple candidate sentences can be retained; secondly, by extracting different words in multiple candidate sentences into the same text, the differences between different candidate sentences can be effectively strengthened, and then by different tags, it can be naturally distinguished whether the current text needs to be corrected; finally, by removing the same parts in multiple candidate sentences, the waste of computing and storage resources can be effectively reduced.

[0076] According to the embodiment of the present invention, by converting different words in multiple candidate sentences generated by the upstream speech recognition system into sequences, the difference parts of multiple candidate sentences are modeled in a targeted manner at the word level granularity, thereby improving the error tolerance of translation. In addition, by retaining the acoustic and linguistic sorting information of multiple candidate sentences in the upstream speech recognition system, the loss of subsequent computing resources and memory resources is reduced, and the real-time performance of the speech simultaneous interpretation system is improved.

[0077] In certain embodiments of the present application, the fault-tolerant translation model includes an encoder with a self-attention mechanism and a decoder with a self-attention mechanism, and the fault-tolerant translation model is used to encode, pre-select and decode the fused candidate sentences to obtain the target language sentence, including: using the encoder to encode and pre-select the fused candidate sentences to obtain the semantic representation of the fused candidate sentences; using the decoder to decode the semantic representation to obtain the target language sentence.

[0078] In a simultaneous cascade translation system, error propagation between various cascade subsystems (e.g., speech recognition module, encoding module, and decoding module) is a key factor affecting the overall system effect. Specifically, in this scenario, the errors generated by the speech recognition system will have a substantial impact on the downstream translation tasks. Therefore, in order to achieve fault-tolerant translation of recognition errors and improve the accuracy of translation in a simultaneous cascade translation system, the embodiments of the present application start from the translation source end (encoder end) and the translation model target end (decoder end) at the same time, and introduce fault-tolerant targets based on multiple candidate words in all directions in the encoding stage with multiple candidate sentences as input and the decoding stage of the translation model.

[0079] In certain embodiments of the present application, the encoder and the decoder are jointly trained based on a first loss function and a second loss function, wherein the first loss function is set for the task of semantically matching multiple candidate words based on the context information of the fused candidate sentences, and the second loss function is set for the machine translation task.

[0080] In some embodiments of the present application, the first loss function is expressed by the following formula:

[0081]

[0082] Among them, T represents the number of multiple candidate intervals in the sentence, K j represents the number of multiple candidate words in the jth candidate interval, and They represent the correct and i-th multiple candidate words or phrases in the j-th candidate interval respectively.

[0083] This loss function will be used together with the loss function of the translation model to jointly optimize the semantic representation of multiple candidates and context. By introducing this loss function, different candidate words in the sentence can be given Different weights.

[0084] The source text is represented as X=(x1,[start],y1,[mid],Y2,[end,…,x n ), where y i Indicates the identification of the i-th candidate word or phrase among multiple candidates, x i Represents the recognition of the same sentence composition in multiple candidate sentences. By serializing different candidate words into the context, the encoder i When performing semantic representation, we can make full use of context information x i and other candidate information y j In this way, the self-attention mechanism can be effectively used to achieve deep fusion of information between multiple candidate words and between multiple candidate words and context.

[0085] In order to enhance the recognition error tolerance of the translation system, the embodiment of the present application adopts a multi-task learning framework. Specifically, while optimizing the translation target, a multi-candidate pre-selection target is introduced on the encoder side, which can be regarded as a task of semantically matching multiple candidate words based on context information.

[0086] In certain embodiments of the present application, the error-tolerant translation method of the first aspect further includes: randomly selecting a second word in a sample sentence, and replacing the second word in the sample sentence position with a neighbor candidate of the second word in a candidate dictionary to generate a candidate training corpus for an error-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by frequency of occurrence in speech recognition.

[0087] Since the candidate semantic matching task goal needs to rely on the annotation of the correct candidate words or phrases, the embodiment of the present application can use a semi-supervised training method when training the model. First, the present application constructs a candidate dictionary in the form of key-value pairs based on multiple candidate words of the recognition system, in which the easily confused neighbor candidate words that often appear in the speech recognition system of common words are saved, and the neighbor candidate words are sorted according to the frequency of occurrence in the speech recognition system. Then, in the training stage, the words in the sample sentences are randomly selected, and the neighbor candidate words in the candidate dictionary are serialized and then replaced, thereby forging the training corpus of multiple candidate words.

[0088] According to an embodiment of the present application, by introducing a candidate dictionary and utilizing a semi-supervised training method, the reliance on labeled data can be reduced.

[0089] In some embodiments of the present application, the fused candidate sentence further includes: at least one third word, each of the plurality of candidate sentences includes at least one third word, and the position of the at least one third word in the fused candidate sentence is the same as the position of the at least one third word in the source language sentence. The third word here can be the same word in the candidate sentence.

[0090] In certain embodiments of the present application, the label information includes at least one of start information, end information and intermediate information, wherein the start information is used to indicate the starting position of multiple candidate words in the fused candidate sentence, the end information is used to indicate the ending position of multiple candidate words in the fused candidate sentence, and the intermediate information is used to indicate the intermediate position of multiple candidate words in the fused candidate sentence.

[0091] Specifically, the boundaries between multiple candidate words and contexts are indicated by the start and end information, and the boundaries between different candidate words are indicated by the middle information, so that clear error correction information can be given during subsequent decoding. For example, the start information can be represented by the English word [start], the end information can be represented by the English word [end], and the middle information can be represented by the English abbreviation [mid]. It should be understood that the above-mentioned representation of the label information is only an example, and other letters, numbers, or even Chinese characters can also be used to represent it.

[0092] In some embodiments of the present application, determining a fused candidate sentence based on multiple candidate sentences includes: using a longest common substring algorithm to determine a fused candidate sentence based on the multiple candidate sentences.

[0093] Specifically, the longest common substring algorithm can be used to determine the largest common substring between different candidate sentences, and the part between the largest common substrings or the part outside the most common substring can be determined as a candidate word. The longest common substring algorithm refers to an algorithm that finds the largest common substring between two strings, for example, by using dynamic programming or a suffix array. It should be understood that the embodiments of the present invention are not limited thereto, and other similar algorithms such as the longest common substring and string similarity algorithm can also be used.

[0094] Since there may be multiple different word intervals between the recognition candidate sentences, the embodiment of the present application can efficiently merge multiple candidate words by recursively utilizing the longest common substring algorithm.

[0095] In some embodiments of the present application, identifying multiple candidate sentences of a source language sentence includes: performing speech recognition on a speech signal of the source language sentence to obtain multiple candidate sentences of the source language sentence.

[0096] Specifically, the speech recognition module in the simultaneous cascade translation system may be used to perform speech recognition on the speech information of the source language sentence.

[0097] Combine the following Figure 3 , Figure 4A and Figure 4B The simultaneous cascade translation system of an embodiment of the present application is explained by taking the translation process of the sentence "Its cost of indoor arrangement is very low" as an example.

[0098] Figure 3 Shown is a flow chart of a fault-tolerant translation method provided by another exemplary embodiment of the present application. Figure 4A Shown is a schematic diagram of a process of fusing multiple candidate sentences provided by an exemplary embodiment of the present application. Figure 4B Shown is a schematic diagram of a process of fusing multiple candidate sentences provided by another exemplary embodiment of the present application.

[0099] 310 , performing acoustic recognition on the speech data containing the source language sentence to obtain a plurality of candidate sentences.

[0100] See also Figure 4A and Figure 4B For example, the speech data containing "Its cost to arrange indoors is very low" is recognized and converted into text by the acoustic recognition system, and three candidate sentences are obtained: "1. Its cost of walking indoors is very low", "1. Its cost of arranging indoors is very low" and "1. Its cost of valuation indoors is very low".

[0101] 320, determining multiple candidate words that are different words between the multiple candidate sentences, and determining the same sentence components in the multiple candidate sentences.

[0102] For example, the longest common substring algorithm may be used to sequentially locate different words or phrases in multiple candidate sentences.

[0103] 330, determine label information of multiple candidate words.

[0104] Specifically, the tag information may be information identifying the starting position, the ending position, and the middle position of the candidate word.

[0105] 340, recognizing the acoustic scores of the speech data by the acoustic recognition system, and sorting the candidate words according to the acoustic scores.

[0106] After the acoustic recognition system recognizes the speech data, the acoustic score of each candidate word can be obtained, and then multiple candidate words can be sorted according to the acoustic score. For example, the acoustic score of "arrangement" is greater than the acoustic scores of "step" and "valuation". Therefore, the arrangement is placed in the center so that the translation module can learn the acoustic information in the subsequent translation process, which helps to improve the accuracy of the translation.

[0107] It should be understood that the embodiments of the present application do not limit the execution order of 330 to 340, and vice versa is also possible.

[0108] 350, keep the same words in the candidate sentences in their original positions, convert these candidate words into sequences, insert the positions of the different words, and insert the label information into the context to obtain the fused candidate sentences.

[0109] For example, the two parts "it is indoors" and "the cost is very low" are retained in the original positions, while the three differential words "step", "layout" and "valuation" are arranged in sequence and placed in the position of the differential words, which naturally retains the sorting information of the acoustic recognition system.

[0110] For example, see Figure 4A, add [start] and [end] tags before and after the candidate word sequence, and insert [mid] tags between different candidate words. Figure 4B , add a number [2] before each candidate word sequence, for example, three [2] are added before the step, arrangement and valuation, where the number is the same as or related to the number of characters in the candidate word. This can distinguish between different words and the same words in the candidate sentences; and can naturally distinguish the fused candidate sentences from normal parallel sentence pairs through the added labels. By introducing these label information, the translation model in the cascade system can learn through training when to selectively translate multiple candidate phrases and when to maintain the integrity of the source information for translation, thereby ensuring the fidelity of the translation.

[0111] 360, an encoder with a self-attention mechanism is used to encode and pre-select the fusion candidate sentences to obtain the semantic representation of the fusion candidate sentences.

[0112] For example, the source text is represented as X = (x1, [start], y1, [mid], y2, [end], ..., x n ), where y i represents the i-th candidate word in the fusion candidate sentence, x i Indicates the same sentence components in the fusion candidate sentence. By serializing different candidate words into the context, the encoder can i When performing semantic representation, we can make full use of context information x i and other candidate words y j , this can effectively utilize the self-attention mechanism to achieve deep fusion of information between multiple candidate words and between multiple candidate words and context.

[0113] In the encoding stage, the self-attention mechanism is used to realize the deep fusion between the context and multiple candidate words and between multiple candidate words from the bottom up for the fused candidate sentences, and the preliminary scoring selection of multiple candidate words in the current context is realized based on the hidden output representation of the encoder to obtain preliminary scoring information. The preliminary scoring information of the encoder will be further used in the decoder.

[0114] 370, using a decoder with error tolerance to translate the semantic representation of the fused candidate sentence to obtain a final translation.

[0115] In the decoding stage, the decoder (or translation model) performs a second scoring selection based on the preliminary scoring information obtained by the encoder, and uses the translated text as the second source of information to implicitly participate in the selection of multiple final candidate words, which are finally presented in the form of translation. This ensures that robust and fault-tolerant translation can be achieved with only a single translation model with minimal consumption of computing and memory resources.

[0116] The simultaneous cascade translation system of the present application can greatly retain the acoustic and linguistic ordering information of multiple candidate sentences in the upstream speech recognition system, improve the accuracy of the simultaneous cascade translation system, and can also reduce the loss of computing resources and memory resources required for semantic representation of subsequent modules, thereby improving the real-time performance of the speech simultaneous interpretation system; in addition, compared with the sentence granularity, the word-level granularity is used to specifically model the different parts of multiple candidate sentences obtained through speech recognition, which is more targeted and helps to improve the accuracy of the error tolerance of multiple candidate word recognition.

[0117] Figure 5 Shown is a flowchart of a method for training an error-tolerant translation model provided by an exemplary embodiment of the present application. Figure 5 The method may be performed by a computing device (eg, a server). Figure 5 As shown, the method includes the following contents.

[0118] 510: Identify multiple candidate sentences corresponding to the source language sentence sample.

[0119] 520: Determine a fused candidate sentence according to the plurality of candidate sentences, where the fused candidate sentence includes a plurality of candidate words corresponding to the first word in the source language sentence sample and label information for indicating the plurality of candidate words.

[0120] 530: Encode and pre-select the fused candidate sentences using an encoder to obtain a semantic representation of the fused candidate sentences, and determine a first loss function based on the semantic representation; decode the semantic representation using a decoder to obtain a target language sentence, and determine a second loss function based on the target language sentence.

[0121] 540: Jointly train the encoder and the decoder based on the first loss function and the second loss function to obtain a fault-tolerant translation model.

[0122] Specifically, based on the multi-task learning framework, a loss function for pre-selecting multiple candidate words can be introduced in the encoder module, and the selection task of identifying multiple candidate words and the machine translation task can be jointly modeled end-to-end together with the loss function of the decoder (or translation model), which helps to alleviate the error propagation problem in cascade modeling. In addition, by implementing explicit and implicit selection of multiple candidates at the source encoder and the target decoder, that is, implicitly implementing selective translation of multiple candidate words based on information in the two channels of the source language and the target language, it helps to improve the selection accuracy of identifying multiple candidates.

[0123] According to an embodiment of the present application, by identifying multiple candidate sentences corresponding to a source language sentence sample, determining a fused candidate sentence based on the multiple candidate sentences, encoding and pre-selecting the fused candidate sentence using an encoder, obtaining a semantic representation of the fused candidate sentence, determining a first loss function based on the semantic representation, decoding the semantic representation using a decoder, obtaining a target language sentence, determining a second loss function based on the target language sentence, and jointly training the encoder and decoder based on the first loss function and the second loss function to obtain a fault-tolerant translation model. Since the fault-tolerant translation model can learn the difference parts of the identified multiple candidate sentences in a targeted manner at a word-level granularity, it helps to improve the accuracy of fault-tolerant translation. In addition, since the fusion of multiple candidate sentences can reduce the subsequent loss of computing resources and memory resources required for semantic representation of multiple candidates, the real-time performance of the translation is improved.

[0124] In some embodiments of the present application, the first loss function is expressed by the following formula:

[0125]

[0126] Among them, T represents the number of multiple candidate intervals in the sentence, K j represents the number of multiple candidate words in the jth candidate interval, and They represent the correct and i-th multiple candidate words or phrases in the j-th candidate interval respectively.

[0127] In certain embodiments of the present application, Figure 5 The method also includes: randomly selecting a second word in the sample sentence, and replacing the second word in the sample sentence position with a neighbor candidate of the second word in a candidate dictionary to generate a candidate training corpus for the error-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by frequency of occurrence in speech recognition.

[0128] By introducing a candidate dictionary, the use of semi-supervised training can reduce the dependence on labeled data.

[0129] Combination Figure 6 and Figure 7To illustrate the end-to-end fault-tolerant translation system according to an embodiment of the present application.

[0130] Figure 6 Shown is a schematic diagram of the structure of a fault-tolerant translation system provided by another exemplary embodiment of the present application.

[0131] See also Figure 6 The error-tolerant translation system mainly includes: an identification module 610, a fusion module 620 and an error-tolerant translation module 630. The error-tolerant translation module 630 includes an encoding module 631 and a decoding module 632.

[0132] Specifically, the recognition module 610 is used to receive the voice data collected by the voice collection device, and convert the voice data into text data, which can include multiple candidate sentences. The fusion module 610 is connected with the recognition module 610, and is used for fusion processing of multiple candidate sentences to obtain fusion candidate sentences. The encoding module 631 is connected with the fusion module 620, and is used for encoding and pre-selecting the fusion candidate sentences to obtain semantic representation. The decoding module 632 is connected with the encoding module, and is used for decoding the semantic representation to obtain the final translation text. It can be seen from this that the above-mentioned fault-tolerant translation module is a cascade translation system formed by cascading the recognition module, the encoding module and the decoding module.

[0133] like Figure 6 As shown, {S} i=1…N It indicates that acoustic recognition obtains multiple candidate sentences, S0 indicates the fused candidate sentence, H indicates the hidden layer distributed representation, i.e., the semantic representation, and T indicates the translated text.

[0134] Figure 7 Shown is a schematic diagram of the overall framework structure of a Transformer model 700 provided by an exemplary embodiment of the present application.

[0135] like Figure 7 As shown, the error-tolerant translation model mainly includes: a distributed encoding module 750 incorporating multiple candidate information and an implicit multiple candidate error-tolerant decoding module 740 .

[0136] The Transformer model may include an embedding (Input Embedding) layer 710 , an embedding (InputEmbedding) (Input Embedding) layer 720 , an embedding (OutInputEmbedding) layer 730 , a matching block (MatchBlock) 760 , and a Linear&Softmax module 770 .

[0137] The decoding module 740 includes a self-attention layer 741, a normalization layer 742, a cross attention layer 743, a normalization layer 744, a feed forward neural network layer 746, and a normalization layer 747. The decoding module 740 is used to use the self-attention mechanism to semantically represent the fused candidate sentences.

[0138] The encoding module 750 includes a self-attention layer 751, a normalization layer 752, a feed forward neural network layer 753, and a normalization layer 754. The encoding module 750 is used to translate the semantic representation of the fused candidate sentence using a cross attention mechanism.

[0139] Match Block 760 is used to implement the above-mentioned task of semantically matching multiple candidate words based on context information.

[0140] The linear & classification layer 770 includes a linear layer and a Softmax layer, wherein the Softmax layer is used for classification using the activation function Softmax.

[0141] Exemplary Devices

[0142] Figure 8 FIG. 8 is a schematic diagram of a structure of a fault-tolerant translation device 800 provided by an exemplary embodiment of the present application. Figure 8 As shown, the error-tolerant translation device 800 includes: an identification module 810 , a determination module 820 and a coding and decoding module 830 .

[0143] The recognition module 810 is used to identify multiple candidate sentences corresponding to the source language sentence. The determination module 820 is used to determine the fused candidate sentence based on the multiple candidate sentences, and the fused candidate sentence includes multiple candidate words corresponding to the first word in the source language sentence and label information for indicating the multiple candidate words. The encoding and decoding module 830 is used to encode, pre-select and decode the fused candidate sentence using the fault-tolerant translation model to obtain the target language sentence.

[0144] According to an embodiment of the present application, by identifying multiple candidate sentences corresponding to a source language sentence, determining a fused candidate sentence according to multiple candidate sentences, and utilizing a fault-tolerant translation model to encode, pre-select and decode the fused candidate sentence, a target language sentence is obtained, and the fault tolerance rate and real-time performance of the translation can be improved. Since the fault-tolerant translation model can learn the difference parts of the identified multiple candidate sentences in a targeted manner with word-level granularity, it is helpful to improve the accuracy of fault-tolerant translation. In addition, since the fusion of multiple candidate sentences can reduce the subsequent loss of computing resources and memory resources required for semantic representation of sentences, the real-time performance of the translation is improved.

[0145] In some embodiments of the present application, the determination module 820 prioritizes the recognition scores of multiple candidate words and embeds the multiple candidate words as a sequence in the fused candidate sentence, wherein the position of the sequence in the fused candidate sentence is the same as the position of the first word in the source language sentence.

[0146] In certain embodiments of the present application, the error-tolerant translation method is a simultaneous interpretation method, and the recognition score includes an acoustic recognition score.

[0147] In certain embodiments of the present application, the fault-tolerant translation model includes an encoder with a self-attention mechanism and a decoder with a self-attention mechanism, and the fault-tolerant translation model is used to encode, pre-select and decode the fused candidate sentences to obtain the target language sentence, including: using the encoder to encode and pre-select the fused candidate sentences to obtain the semantic representation of the fused candidate sentences; using the decoder to decode the semantic representation to obtain the target language sentence.

[0148] In certain embodiments of the present application, the encoder and the decoder are jointly trained based on a first loss function and a second loss function, wherein the first loss function is set for the task of semantically matching multiple candidate words based on the context information of the fused candidate sentences, and the second loss function is set for the machine translation task.

[0149] In some embodiments of the present application, the first loss function is expressed by the following formula:

[0150]

[0151] Among them, T represents the number of multiple candidate intervals in the sentence, K j represents the number of multiple candidate words in the jth candidate interval, and They represent the correct and i-th multiple candidate words or phrases in the j-th candidate interval respectively.

[0152] In certain embodiments of the present application, the fault-tolerant translation method of the first aspect also includes: a generation module, which is used to randomly select a second word in a sample sentence, and replace the second word in the sample sentence position with a neighbor candidate of the second word in a candidate dictionary to generate a candidate training corpus for a fault-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by frequency of occurrence in speech recognition.

[0153] In certain embodiments of the present application, the label information includes at least one of start information, end information and intermediate information, wherein the start information is used to indicate the starting position of multiple candidate words in the fused candidate sentence, the end information is used to indicate the ending position of multiple candidate words in the fused candidate sentence, and the intermediate information is used to indicate the intermediate position of multiple candidate words in the fused candidate sentence.

[0154] In some embodiments of the present application, the determination module 820 uses the longest common substring algorithm to determine the fused candidate sentence according to multiple candidate sentences.

[0155] In some embodiments of the present application, the recognition module 810 performs speech recognition on the speech signal of the source language sentence to obtain a plurality of candidate sentences of the source language sentence.

[0156] In certain embodiments of the present application, the fused candidate sentence also includes: at least one third word, each of the multiple candidate sentences includes at least one third word, and the position of the at least one third word in the fused candidate sentence is the same as the position of the at least one third word in the source language sentence.

[0157] In certain embodiments of the present application, the error-tolerant translation model includes a neural network model with a self-attention mechanism.

[0158] In some embodiments of the present application, the neural network model includes a transformer model.

[0159] It should be understood that the operations and functions of the identification module 810, the determination module, and the encoding and decoding module 830 in the above embodiment can refer to the above Figure 2 or Figure 3 The description of the fault-tolerant translation method provided in the embodiment is not repeated here to avoid repetition.

[0160] Fig. 9 FIG. 9 is a schematic diagram of a structure of an apparatus 900 for training a fault-tolerant translation model provided by an exemplary embodiment of the present application. Fig. 9 As shown, the apparatus 900 includes: an identification module 910 , a determination module 920 , a coding and decoding module 930 , and a training module 940 .

[0161] The identification module 910 is used to identify multiple candidate sentences corresponding to the source language sentence sample; the determination module 920 is used to determine the fused candidate sentence based on the multiple candidate sentences, and the fused candidate sentence includes multiple candidate words corresponding to the first word in the source language sentence sample and label information for indicating the multiple candidate words. The encoding and decoding module 930 is used to encode and pre-select the fused candidate sentence using an encoder to obtain a semantic representation of the fused candidate sentence, determine a first loss function based on the semantic representation, decode the semantic representation using a decoder to obtain a target language sentence, and determine a second loss function based on the target language sentence. The training module 940 is used to jointly train the encoder and decoder based on the first loss function and the second loss function to obtain a fault-tolerant translation model.

[0162] According to an embodiment of the present application, multiple candidate sentences corresponding to source language sentence samples are identified; a fused candidate sentence is determined based on the multiple candidate sentences, an encoder is used to encode and pre-select the fused candidate sentence, a semantic representation of the fused candidate sentence is obtained, a first loss function is determined based on the semantic representation, a decoder is used to decode the semantic representation, a target language sentence is obtained, and a second loss function is determined based on the target language sentence; the encoder and the decoder are jointly trained based on the first loss function and the second loss function to obtain a fault-tolerant translation model. Since the fault-tolerant translation model can learn the difference parts of the identified multiple candidate sentences in a targeted manner at a word-level granularity, it helps to improve the accuracy of fault-tolerant translation. In addition, since the fusion of multiple candidate sentences can reduce the loss of computing resources and memory resources required for the subsequent semantic representation of the candidate sentences, the real-time performance of the translation is improved.

[0163] In some embodiments of the present application, the first loss function is expressed by the following formula:

[0164]

[0165] Among them, T represents the number of multiple candidate intervals in the sentence, K j represents the number of multiple candidate words in the jth candidate interval, and They represent the correct and i-th multiple candidate words or phrases in the j-th candidate interval respectively.

[0166] In certain embodiments of the present application, the device 900 includes: a generation module 950, which is used to randomly select a second word in a sample sentence and replace the second word in the sample sentence position with a neighbor candidate of the second word in a candidate dictionary to generate a candidate training corpus for a fault-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by frequency of occurrence in machine speech recognition.

[0167] It should be understood that the operations and functions of the identification module 910, the determination module 920, the encoding and decoding module 930, and the training module 940 in the above embodiment can refer to the above Figure 5 The description of the method for training the fault-tolerant translation model provided in the embodiment will not be repeated here to avoid repetition.

[0168] Fig.10 FIG. 1 is a block diagram of an electronic device 1000 for executing an error-tolerant translation method or a method for training an error-tolerant translation model provided by an exemplary embodiment of the present application.

[0169] Reference Fig.10 , the electronic device 1000 includes a processing component 1010, which further includes one or more processors, and a memory resource represented by a memory 1020 for storing instructions executable by the processing component 1010, such as an application. The application stored in the memory 1020 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1010 is configured to execute instructions to perform the above-mentioned error-tolerant translation method or the method for training an error-tolerant translation model.

[0170] The electronic device 1000 may also include a power supply component configured to perform power management of the electronic device 1000, a wired or wireless network interface configured to connect the electronic device 1000 to a network, and an input / output (I / O) interface. The electronic device 1000 may be operated based on an operating system stored in the memory 1020, such as Windows Server 2003. TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or similar.

[0171] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device 1000, enables the electronic device 1000 to perform the error-tolerant translation method or the method for training an error-tolerant translation model.

[0172] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, and will not be described one by one here.

[0173] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0175] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0176] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0178] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program check codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0179] It should be noted that, in the description of this application, the terms "first", "second", "third", etc. are used for descriptive purposes only and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.

[0180] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A fault-tolerant translation method, characterized in that: include: identifying a plurality of candidate sentences corresponding to the source language sentence; Determining a fused candidate sentence according to the multiple candidate sentences, the fused candidate sentence including multiple candidate words corresponding to the first word in the source language sentence and label information for indicating the multiple candidate words; The fused candidate sentence with the label information and the multiple candidate words is input into a fault-tolerant translation model to encode, pre-select and decode the fused candidate sentence to obtain a target language sentence, wherein the fault-tolerant translation model is used to selectively translate the multiple candidate words.

2. The error-tolerant translation method according to claim 1, characterized in that: The determining of a fused candidate sentence according to the multiple candidate sentences includes: The plurality of candidate words are embedded as a sequence in the fused candidate sentence according to the priority order of the recognition scores of the plurality of candidate words, wherein the position of the sequence in the fused candidate sentence is the same as the position of the first word in the source language sentence.

3. The error-tolerant translation method according to claim 1, characterized in that: The error-tolerant translation model includes an encoder with a self-attention mechanism and a decoder with a self-attention mechanism. The error-tolerant translation model is used to encode, pre-select and decode the fusion candidate sentences to obtain the target language sentence, including: Encoding and pre-selecting the fused candidate sentences using the encoder to obtain semantic representations of the fused candidate sentences; The semantic representation is decoded by using the decoder to obtain the target language sentence.

4. The error-tolerant translation method according to claim 3, characterized in that: The encoder and the decoder are jointly trained based on a first loss function and a second loss function, wherein the first loss function is set for a task of semantically matching the multiple candidate words based on context information of the fused candidate sentence, and the second loss function is set for a machine translation task.

5. The error-tolerant translation method according to claim 1, characterized in that: Also includes: A second word in a sample sentence is randomly selected, and a neighbor candidate of the second word in a candidate dictionary replaces the second word at the position of the sample sentence to generate a candidate training corpus for the error-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by occurrence frequency in speech recognition.

6. The error-tolerant translation method according to claim 1, characterized in that: The label information includes at least one of start information, end information and intermediate information, wherein the start information is used to indicate the start positions of the multiple candidate words in the fused candidate sentence, the end information is used to indicate the end positions of the multiple candidate words in the fused candidate sentence, and the intermediate information is used to indicate the intermediate positions of the multiple candidate words in the fused candidate sentence.

7. The error-tolerant translation method according to claim 1, characterized in that: The determining of a fused candidate sentence according to the multiple candidate sentences includes: The fused candidate sentence is determined according to the multiple candidate sentences by using the longest common substring algorithm.

8. The error-tolerant translation method according to claim 1, characterized in that: The identifying of multiple candidate sentences of the source language sentence includes: Speech recognition is performed on the speech signal of the source language sentence to obtain a plurality of candidate sentences of the source language sentence.

9. The error-tolerant translation method according to any one of claims 1 to 8, characterized in that: The fused candidate sentence further includes: at least one third word, each of the plurality of candidate sentences includes the at least one third word, and a position of the at least one third word in the fused candidate sentence is the same as a position of the at least one third word in the source language sentence.

10. A method for training an error-tolerant translation model, characterized in that: include: Identify multiple candidate sentences corresponding to the source language sentence sample; Determine a fused candidate sentence according to the multiple candidate sentences, the fused candidate sentence including multiple candidate words corresponding to the first word of the source language sentence sample and label information for indicating the multiple candidate words; Encoding and pre-selecting the fused candidate sentence with the label information and the plurality of candidate words using an encoder to obtain a semantic representation of the fused candidate sentence, and determining a first loss function based on the semantic representation; Decoding the semantic representation using a decoder to obtain a target language sentence, and determining a second loss function based on the target language sentence; The encoder and the decoder are jointly trained based on a first loss function and a second loss function to obtain an error-tolerant translation model, wherein the error-tolerant translation model is used to selectively translate the multiple candidate words.

11. The method for training an error-tolerant translation model according to claim 10, characterized in that: Also includes: A second word in a sample sentence is randomly selected, and a neighbor candidate of the second word in a candidate dictionary replaces the second word at the position of the sample sentence to generate a candidate training corpus for the error-tolerant translation model, wherein the candidate dictionary stores neighbor candidates sorted by occurrence frequency in speech recognition.

12. A fault-tolerant translation device, characterized in that: include: A recognition module, used for identifying multiple candidate sentences corresponding to the source language sentence; A determination module, configured to determine a fused candidate sentence according to the plurality of candidate sentences, wherein the fused candidate sentence includes a plurality of candidate words corresponding to the first word in the source language sentence and label information indicating the plurality of candidate words; The encoding and decoding module is used to input the fused candidate sentence with the label information and the multiple candidate words into the fault-tolerant translation model to encode, pre-select and decode the fused candidate sentence to obtain a target language sentence, wherein the fault-tolerant translation model is used to selectively translate the multiple candidate words.

13. A device for training an error-tolerant translation model, characterized in that: include: A recognition module, used for identifying multiple candidate sentences corresponding to the source language sentence sample; A determination module, configured to determine a fused candidate sentence according to the plurality of candidate sentences, wherein the fused candidate sentence includes a plurality of candidate words corresponding to the first word in the source language sentence sample and label information for indicating the plurality of candidate words; an encoding and decoding module, configured to encode and pre-select the fused candidate sentence with the label information and the multiple candidate words by using an encoder, obtain a semantic representation of the fused candidate sentence, determine a first loss function based on the semantic representation, decode the semantic representation by using a decoder, obtain a target language sentence, and determine a second loss function based on the target language sentence; A training module is used to jointly train the encoder and the decoder based on a first loss function and a second loss function to obtain an error-tolerant translation model, wherein the error-tolerant translation model is used to selectively translate the multiple candidate words.

14. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor, The processor is used to execute the error-tolerant translation method described in any one of claims 1 to 9 or the method for training an error-tolerant translation model described in any one of claims 10 to 11.

15. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the error-tolerant translation method described in any one of claims 1 to 9 or the method for training an error-tolerant translation model described in any one of claims 10 to 11.

Citation Information

Patent Citations

  • Pre-training method and device of intelligent translation model and storage medium

    CN111460838A

  • Machine translation method and device, electronic equipment and storage medium

    CN111860001A