Text translation method and device, equipment, medium and product
By adding multiple translations to enable discrimination after each round of translation process, the semantic stability of the hidden state is determined, and the applicability problem of the translation model with the hidden layer only includes the decoder in real-time translation scenarios is solved, achieving smooth and accurate real-time translation effects.
Patent Information
- Application Number
- CN202510399952.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The translation model that only includes the decoder is difficult to ensure accuracy and fluency in real-time translation scenarios and cannot be effectively applied to real-time translation tasks.
By adding multiple translations to enable discrimination after each round of translation process, we can judge whether the hidden state of the text sequence is semantically stable, and use the target translation model to perform a new round of translation to ensure the accuracy and fluency of each round of translation process.
The application of the translation model with hidden layer only including decoder in real-time translation scenarios is realized, which improves the accuracy and fluency of translation and ensures the effect of real-time translation.
Smart Images

Figure CN120337946A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a text translation method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] With the continuous development of computer technology, language models that can be applied to different scenarios have emerged. For example, language models can be applied to the translation scenario. In the translation scenario, a language model can also be called a translation model, and the translation model is used to translate text from the source language into the target language in real time.
[0003] A translation model usually consists of an input layer, multiple hidden layers, and an output layer. The hidden layers of some translation models only include a decoder, while the hidden layers of other translation models include an encoder and a decoder. In the related art, in the real-time translation scenario, a language model with an encoder and a decoder in the hidden layer is usually used for real-time translation of text. Since the language model with only a decoder in the hidden layer lacks an encoder, it is difficult to ensure the accuracy and fluency of real-time translation and is difficult to be applied to the real-time translation scenario. Summary of the Invention
[0004] This application provides a text translation method. This method can apply a language model with only a decoder in the hidden layer to the real-time translation scenario and achieve fluent real-time translation. This application also provides a corresponding apparatus, electronic device, computer-readable storage medium, and computer program product for the above method.
[0005] In a first aspect, this application provides a text translation method, which includes:
[0006] Continuously receive the first text described in the source language input by the user; where the first text corresponds to multiple tokens;
[0007] In response to the stop of the nth round of translation process, perform M translation start discriminations; where the nth round of translation process is used to translate the ith to jth tokens corresponding to the first text, 1 ≤ i < j;
[0008] In response to the similarity degree between the hidden state corresponding to the Mth first text sequence and the hidden state corresponding to the M - Nth first text sequence satisfying the translation start condition, use the target translation model to perform the (n + 1)th round of translation process; where the (n + 1)th round of translation process is used to translate the Mth first text sequence, N < M;
[0009] Among them, the m-th translation start determination includes: determining the m-th first text sequence, sending the m-th first text sequence to the target translation model, and receiving the hidden state corresponding to the m-th first text sequence output by the hidden layer of the target translation model, where 1 ≤ m ≤ M, and the hidden layer of the target translation model consists of multiple decoders;
[0010] The m-th first text sequence includes: the (j + 1)-th to the (j + m)-th tokens corresponding to the first text.
[0011] In a second aspect, the present application provides a text translation device, which includes:
[0012] An acquisition module, configured to continuously receive the first text described in the source language input by the user; among them, the first text corresponds to multiple tokens;
[0013] A discrimination module, configured to perform M times of translation start discrimination in response to the stop of the n-th round of translation process; where the n-th round of translation process is used to translate the i-th to the j-th tokens corresponding to the first text, where 1 ≤ i < j;
[0014] A translation module, configured to execute the (n + 1)-th round of translation process by using the target translation model in response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence satisfying the translation start condition; where the (n + 1)-th round of translation process is used to translate the M-th first text sequence, and N < M;
[0015] Among them, the m-th translation start determination includes: determining the m-th first text sequence, sending the m-th first text sequence to the target translation model, and receiving the hidden state corresponding to the m-th first text sequence output by the hidden layer of the target translation model, where 1 ≤ m ≤ M, and the hidden layer of the target translation model consists of multiple decoders;
[0016] The m-th first text sequence includes: the (j + 1)-th to the (j + m)-th tokens corresponding to the first text.
[0017] In a third aspect, the present application provides an electronic device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute the instructions stored in the memory so that the electronic device executes the text translation method as described in the first aspect or any implementation manner of the first aspect.
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct the electronic device to execute the text translation method described in the above first aspect or any implementation manner of the first aspect.
[0019] In a fifth aspect, the present application provides a computer program product containing instructions, which, when running on an electronic device, causes the electronic device to execute the text translation method described in the first aspect or any implementation manner of the first aspect above.
[0020] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
[0021] As can be seen from the above technical solutions, the present application has the following advantages:
[0022] The present application provides a text translation method, which continuously receives a first text described in a source language input by a user, where the first text corresponds to multiple tokens. In response to the stop of the nth round of translation process, M translation start discriminations are performed, where the nth round of translation process is used to translate the ith to jth tokens corresponding to the first text. In response to the similarity degree between the hidden state corresponding to the Mth first text sequence and the hidden state corresponding to the M - Nth first text sequence satisfying the translation start condition, the (n + 1)th round of translation process is performed using a target translation model, where the (n + 1)th round of translation process is used to translate the Mth first text sequence.
[0023] Wherein, the mth translation start discrimination includes: determining the mth first text sequence, sending the mth first text sequence to the target translation model, receiving the hidden state corresponding to the mth first text sequence output by the hidden layer of the target translation model, the hidden layer of the target translation model consists of multiple decoders, and the mth first text sequence includes: the (j + 1)th to (j + m)th tokens corresponding to the first text.
[0024] In this method, for a translation model whose hidden layer only includes decoders, by adding multiple translation start discriminations after each round of translation process, when the hidden states of the text sequences in two translation start discriminations represent that the text sequence that has been received but not yet translated (i.e., the Mth text sequence) has a stable semantics, a new round of translation process is started to translate the text sequence that has been received but not yet translated. In this way, the start timing of each round of translation process is accurately judged, and in the process of the user continuously inputting the text to be translated, smooth real-time translation is achieved through multiple rounds of translation processes, applying the translation model whose hidden layer only includes decoders to the real-time translation scenario, and improving the applicability of the translation model whose hidden layer only includes decoders. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below.
[0026] Figure 1 A structural schematic diagram of a translation model provided for an embodiment of the present application;
[0027] Figure 2 A flowchart of a text translation method provided by an embodiment of the present application;
[0028] Figure 3 A flowchart of a text translation method provided by an embodiment of the present application;
[0029] Figure 4 A structural schematic diagram of a text translation device provided by an embodiment of the present application;
[0030] Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0031] The terms "first" and "second" in the embodiments of the present application are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0032] First, some technical terms and application scenarios involved in the embodiments of the present application are introduced.
[0033] A language model can be understood as a natural language processing model based on deep learning technology. A language model usually has the ability to understand, process, and generate natural language. The language model can be applied to different scenarios. For example, the language model can be applied to scenarios such as question answering, content generation, and translation.
[0034] When applying the language model to the translation scenario, the language model can also be called a translation model. The translation model can provide translation services in different forms such as cloud services, hardware translation devices, and software translation applications. Specifically, real-time translation service is a type of translation service. In the real-time translation scenario, the user continuously inputs text described in the source language. During the process of the user inputting text, the translation model real-time translates the text from the source language into the target language, reducing the waiting time of the user.
[0035] For example, the user continuously inputs the text described in English, "Newton's laws of motion describe the relationship between forces and the motion of an object." When the user inputs "Newton's laws of motion describe", the translation model can output the translation result of the Chinese description, "Newton's laws of motion describe." Before the text described in the source language has formed a complete sentence, the translation of part of the text content begins, realizing low-latency real-time translation of "translation while inputting".
[0036] Usually, the translation model generates text based on "next token prediction", that is, the translation model predicts the next most likely token based on the existing text. Specifically, the translation model is usually based on the transformer architecture, such as Figure 1 A schematic diagram of the structure of a translation model is shown, where the translation model may include an input layer, multiple hidden layers, and an output layer.
[0037] Among them, the existing text is tokenized and decomposed into the smallest units (i.e., multiple tokens) that the translation model can understand and process, and the multiple tokens corresponding to the existing text are input into the input layer of the translation model. In the input layer, the embedding vector (embedding) of each token is determined. Therefore, the input layer can also be called the embedding layer.
[0038] Next, the embedding vectors of each word are input into multiple hidden layers of the translation model. In the multiple hidden layers, the embedding vectors of multiple words are calculated based on the self-attention mechanism, and the last hidden layer outputs the hidden state of multiple words.
[0039] Finally, the hidden states of multiple words are input into the output layer of the translation model. In the output layer, a linear neural network and normalization processing are used to calculate the probability of different words being the next word based on the existing text. Finally, the word with the highest probability is generated and output as the next word, completing a round of model reasoning.
[0040] Translation models are divided into two types: the hidden layer of the first type of translation model only includes a decoder, and the hidden layer of the second type of translation model includes an encoder and a decoder. In the first type of translation model, the embedding vectors of multiple tokens first pass through the encoder. The encoder extracts the context information of multiple tokens through a bidirectional attention mechanism and passes the context information of multiple tokens to the decoder. The decoder determines the hidden states of multiple tokens by combining the context information provided by the encoder. In the second type of translation model, the embedding vectors of multiple tokens directly enter the decoder. The decoder captures the relationships between multiple tokens through a self-attention mechanism and determines the hidden states of multiple tokens. Compared with the first type of translation model, the decoder in the hidden layer of the second type of translation model cannot obtain context information.
[0041] In a real-time translation scenario, since the second type of translation model lacks an encoder and cannot obtain high-level semantic information and context representation, it is difficult to ensure the accuracy and fluency of real-time translation. Moreover, the second type of translation model is difficult to effectively align the source language text and the target language translation result and is not suitable for real-time translation tasks.
[0042] In view of this, the present application provides a text translation method. This method continuously receives the first text described in the source language input by the user, where the first text corresponds to multiple tokens. In response to the stop of the nth round of translation process, M translation start discriminations are performed. The nth round of translation process is used to translate the i-th to j-th tokens corresponding to the first text. In response to the similarity degree between the hidden state corresponding to the Mth first text sequence and the hidden state corresponding to the M - Nth first text sequence satisfying the translation start condition, the (n + 1)th round of translation process is performed using the target translation model, where the (n + 1)th round of translation process is used to translate the Mth first text sequence.
[0043] Among them, the mth translation start discrimination includes: determining the mth first text sequence, sending the mth first text sequence to the target translation model, and receiving the hidden state corresponding to the mth first text sequence output by the hidden layer of the target translation model. The hidden layer of the target translation model consists of multiple decoders. The mth first text sequence includes: the (j + 1)-th to (j + m)-th tokens corresponding to the first text.
[0044] In this method, for a translation model whose hidden layer only includes a decoder, by adding multiple translation start judgments after each round of translation process, when the hidden state representation of the text sequence in the two translation start judgments has received but not yet translated the text sequence (i.e., the Mth text sequence) and its semantics is stable, a new round of translation process is started to translate the received but not yet translated text sequence. In this way, the start timing of each round of translation process is accurately judged. During the process of the user continuously inputting the text to be translated, smooth real-time translation is achieved through multiple rounds of translation process, and the translation model whose hidden layer only includes a decoder is applied to the real-time translation scenario, improving the applicability of the translation model whose hidden layer only includes a decoder.
[0045] To facilitate understanding of the technical solutions provided in the embodiments of the present application, the following will be described with reference to the accompanying drawings. Refer to Figure 2 The flowchart of a text translation method shown in the figure, the text translation method provided in the embodiments of the present application can be applied to cloud-based translation services, offline translation applications, hardware translation devices, and mobile translation applications. The method specifically includes:
[0046] S201: Continuously receive the first text described in the source language input by the user.
[0047] Among them, the first text can be understood as the text with a translation requirement. The first text is described in the source language, and the translation requirement is to translate the first text described in the source language into the second text in the target language. That is to say, the second text can be understood as the translation result after translating the first text. For example, the source language can be English and the target language can be Chinese.
[0048] In the embodiments of the present application, the first text is translated using a target translation model. Among them, the target translation model can be understood as a language model with text translation capabilities. The hidden layer of the target translation model consists of multiple decoders, that is, the hidden layer of the target translation model only includes multiple decoders and does not include an encoder.
[0049] Since the target translation model performs text translation based on "next token prediction", the target translation model needs to process multiple tokens corresponding to the first text. Among them, the multiple tokens corresponding to the first text can be understood as the multiple tokens obtained after tokenizing the first text.
[0050] In the embodiments of the present application, there is no limitation on the way for the user to input the first text. In some embodiments, the user inputs the text in a text input manner. In this case, the text input by the user is the first text. In other embodiments, the user inputs the text in a voice input manner. In this case, through speech recognition of the voice input by the user, the voice input by the user is converted into text. In this case, the converted text is the first text.
[0051] In the embodiments of the present application, real-time translation is performed on the first text, that is, during the process of the user inputting the first text and before the entire first text is input, translation starts and a translation result is generated. Therefore, continuously receive the first text input by the user, for example, continuously receive each character in the first text input by the user, so as to start real-time translation.
[0052] S202: In response to the stop of the nth round of translation process, perform M times of translation start determination.
[0053] In the embodiments of the present application, since real-time translation needs to be performed on the first text, the process of translating the first text can be divided into multiple rounds of translation processes, and each round of translation process translates the corresponding partial tokens of the first text.
[0054] See Figure 2 As shown, after each round of translation process for the first text input by the user stops, multiple translation start determinations will be performed to determine the start timing of the next round of translation process.
[0055] Taking the nth round of translation process as an example for illustration, the nth round of translation process can be understood as any round of translation process for real-time translation of the first text. The nth round of translation process can be used to translate the ith to jth tokens corresponding to the first text, where 1 ≤ i < j. That is to say, when n is 1, in the nth round of translation process, the target translation model completes the translation of the first j tokens of the first text. When n is not 1, in the previous n - 1 rounds of translation processes, the target translation model completes the translation of the first i - 1 tokens in the first text. In the nth round of translation process, the target translation model completes the translation of the ith to jth tokens of the first text. After the nth round of translation process stops, the target translation model completes the translation of the first j tokens of the first text.
[0056] The translation start determination can be used to determine whether to start the next round of translation process. In other words, after the nth round of translation process stops, perform multiple translation start determinations. When the Mth translation start determination indicates starting the next round of translation process, start the (n + 1)th round of translation process, so that the target translation model continues to translate from the (j + 1)th token of the first text. In this way, through multiple rounds of translation processes, real-time translation of the first text is achieved.
[0057] Taking the determination of the m-th translation start as an example for illustration, where 1 ≤ m ≤ M, the determination of the m-th translation start includes: determining the m-th first text sequence, sending the m-th first text sequence to the target translation model, and receiving the hidden state corresponding to the m-th first text sequence output by the hidden layer of the target translation model.
[0058] Among them, the m-th first text sequence includes: the (j + 1)-th to the (j + m)-th tokens corresponding to the first text. That is to say, the first token of the first text sequence is the first untranslated token in the first text (i.e., the (j + 1)-th token). In each determination of the translation start, the subsequent tokens of the first text are sequentially added to the first text sequence. That is, the 1st first text sequence only includes the (j + 1)-th token, the 2nd first text sequence includes the (j + 1)-th token and the (j + 2)-th token, ……, the m-th first text sequence includes the (j + 1)-th token to the (j + m)-th token.
[0059] In the determination of the m-th translation start, from multiple hidden layers of the target translation model, obtain the hidden state corresponding to the m-th first text sequence, that is, determine the result after the m-th first text sequence passes through the attention mechanism calculation. In the attention mechanism calculation, the target translation model captures the dependency relationships between various positions in the m-th first text sequence. The hidden state corresponding to the m-th first text sequence can represent the semantic information of each token in the m-th first text sequence and the relationship information with other tokens. Therefore, by receiving the hidden state corresponding to the m-th first text sequence, the overall semantics represented by the m-th first text sequence can be determined.
[0060] S203: In response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence satisfying the translation start condition, use the target translation model to execute the (n + 1)-th round of translation process.
[0061] The determination of the m-th translation start further includes: judging whether the similarity degree between the hidden state corresponding to the m-th first text sequence and the hidden state corresponding to the (m - N)-th first text sequence satisfies the translation start condition, where N < M. Among them, the similarity degree can be used to measure whether the hidden state corresponding to the m-th first text sequence is similar to the hidden state corresponding to the (m - N)-th first text sequence, and the translation start condition can be used to represent the condition for starting the next round of translation process. For example, the translation start condition can be that the similarity degree between the hidden state corresponding to the m-th first text sequence and the hidden state corresponding to the (m - N)-th first text sequence is less than the first set threshold.
[0062] When the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence satisfies the translation start condition, it indicates that the hidden state corresponding to the M-th first text sequence received in the M-th translation start determination is similar to the hidden state corresponding to the M-th first text sequence received in the (M - N)-th translation start determination. That is, the overall semantics represented by the M-th first text sequence is similar to the overall semantics represented by the (M - N)-th first text sequence. Therefore, it can be determined that the semantics represented by the M-th first text sequence is stable and there is no content that may cause ambiguity. At this time, the next round of the translation process is started, and an accurate translation result can be obtained.
[0063] In some possible implementation manners, determine the similarity degree between the hidden state corresponding to the M-th first text sequence represented in at least one of the following forms and the hidden state corresponding to the (M - N)-th first text sequence: the change amplitude between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence, the linear transformation function for converting the hidden state corresponding to the M-th first text sequence into the hidden state corresponding to the (M - N)-th first text sequence, and the non-linear transformation function for converting the hidden state corresponding to the M-th first text sequence into the hidden state corresponding to the (M - N)-th first text sequence.
[0064] Among them, the change amplitude can be represented by the norm of a vector, for example, it is the 1-norm or 2-norm of the vector. The linear transformation function can be represented by a transformation matrix and a bias value, and the non-linear transformation function can be represented by a function corresponding to boosting and a random forest model or a neural network model.
[0065] Denote the hidden state corresponding to the M-th first text sequence as E(S j+1:j+M )[j + M], and denote the hidden state corresponding to the (M - N)-th first text sequence as E(S j+1:j+M-N )[j + M - N]. Then there is:
[0066] The change amplitude between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence is: |E(S j+1:j+M )[j + M] - E(S j+1:j+M-N )[j + M - N]|;
[0067] The linear transformation function for converting the hidden state corresponding to the M-th first text sequence into the hidden state corresponding to the (M - N)-th first text sequence is: W T E(S j+1:j+M )[j + M - N:j + M] + b, where W T is the transformation matrix and b is the bias value;
[0068] The non - linear transformation function for converting the hidden state corresponding to the M - th first text sequence into the hidden state corresponding to the (M - N) - th first text sequence is: g(E(S j+1:j+M )[j + M - N:j + M]) or h(E(S j+1:j+M )[j + M - N:j + M]), where the function g is the transformation function corresponding to the boosting and random forest model, and the function h is the transformation function corresponding to the neural network model.
[0069] Through one or more of the above forms, the similarity degree between the hidden state corresponding to the M - th first text sequence and the hidden state corresponding to the (M - N) - th first text sequence is quantitatively characterized, and based on the similarity degree, it is judged whether to start the next round of translation process. When the hidden state corresponding to the M - th first text sequence is relatively similar to the hidden state corresponding to the (M - N) - th first text sequence, the (n + 1) - th round of translation process is started.
[0070] In this case, when the timing to start the next round of translation process is reached, the target translation model is used to execute the (n + 1) - th round of translation process. In the (n + 1) - th round of translation process, the M - th first text sequence is translated, that is, the (j + 1) - th to the (j + m) - th tokens corresponding to the first text are translated.
[0071] In this way, although the first text input by the user is continuously received, when the timing to start the next round of translation process is reached, the target translation model is used to translate the text content (i.e., the M - th first text sequence) required for the next round of translation process, ensuring that in each round of translation process, the target translation model can translate short sentences with stable semantics and less prone to ambiguity, improving the accuracy of the translation result.
[0072] Taking n as 2 as an example, in the first round of translation process, the target translation model translates the 1 - 5th tokens corresponding to the first text. In the second round of translation process, the target translation model translates the 6 - 9th tokens corresponding to the first text, that is, i is 6 and j is 9. When M is 3 and N is 1, the M - th first text sequence is the 10 - 12th tokens corresponding to the first text, the (M - N) - th first text sequence is the 10 - 11th tokens corresponding to the first text, and the similarity degree between the hidden state corresponding to the M - th first text sequence and the hidden state corresponding to the (M - N) - th first text sequence meets the translation start condition, that is, the 10 - 12th tokens corresponding to the first text have stable semantics and belong to short sentences that can be translated. In this case, the target translation model is used to execute the third round of translation process, and in the third round of translation process, the 10 - 12th tokens corresponding to the first text are translated.
[0073] In specific implementation, the hidden state corresponding to the Mth first text sequence is sent to the output layer of the target translation model, and the translation result of the (n + 1)th round of translation process output by the output layer of the target translation model is received.
[0074] That is to say, since in the Mth translation start discrimination, the hidden state corresponding to the Mth first text sequence output by the hidden layer of the target translation model has been received, therefore, in the (n + 1)th round of translation process, only the hidden state corresponding to the Mth first text sequence needs to be directly sent to the output layer of the target translation model, and the translation result of the (n + 1)th round of translation process output by the target translation model can be obtained, that is, the result of translating the (j + 1)th to the (j + M)th tokens corresponding to the first text into the target language.
[0075] In the embodiment of the present application, the translation result of the (n + 1)th round of translation process includes: the pth to the qth tokens corresponding to the second text described in the target language, where 1 ≤ p < q. That is to say, in the previous n rounds of translation process, the first p - 1 tokens of the second text are translated, and by performing the (n + 1)th round of translation process, the pth to the qth tokens corresponding to the second text are translated.
[0076] Furthermore, the translation result of the (n + 1)th round of translation process is output in a streaming manner.
[0077] Since the text translation method provided by the embodiment of the present application can achieve real-time translation of the first text, therefore, after the target translation model completes the translation of the Mth first text sequence in the (n + 1)th round of translation process, the translation result of the (n + 1)th round of translation process (that is, the pth to the qth tokens corresponding to the second text) is presented to the user in a streaming output manner. In this way, the translation result of each round of translation process is presented in real time, achieving the real-time translation effect of "the user inputs while the translation result is presented".
[0078] In this method, for a translation model whose hidden layer only includes a decoder, by adding multiple translation start discriminations after each round of translation process, when the hidden state of the text sequence in two translation start discriminations represents that the text sequence that has been received but not yet translated (that is, the Mth text sequence) is semantically stable, a new round of translation process is started to translate the text sequence that has been received but not yet translated. In this way, the start timing of each round of translation process is accurately judged, and in the process of the user continuously inputting the text to be translated, smooth real-time translation is achieved through multiple rounds of translation process, and the translation model whose hidden layer only includes a decoder is applied to the real-time translation scenario, improving the applicability of the translation model whose hidden layer only includes a decoder.
[0079] Continue to refer to Figure 2 , after each round of translation process is started, multiple translation stop discriminations will also be performed to determine the stop timing of the current round of translation process.
[0080] Taking the (n + 1)-th round of the translation process as an example for illustration, the (n + 1)-th round of the translation process translates the M-th first text sequence into the p-th to q-th tokens corresponding to the second text. For the target translation model, although only the M-th first text sequence is sent to the target translation model for translation task processing in the (n + 1)-th round of the translation process, during the translation task processing of the target translation model, if the target translation model is not controlled to stop, the target translation model will continue to generate and output new tokens. That is to say, after the target translation model outputs the p-th to q-th tokens corresponding to the second text, it will continue to output other tokens that have nothing to do with the translation result of the first text.
[0081] In a real-time translation scenario, users do not want to receive other tokens that have nothing to do with the translation result of the first text. Therefore, in the embodiments of the present application, by performing multiple translation stop discriminations, the timing when the target translation model completes the (n + 1)-th round of the translation process is judged, and the (n + 1)-th round of the translation process is stopped to prevent the target translation model from generating other tokens that have nothing to do with the translation result of the first text.
[0082] Specifically, perform q - p + 2 times of translation stop discrimination. In response to the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence satisfying the translation stop condition, stop the (n + 1)-th round of the translation process.
[0083] Among them, the translation stop discrimination can be used to judge whether the target translation model has completed the translation of the first text sequence in the current round of the translation process and whether to stop the current round of the translation process. In other words, after the (n + 1)-th round of the translation process is started, perform multiple translation stop discriminations. When the (q - p + 2)-th translation stop discrimination indicates to stop the current round of the translation process, stop the (n + 1)-th round of the translation process, so that the target translation model stops generating and outputting new tokens.
[0084] Taking the k-th translation stop discrimination as an example for illustration, 1 ≤ k ≤ q - p + 2, the k-th translation stop discrimination includes: determining the k-th second text sequence, sending the k-th second text sequence to the target translation model, and receiving the hidden state corresponding to the k-th second text sequence output by the hidden layer of the target translation model.
[0085] Among them, the k-th second text sequence includes: the M-th first text sequence and k - 1 tokens corresponding to the second text already generated in the (n + 1)-th round of the translation process. That is to say, the second text sequence is composed of the tokens to be translated in the (n + 1)-th round of the translation process (i.e., the M-th first text sequence) and the tokens already translated by the target translation model in the (n + 1)-th round of the translation process.
[0086] Taking n as 2, p as 8, and q as 11 as an example for illustration, in the third round of the translation process, the M-th first text sequence is translated into the 8th to 11th tokens corresponding to the second text. In the first translation stop determination, the first second text sequence includes the M-th first text sequence. In the second translation stop determination, the second second text sequence includes the M-th first text sequence and 1 token corresponding to the second text generated in the third round of the translation process (i.e., the 8th token corresponding to the second text). In the third translation stop determination, the third second text sequence includes the M-th first text sequence and 2 tokens corresponding to the second text generated in the third round of the translation process (i.e., the 8th - 9th tokens corresponding to the second text), ……, in the fifth translation stop determination, the fifth second text sequence includes the M-th first text sequence and 4 tokens corresponding to the second text generated in the third round of the translation process (i.e., the 8th - 11th tokens corresponding to the second text).
[0087] The k-th translation stop determination further includes: determining whether the similarity between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence meets the translation stop condition. Among them, the similarity can be used to measure whether the hidden state corresponding to the (q - p + 2)-th second text sequence is similar to the hidden state corresponding to the (q - p + 2 - N)-th second text sequence, and the translation stop condition can be used to represent the condition for stopping the current round of the translation process. For example, the translation stop condition can be that the similarity between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence is greater than the second set threshold.
[0088] When the similarity between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence meets the translation stop condition, it indicates that the hidden state corresponding to the (q - p + 2)-th second text sequence received in the (q - p + 2)-th translation stop determination is not similar to the hidden state corresponding to the (q - p + 2 - N)-th second text sequence received in the (q - p + 2 - N)-th translation stop determination. That is, the overall semantics represented by the (q - p + 2)-th second text sequence is inconsistent with the overall semantics represented by the (q - p + 2 - N)-th second text sequence. Therefore, it can be determined that the target translation model has completed the translation of the M-th first text sequence and has started to generate other tokens that are not related to the translation result of the first text. At this time, the current round of the translation process is stopped to prevent the target translation model from generating other content after completing the content that should be translated in the (n + 1)-th round of the translation process.
[0089] In some possible implementation manners, the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence characterized in at least one of the following forms and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence is determined: the change amplitude between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence, the linear transformation function for converting the hidden state corresponding to the (q - p + 2)-th second text sequence into the hidden state corresponding to the (q - p + 2 - N)-th second text sequence, and the non-linear transformation function for converting the hidden state corresponding to the (q - p + 2)-th second text sequence into the hidden state corresponding to the (q - p + 2 - N)-th second text sequence.
[0090] Wherein, the change amplitude can be represented by the norm of a vector, for example, the 1-norm or 2-norm of the vector. The linear transformation function can be represented by a transformation matrix and a bias value. The non-linear transformation function can be represented by a function corresponding to boosting and a random forest model or a neural network model. The specific content is similar to the content for characterizing the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence in the foregoing, and will not be elaborated herein.
[0091] In this way, when reaching the time to stop the current round of the translation process, the target translation model is timely stopped from continuing to generate new tokens. On the basis of ensuring that the target translation model completes the translation of the M-th first text sequence in the (n + 1)-th round of the translation process, the target translation model is prevented from outputting content irrelevant to the translation task of the first text.
[0092] Further, in response to the stop of the (n + 1)-th round of the translation process and the non-receipt of the (j + M + 1)-th token corresponding to the first text, the translation of the first text is completed.
[0093] That is to say, after the target translation model completes the translation of the (j + 1)-th to (j + M)-th tokens corresponding to the first text in the (n + 1)-th round of the translation process, if the user does not continue to input a new first text, it indicates that the target translation model has completed the translation of all tokens corresponding to the first text, and the translation task of the first text ends.
[0094] As described above in conjunction with Figures 1 to 3 the text translation method provided by the embodiments of the present application has been introduced in detail. Next, the apparatuses and devices provided by the embodiments of the present application will be introduced in conjunction with the accompanying drawings.
[0095] See Figure 4 the structural schematic diagram of the text translation apparatus shown. The apparatus 40 includes:
[0096] An obtaining module 401, configured to continuously receive the first text described in the source language input by the user; wherein, the first text corresponds to a plurality of tokens.
[0097] A discrimination module 402, configured to perform M translation start discriminations in response to the stop of the n-th round of translation process; wherein, the n-th round of translation process is used to translate the i-th to j-th tokens corresponding to the first text, 1 ≤ i < j;
[0098] A translation module 403, configured to perform the (n + 1)-th round of translation process by using a target translation model in response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence satisfying a translation start condition; wherein, the (n + 1)-th round of translation process is used to translate the M-th first text sequence, N < M;
[0099] Wherein, the m-th translation start discrimination includes: determining the m-th first text sequence, sending the m-th first text sequence to the target translation model, and receiving the hidden state corresponding to the m-th first text sequence output by the hidden layer of the target translation model, 1 ≤ m ≤ M, and the hidden layer of the target translation model is composed of multiple decoders;
[0100] The m-th first text sequence includes: the (j + 1)-th to (j + m)-th tokens corresponding to the first text.
[0101] In some possible implementation manners, the translation module 403 is specifically configured to:
[0102] Send the hidden state corresponding to the M-th first text sequence to the output layer of the target translation model, and receive the translation result of the (n + 1)-th round of translation process output by the output layer of the target translation model; wherein, the translation result of the (n + 1)-th round of translation process includes: the p-th to q-th tokens corresponding to the second text described in the target language, 1 ≤ p < q.
[0103] In some possible implementation manners, the apparatus 40 further includes an output module, and the output module is configured to:
[0104] Stream out the translation result of the (n + 1)-th round of translation process.
[0105] In some possible implementation manners, the (n + 1)-th round of translation process translates the M-th first text sequence into the p-th to q-th tokens corresponding to the second text, 1 ≤ p < q; the discrimination module 402 is further configured to:
[0106] Perform q - p + 2 translation stop discriminations;
[0107] Stop the (n + 1)-th round of translation process in response to the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence satisfying a translation stop condition;
[0108] Among them, the k-th translation stop determination includes: determining the k-th second text sequence, sending the k-th second text sequence to the target translation model, and receiving the hidden state corresponding to the k-th second text sequence output by the hidden layer of the target translation model, where 1 ≤ k ≤ q - p + 2;
[0109] The k-th second text sequence includes: the M-th first text sequence and k - 1 word elements corresponding to the second text generated in the (n + 1)-th translation process.
[0110] In some possible implementation manners, the translation module 403 is specifically configured to:
[0111] Determine the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence represented in at least one of the following forms: the change amplitude of the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence, the linear transformation function for converting the hidden state corresponding to the M-th first text sequence into the hidden state corresponding to the (M - N)-th first text sequence, and the non-linear transformation function for converting the hidden state corresponding to the M-th first text sequence into the hidden state corresponding to the (M - N)-th first text sequence;
[0112] In response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence satisfying the translation start condition, use the target translation model to perform the (n + 1)-th translation process.
[0113] In some possible implementation manners, the discrimination module 402 is specifically configured to:
[0114] Determine the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence represented in at least one of the following forms: the change amplitude of the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence, the linear transformation function for converting the hidden state corresponding to the (q - p + 2)-th second text sequence into the hidden state corresponding to the (q - p + 2 - N)-th second text sequence, and the non-linear transformation function for converting the hidden state corresponding to the (q - p + 2)-th second text sequence into the hidden state corresponding to the (q - p + 2 - N)-th second text sequence;
[0115] In response to the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence satisfying the translation stop condition, stop the (n + 1)-th translation process.
[0116] In some possible implementations, the discrimination module 402 is further configured to:
[0117] In response to the stop of the (n + 1)-th round of translation process and the non-receipt of the (j + M + 1)-th token corresponding to the first text, complete the translation of the first text.
[0118] The text translation device 40 according to the embodiments of the present application may correspond to executing the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module / unit of the text translation device 40 are respectively for implementing Figure 2 the corresponding processes of the respective methods in the illustrated embodiments. For the sake of brevity, they will not be elaborated herein.
[0119] The embodiments of the present application further provide an electronic device. Specifically, the electronic device is used to implement the functions of the text translation device 40 in the Figure 4 illustrated embodiments.
[0120] Figure 5 A schematic structural diagram of an electronic device 500 is provided, as Figure 5 shown. The electronic device 500 includes a bus 501, a processor 502, a communication interface 503, and a memory 504. The processor 502, the memory 504, and the communication interface 503 communicate with each other through the bus 501.
[0121] The bus 501 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a thick line is used herein, but it does not mean that there is only one bus or one type of bus.
[0122] The processor 502 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0123] The communication interface 503 is used for external communication. For example, the communication interface 503 may be used for communicating with a terminal.
[0124] The memory 504 may include volatile memory, such as random access memory (RAM). The memory 504 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0125] Executable code is stored in the memory 504, and the processor 502 executes the executable code to perform the foregoing text translation method.
[0126] Specifically, in the case of implementing Figure 4 the illustrated embodiments, and Figure 4 when each module or unit of the text translation device 40 described in the embodiments is implemented by software, the software or program code required to execute Figure 4 the functions of each module / unit in may be partially or wholly stored in the memory 504. The processor 502 executes the program code corresponding to each unit stored in the memory 504 to perform the foregoing text translation method.
[0127] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center including one or more available media. The available media may be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid state drives), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the text translation method applied to the text translation device 40 described above.
[0128] An embodiment of the present application also provides a computer program product that includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions according to the embodiments of the present application are generated in whole or in part.
[0129] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).
[0130] When the computer program product is executed by a computer, the computer executes any one of the foregoing text translation methods. The computer program product may be a software installation package. In the case where any one of the foregoing text translation methods needs to be used, the computer program product can be downloaded and executed on the computer.
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0132] The units involved in the embodiments described in the present application can be implemented in software or in hardware. Among them, the name of the unit / module does not constitute a limitation to the unit itself in some cases.
[0133] The functions described above herein can be at least partially performed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0134] In the context of the embodiments of the present application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0135] It should be noted that the embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems or apparatuses disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method section.
[0136] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0137] It should also be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0138] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the art.
[0139] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A text translation method, characterized in that, The method includes: Continuously receiving a first text described in the source language input by the user; wherein, the first text corresponds to a plurality of tokens; In response to the stop of the n-th round of translation process, performing M translation start discriminations; wherein, the n-th round of translation process is used to translate the i-th to j-th tokens corresponding to the first text, 1 ≤ i < j; In response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the M-N-th first text sequence satisfying the translation start condition, using the target translation model to perform the (n + 1)-th round of translation process; wherein, the (n + 1)-th round of translation process is used to translate the M-th first text sequence, N < M; Wherein, the m-th translation start discrimination includes: determining the m-th first text sequence, sending the m-th first text sequence to the target translation model, and receiving the hidden state corresponding to the m-th first text sequence output by the hidden layer of the target translation model, 1 ≤ m ≤ M, and the hidden layer of the target translation model consists of a plurality of decoders; The m-th first text sequence includes: the (j + 1)-th to (j + m)-th tokens corresponding to the first text.
2. The method according to claim 1, wherein The using the target translation model to perform the (n + 1)-th round of translation process includes: Sending the hidden state corresponding to the M-th first text sequence to the output layer of the target translation model, and receiving the translation result of the (n + 1)-th round of translation process output by the output layer of the target translation model; wherein, the translation result of the (n + 1)-th round of translation process includes: the p-th to q-th tokens corresponding to the second text described in the target language, 1 ≤ p < q.
3. The method according to claim 2, wherein The method further includes: Streaming the translation result of the (n + 1)-th round of translation process.
4. The method according to claim 1, characterized in that, The (n + 1)-th round of translation process translates the M-th first text sequence into the p-th to q-th tokens corresponding to the second text, 1 ≤ p < q; the method further includes: Performing q - p + 2 translation stop discriminations; In response to the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence satisfying the translation stop condition, stopping the (n + 1)-th round of translation process; Wherein, the k-th translation stop discrimination includes: determining the k-th second text sequence, sending the k-th second text sequence to the target translation model, and receiving the hidden state corresponding to the k-th second text sequence output by the hidden layer of the target translation model, 1 ≤ k ≤ q - p + 2; The k-th second text sequence includes: the M-th first text sequence and the k - 1 tokens corresponding to the second text generated in the (n + 1)-th round of translation process.
5. The method according to any one of claims 1 to 4, characterized in that The in response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the M-N-th first text sequence satisfying the translation start condition, using the target translation model to perform the (n + 1)-th round of translation process includes: Determine the similarity degree between the hidden state corresponding to the M-th first text sequence characterized in at least one of the following forms and the hidden state corresponding to the (M - N)-th first text sequence: the change amplitude between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence, the linear transformation function for converting the hidden state corresponding to the M-th first text sequence into the hidden state corresponding to the (M - N)-th first text sequence, and the non-linear transformation function for converting the hidden state corresponding to the M-th first text sequence into the hidden state corresponding to the (M - N)-th first text sequence; In response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence satisfying the translation start condition, use the target translation model to perform the (n + 1)-th translation process.
6. The method according to claim 4, wherein The stopping of the (n + 1)-th translation process in response to the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence satisfying the translation stop condition includes: Determine the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence characterized in at least one of the following forms and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence: the change amplitude between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence, the linear transformation function for converting the hidden state corresponding to the (q - p + 2)-th second text sequence into the hidden state corresponding to the (q - p + 2 - N)-th second text sequence, and the non-linear transformation function for converting the hidden state corresponding to the (q - p + 2)-th second text sequence into the hidden state corresponding to the (q - p + 2 - N)-th second text sequence; In response to the similarity degree between the hidden state corresponding to the (q - p + 2)-th second text sequence and the hidden state corresponding to the (q - p + 2 - N)-th second text sequence satisfying the translation stop condition, stop the (n + 1)-th translation process.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: In response to the (n + 1)-th translation process stopping and not receiving the (j + M + 1)-th token corresponding to the first text, complete the translation of the first text.
8. A language model training device, characterized in that, The device includes: An acquisition module, configured to continuously receive the first text described in the source language input by the user; wherein, the first text corresponds to multiple tokens; A discrimination module, configured to perform M translation start discriminations in response to the stop of the n-th translation process; wherein, the n-th translation process is used to translate the i-th to j-th tokens corresponding to the first text, 1 ≤ i < j; A translation module, configured to use the target translation model to perform the (n + 1)-th translation process in response to the similarity degree between the hidden state corresponding to the M-th first text sequence and the hidden state corresponding to the (M - N)-th first text sequence satisfying the translation start condition; wherein, the (n + 1)-th translation process is used to translate the M-th first text sequence, N < M; Among them, the determination of the m-th translation start includes: determining the m-th first text sequence, sending the m-th first text sequence to the target translation model, and receiving the hidden state corresponding to the m-th first text sequence output by the hidden layer of the target translation model, where 1 ≤ m ≤ M, and the hidden layer of the target translation model is composed of multiple decoders; The m-th first text sequence includes: the (j + 1)-th to (j + m)-th tokens corresponding to the first text.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory; The processor is configured to execute the instructions stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It includes instructions that direct the electronic device to execute the method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, The computer program product includes computer-readable instructions for implementing the method according to any one of claims 1 to 7.