Passage-level translation method, translation model training method and device

The chapter-level translation method addresses the lack of contextual dependency in NMT by training a model to capture semantic information, improving translation quality and consistency across sentences.

CN114065778BActive Publication Date: 2025-07-15BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010763386.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-07-15
Estimated Expiration
2040-12-25

AI Technical Summary

Technical Problem

In the prior art, the translation effect of chapter-level text is poor and the contextual dependencies are lacking, resulting in inconsistent translation results.

Method used

The target chapter translation model learns the context semantic information in the chapter-level training corpus, combines the neural machine translation model and the context prediction model for joint training or pre-training and fine-tuning to capture the dependence between sentences.

Benefits of technology

It improves the accuracy and consistency of chapter-level text translation, eliminates translation ambiguity, and enhances the consistency of translation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065778B_ABST
    Figure CN114065778B_ABST
Patent Text Reader

Abstract

The present invention discloses a passage translation method, a translation model training method and a device, which are applied to the field of machine translation. For each sentence to be translated in a passage to be translated, a sentence representation containing context semantic information of the sentence to be translated is obtained through a target passage translation model, and the sentence to be translated is translated based on the sentence representation; according to the translation results of each sentence to be translated in the passage to be translated, a passage translation result corresponding to the passage to be translated is obtained. By means of the present invention, the effect of passage-level text translation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural machine translation, and particularly to a passage-level translation method, a translation model training method and a device. Background Art

[0002] In recent years, with the emergence of the Transformer framework, Neural Machine Translation (NMT) has achieved leapfrog development, and the translation quality has also been greatly improved. As more and more enterprises go global, NMT may have a huge impact on the translation industry. Different from traditional statistical machine translation, NMT uses neural network-based technologies to achieve more contextually accurate translations.

[0003] Since NMT can translate an entire sentence at once, the output of NMT can be similar to human translation. Currently, for passage-level translation, usually a single sentence is used as the translation unit, and then the translation results of each sentence are concatenated to obtain the final passage translation result. Due to the lack of context dependency, the sentence-level translation system performs poorly in translating passage-level texts and still cannot meet people's needs. Summary of the Invention

[0004] Embodiments of the present invention provide a passage-level translation method, a translation model training method and a device to solve the technical problem of poor translation effect of passage-level texts in the prior art.

[0005] In a first aspect, an embodiment of the present invention provides a passage translation method, including:

[0006] For each sentence to be translated in the passage to be translated, obtain a sentence representation including context semantic information of the sentence to be translated through a target passage translation model, and translate the sentence to be translated based on the sentence representation;

[0007] Obtain a passage translation result corresponding to the passage to be translated according to the translation results of each sentence to be translated in the passage to be translated.

[0008] Optionally, the target passage translation model is obtained by learning sentence representations including context semantic information in passage-level training corpora, where the passage-level training corpora are passage-level parallel corpora and / or passage-level monolingual corpora.

[0009] Optionally, the obtaining the target passage translation model by learning sentence representations including context semantic information in passage-level training corpora includes:

[0010] For the case where the text-level training corpus is a text-level parallel corpus, jointly train a context prediction model and a neural machine translation model based on the text-level parallel corpus to obtain a target text translation model corresponding to the neural machine translation model, where the context prediction model is used to learn sentence representations containing context semantic information from the source-side sentences of the text-level parallel corpus; or

[0011] For the case where the text-level training corpus is a text-level monolingual corpus, pre-train a pre-trained model based on the text-level monolingual corpus, and fine-tune a target combined model according to the pre-trained pre-trained model to obtain a target text translation model corresponding to the target combined model, where the pre-trained model is used to learn sentence representations containing context semantic information from the source-side sentences of the text-level monolingual corpus, and the target combined model includes a neural machine translation model and a source-side context encoder.

[0012] Optionally, the step of jointly training the context prediction model and the neural machine translation model based on the text-level parallel corpus to obtain a target text translation model corresponding to the neural machine translation model includes:[[]]

[0013] Jointly train the neural machine translation model and the context prediction model using the obtained first text-level parallel corpus until a trained joint model is obtained. The first parallel sentence pairs in the first text-level parallel corpus include the current source-side sentence, the source-side context sentence for the current source-side sentence, and the target-side sentence.

[0014] Extract the trained neural machine translation model from the trained joint model.

[0015] Continue to train the trained neural machine translation model using the obtained second text-level parallel corpus until a target text-level translation model corresponding to the neural machine translation model is obtained, where the second parallel sentence pairs in the second text-level parallel corpus include the current source-side sentence and the target-side sentence for the current source-side sentence.

[0016] Optionally, the neural machine translation model and the context prediction model share the same source-side encoder, the neural machine translation model further includes a target-side decoder, and the context prediction model further includes a source-side context decoder.

[0017] Optionally, the step of jointly training the neural machine translation model and the context prediction model using the obtained first text-level parallel corpus includes multiple joint iterative trainings of the neural machine translation model and the context prediction model; where any one joint iterative training includes:[[]]

[0018] Encode the current source sentence of the first parallel sentence pair in the first chapter-level parallel corpus through the source encoder;

[0019] Decode the encoding result of the same source encoder through the target decoder and the source context decoder respectively to predict the target sentence and the source context sentence corresponding to the current source sentence;

[0020] Jointly determine the joint loss gradient according to the predicted source context sentence, the predicted target prediction sentence, and the first parallel sentence pair;

[0021] Update the model parameters of the neural machine translation model and the model parameters of the context prediction model based on the joint loss gradient.

[0022] Optionally, before jointly training the neural machine translation model and the context prediction model using the obtained first chapter-level parallel corpus, the method further includes:

[0023] Pre-train a pre-trained model using the chapter-level monolingual corpus to obtain a pre-trained pre-trained model, where the pre-trained model is used to learn sentence representations containing context semantic information from the sentences at the source end of the chapter-level monolingual corpus;

[0024] Initialize the neural machine translation model and the context prediction model based on the pre-trained pre-trained model.

[0025] Optionally, the fine-tuning the target combination model according to the pre-trained pre-trained model to obtain a target chapter translation model corresponding to the target combination model includes:

[0026] Initialize the target combination model based on the pre-trained pre-trained model;

[0027] Train the initialized target combination model according to the chapter-level monolingual corpus to obtain a target chapter translation model corresponding to the target combination model.

[0028] In a second aspect, an embodiment of the present invention provides a translation model training method, including:

[0029] Jointly train a context prediction model and a neural machine translation model based on the chapter-level parallel corpus to obtain a target chapter translation model corresponding to the neural machine translation model, where the context prediction model is used to learn sentence representations containing context semantic information from the source sentences of the chapter-level parallel corpus.

[0030] Optionally, the step of jointly training the context prediction model and the neural machine translation model based on the passage-level parallel corpus to obtain a target passage translation model corresponding to the neural machine translation model includes:

[0031] Jointly training the neural machine translation model and the context prediction model by using the obtained first passage-level parallel corpus until a trained joint model is obtained, where the first parallel sentence pairs in the first passage-level parallel corpus include the current source-side sentence, the source-side context sentence, and the target-side sentence for the current source-side sentence;

[0032] Extracting the trained neural machine translation model from the trained joint model;

[0033] Continuing to train the trained neural machine translation model by using the obtained second passage-level parallel corpus until a target passage-level translation model corresponding to the neural machine translation model is obtained, where the second parallel sentence pairs in the second passage-level parallel corpus include the current source-side sentence and the target-side sentence for the current source-side sentence.

[0034] Optionally, before jointly training the neural machine translation model and the context prediction model by using the obtained first passage-level parallel corpus, it further includes:

[0035] Training a pre-training model by using a passage-level monolingual corpus to obtain a pre-trained pre-training model, where the pre-training model is used to learn sentence representations containing context semantic information from the sentences on the source side of the passage-level monolingual corpus;

[0036] Initializing the neural machine translation model and the context prediction model based on the pre-trained pre-training model.

[0037] In a third aspect, an embodiment of the present invention provides a translation model training method, including:

[0038] Pre-training a pre-training model based on a passage-level monolingual corpus;

[0039] Fine-tuning a target combination model according to the pre-trained pre-training model to obtain a target passage translation model corresponding to the target combination model, where the pre-training model is used to learn sentence representations containing context semantic information from the source-side sentences of the passage-level monolingual corpus, and the target combination model includes a neural machine translation model and a source-side context encoder.

[0040] In a fourth aspect, an embodiment of the present invention provides a passage translation device, including:

[0041] A sentence translation unit, which is used for each sentence to be translated in the passage to be translated, to obtain a sentence representation including context semantic information of the sentence to be translated through a target passage translation model, and to translate the sentence to be translated based on the sentence representation;

[0042] A translation result forming unit, which is used to obtain a passage translation result corresponding to the passage to be translated according to the translation results of each sentence to be translated in the passage to be translated.

[0043] Optionally, the device further includes:

[0044] A model training unit, which is used to obtain the target passage translation model by learning sentence representations including context semantic information in passage-level training corpora, where the passage-level training corpora are passage-level parallel corpora and / or passage-level monolingual corpora.

[0045] Optionally, the model training unit includes:

[0046] A first training unit, which is used for the passage-level training corpus being a passage-level parallel corpus, to jointly train a context prediction model and a neural machine translation model based on the passage-level parallel corpus, and to obtain a target passage translation model corresponding to the neural machine translation model, where the context prediction model is used to learn sentence representations including context semantic information from source-side sentences in the passage-level parallel corpus;

[0047] Or the model training unit includes:

[0048] A second training unit, which is used for the passage-level training corpus being a passage-level monolingual corpus, to pre-train a pre-trained model based on the passage-level monolingual corpus;

[0049] A model fine-tuning unit, which is used to fine-tune a target combined model according to the pre-trained pre-trained model to obtain a target passage translation model corresponding to the target combined model, where the pre-trained model is used to learn sentence representations including context semantic information from source-side sentences in the passage-level monolingual corpus, and the target combined model includes a neural machine translation model and a source-side context encoder.

[0050] Optionally, the first training unit includes:

[0051] A joint training subunit, which is used to jointly train the neural machine translation model and the context prediction model by using the obtained first passage-level parallel corpus until a trained joint model is obtained, and the first parallel sentence pairs in the first passage-level parallel corpus include the current source-side sentence and the source-side context sentence and the target-side sentence for the current source-side sentence;

[0052] A model extraction subunit, configured to extract a trained neural machine translation model from the trained joint model;

[0053] A continuous training subunit, configured to continuously train the trained neural machine translation model by using the obtained second passage-level parallel corpus until a target passage-level translation model corresponding to the neural machine translation model is obtained, where the second parallel sentence pair in the second passage-level parallel corpus includes a current source-side sentence and a target-side sentence for the current source-side sentence.

[0054] Optionally, the neural machine translation model and the context prediction model share the same source-side encoder, the neural machine translation model further includes a target-side decoder, and the context prediction model further includes a source-side context decoder.

[0055] Optionally, the joint training subunit is configured to perform multiple joint iterative trainings on the neural machine translation model and the context prediction model; for any one of the joint iterative trainings, the joint training subunit is specifically configured to:

[0056] Encode the current source-side sentence of the first parallel sentence pair in the first passage-level parallel corpus through the source-side encoder;

[0057] Decode the encoding result of the same source-side encoder through the target-side decoder and the source-side context decoder respectively to predict a target-side sentence and a source-side context sentence corresponding to the current source-side sentence;

[0058] Jointly determine a joint loss gradient according to the predicted source-side context sentence, the predicted target-side prediction sentence, and the first parallel sentence pair;

[0059] Update the model parameters of the neural machine translation model and the model parameters of the context prediction model based on the joint loss gradient.

[0060] Optionally, the apparatus further includes:

[0061] A first pre-training unit, configured to train a pre-training model by using the passage-level monolingual corpus to obtain a pre-trained pre-training model, where the pre-training model is configured to learn a sentence representation including context semantic information from the sentences on the source side of the passage-level monolingual corpus;

[0062] An initialization unit, configured to initialize the neural machine translation model and the context prediction model based on the pre-trained pre-training model.

[0063] Optionally, the model fine-tuning unit includes:

[0064] An initialization subunit, configured to initialize the target combined model based on the pre-trained pre-trained model;

[0065] A fine-tuning training unit, configured to train the target combined model according to the passage-level monolingual corpus to obtain a target passage translation model corresponding to the target combined model.

[0066] In a fifth aspect, an embodiment of the present invention provides a translation model training apparatus, including:

[0067] A first training unit, configured to jointly train a context prediction model and a neural machine translation model based on the passage-level parallel corpus to obtain a target passage translation model corresponding to the neural machine translation model, where the context prediction model is used to learn a sentence representation including context semantic information from source-side sentences of the passage-level parallel corpus.

[0068] Optionally, the first training unit includes:

[0069] A joint training subunit, configured to jointly train the neural machine translation model and the context prediction model by using the obtained first passage-level parallel corpus until a trained joint model is obtained, where the first parallel sentence pairs in the first passage-level parallel corpus include a current source-side sentence, a source-side context sentence, and a target-side sentence for the current source-side sentence;

[0070] A model extraction subunit, configured to extract a trained neural machine translation model from the trained joint model;

[0071] A continuous training subunit, configured to continue training the trained neural machine translation model by using the obtained second passage-level parallel corpus until a target passage-level translation model corresponding to the neural machine translation model is obtained, where the second parallel sentence pairs in the second passage-level parallel corpus include a current source-side sentence and a target-side sentence for the current source-side sentence.

[0072] Optionally, the apparatus further includes:

[0073] A first pre-training unit, configured to train a pre-trained model by using a passage-level monolingual corpus to obtain a pre-trained pre-trained model, where the pre-trained model is used to learn a sentence representation including context semantic information from source-side sentences of the passage-level monolingual corpus;

[0074] An initialization unit, configured to initialize the neural machine translation model and the context prediction model based on the pre-trained pre-trained model.

[0075] In a sixth aspect, an embodiment of the present invention provides a translation model training apparatus, including:

[0076] An initialization subunit, configured to initialize the target combined model based on the pre-trained model after pre-training;

[0077] A fine-tuning training unit, configured to train the target combined model according to the passage-level monolingual corpus to obtain a target passage translation model corresponding to the target combined model.

[0078] In a seventh aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the method described in any one of the first aspects is implemented.

[0079] In an eighth aspect, an embodiment of the present invention provides an electronic device, including a memory, and one or more programs, wherein one or more programs are stored in the memory and are configured to be executed by one or more processors to perform the operation instructions included in the one or more programs for performing the method described in any one of the first aspects.

[0080] One or more technical solutions provided by the embodiments of the present invention at least achieve the following beneficial effects:

[0081] For the sentence to be translated in the passage to be translated, the embodiment of the present invention obtains a sentence representation including context semantic information of the sentence to be translated through the target passage translation model, and translates the sentence to be translated based on the sentence representation; since the target passage translation model captures the context semantic information of the sentence to be translated in the passage, it can eliminate translation ambiguities and keep the translation results of the same words more consistent in different sentences of the passage, thereby improving the translation effect for passage-level texts. Description of the Drawings

[0082] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0083] Figure 1 It is a flowchart of the passage translation method in the embodiment of the present invention;

[0084] Figure 2 It is a schematic structural diagram of the combined model in the embodiment of the present invention;

[0085] Figure 3 It is a schematic structural diagram of the pre-trained model in the embodiment of the present invention;

[0086] Figure 4Schematic flowchart of the joint training method in the embodiments of the present invention;

[0087] Figure 5 Schematic structural diagram of the target combined model in the embodiments of the present invention;

[0088] Figure 6 Schematic flowchart of the pre-training + fine-tuning method in the embodiments of the present invention;

[0089] Figure 7 Functional module diagram of the passage translation device in the embodiments of the present invention;

[0090] Figure 8 Schematic structural diagram of the electronic device in the embodiments of the present invention. Detailed implementation manners

[0091] To solve the above technical problems, the general idea of the technical solution provided by the embodiments of the present invention is as follows: for the sentence to be translated in the passage to be translated, the sentence representation including the context semantic information of the sentence to be translated is obtained through the target passage translation model, and the sentence to be translated is translated based on the sentence representation; thus, the target passage translation model captures the inter-sentence dependency relationship of the sentence to be translated at the passage level, so as to improve the translation effect of the passage-level text.

[0092] To make the objectives, technical solutions, and advantages of the embodiments of this specification clearer, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are some, but not all, of the embodiments of this specification. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts shall fall within the protection scope of this specification.

[0093] In a first aspect, the embodiments of the present invention provide a passage translation method. Referring to Figure 1 as shown, the passage translation method includes the following steps S10 to S11:

[0094] S10. For each sentence to be translated in the passage to be translated, obtain the sentence representation including the context semantic information of the sentence to be translated through the target passage translation model, and translate the sentence to be translated based on the sentence representation;

[0095] S11. Obtain the passage translation result corresponding to the passage to be translated according to the translation results of each sentence to be translated in the passage to be translated.

[0096] In the target passage translation model according to the embodiments of the present invention, the model structure at least includes a neural machine translation model, and the training of the target passage translation model is to learn the sentence representations of the passage-level training corpus containing context semantic information. By learning the sentence representations of the passage-level training corpus containing context semantic information, the trained target passage translation model can perform "contextualized" sentence representations on the source-side sentences, thereby helping to improve the ability of the passage-level neural machine to perform passage translation.

[0097] In the embodiments of the present invention, the neural machine translation model in the target passage translation model can be a translation model based on the Transformer framework. Among them, the Transformer architecture is a translation architecture based on attention (self-attention mechanism) composed of an encoder-decoder. Of course, according to actual needs, other types of neural machine translation models can also be selected, such as neural machine translation models with an RNN (recurrent neural network) architecture and other translation models in the encoder-decoder series.

[0098] It should be noted that the training of the target passage translation model in the embodiments of the present invention and the passage translation method described in the above steps S10-S11 belong to independent execution processes and can be performed on different electronic devices. That is, the target passage translation model trained on one electronic device can be applied to another electronic device for the translation of passage-level texts.

[0099] In specific implementation, the passage-level training corpus for training the target passage translation model can be a passage-level parallel corpus or a passage-level monolingual corpus. According to the different passage-level training corpora used, the training process of the target passage translation model is correspondingly different. Below, a training method is given respectively for training the target passage translation model with a passage-level parallel corpus or a passage-level monolingual corpus:

[0100] Training Method 1: Joint Training Method

[0101] Regarding the passage-level training corpus being a passage-level parallel corpus, a joint training method can be used to obtain the target passage translation model. The specific method is: jointly train the neural machine translation model and the context prediction model using the passage-level parallel corpus to obtain the target passage translation model corresponding to the neural machine translation model.

[0102] Specifically, the model structure of the finally obtained target passage translation model is the same as that of the neural machine translation model. Among them, the context prediction model is used to learn the sentence representations containing context semantic information from the source-side sentences of the passage-level parallel corpus.

[0103] Specifically, a target passage translation model corresponding to a neural machine translation model is obtained through a joint training method, which specifically includes the following steps A1 to A3:

[0104] Step A1: Jointly train the neural machine translation model and the context prediction model using the obtained first passage-level parallel corpus until a trained joint model is obtained.

[0105] Specifically, the first passage-level parallel corpus contains a certain number of first parallel sentence pairs. Among them, the first parallel sentence pair includes a source sentence, a source context sentence for the source sentence, and a target sentence. The source context sentence includes the previous source sentence and the next source sentence of the current source sentence. For example: The first parallel sentence pair is expressed as (s i , s i-1 , s i+1 , y i ), where s i is the current source sentence. s i-1 is the previous source sentence of s i , s i+1 is the next source sentence of s i , and y i is the target sentence of s i .

[0106] The joint model includes a neural machine translation model and a context prediction model. During the joint training process, the target sentence of the current source sentence is predicted through the neural machine translation model, and the source context sentence of the current source sentence is predicted through the context prediction model.

[0107] Among them, as shown in Figure 2 , Figure 2 is a schematic structural diagram of the joint model in the embodiment of the present invention. The joint model includes a source encoder and three decoders sharing the source encoder. Specifically, the neural machine translation model and the context prediction model share the same source encoder. The neural machine translation model further includes a target decoder, and the context prediction model further includes a source context decoder. The source context decoder of the context prediction model includes a pre-encoder and a next-encoder. The training of the joint model is a process of jointly training the neural machine translation model and the context prediction model, specifically including multiple joint iterative trainings of the neural machine translation model and the context prediction model; as shown in Figure 4 , Figure 4 is a schematic flow diagram of the joint training method in the embodiment of the present invention. Any joint iterative training includes the following steps A11 to A13:

[0108] Step A11: Encode the current source sentence of the first parallel sentence pair in the first chapter-level parallel corpus through the source encoder, and decode the encoding results of the same source encoder through the target decoder and the source context decoder respectively to predict the target sentence and the source context sentence corresponding to the current source sentence.

[0109] More specifically, as shown in Figure 2 , based on the context prediction model in the embodiments of the present invention including a front decoder and a rear decoder, step A11 specifically includes: encoding the current source sentence of the first parallel sentence pair through the source encoder to obtain a source sentence encoding vector; decoding the source sentence encoding vector through the front decoder to predict the previous source sentence, decoding the source sentence encoding vector through the rear decoder to predict the next source sentence, and decoding the source sentence encoding vector through the target decoder to predict the target sentence.

[0110] In the specific implementation process, the front and rear decoders of the context prediction model and the source encoder share the word vectors of the source sentence, that is, after the source sentence is represented as word vectors, it will be passed into the source encoder, the front decoder and the rear decoder as shown in Figure 2 .

[0111] Step A12: jointly determine the joint loss gradient according to the predicted source context sentence, the predicted target prediction sentence, and the first parallel sentence pair.

[0112] It should be noted that M first parallel sentence pairs are used in one joint iteration process, and M is a positive integer. The calculation formula of the joint loss function is as follows:

[0113] Loss = Loss_tgt + μ * Loss_pre + λ * Loss_next

[0114] Among them, Loss_tgt is the loss of the target decoder predicting M target sentences; Loss_pre is the loss of the front decoder predicting M previous source sentences corresponding to the M target sentences, and Loss_next is the loss of the rear decoder predicting M next source sentences corresponding to the M target sentences. The joint loss gradient is determined according to the joint loss function.

[0115] Step A13: update the model parameters of the neural machine translation model and the context prediction model based on the joint loss gradient. That is, update the model parameters of the joint model as shown in Figure 2 .

[0116] The joint model is continuously iteratively trained by repeating the above steps A11 to A13 until the training is ended when the first preset end condition is satisfied, and a trained joint model is obtained. Specifically, the joint training can be ended when a preset iteration number threshold is reached or convergence occurs, and a trained joint model is obtained.

[0117] After the trained joint model is obtained, step A2 is executed: extracting the trained neural machine translation model from the trained joint model. Specifically, as Figure 2 shown, the front decoder and the back decoder can be removed from the joint model, and what remains is an encoder-decoder network composed of a source-side decoder and a target-side decoder, thereby obtaining the trained neural machine translation model.

[0118] After step A2, step A3 is executed: continuing to train the extracted trained neural machine translation model by using the obtained second chapter parallel corpus until a target chapter-level translation model corresponding to the neural machine translation model is obtained.

[0119] Specifically, the second chapter parallel corpus includes a certain number of second parallel sentence pairs, where the second parallel sentence pair includes the current source-side sentence and the target-side sentence for the current source-side sentence. For example, the second parallel sentence pair can be expressed as {(s i ,y i )}, where s i is the current source-side sentence, and y i is the target-side sentence of s i .

[0120] Step A3 specifically includes: A31, encoding the current source-side sentence in the second parallel sentence pair through the source-side encoder; A32, decoding the encoding result through the target-side decoder to predict the target-side sentence of the current source-side sentence; A33, calculating the loss Loss_tgt of the target-side decoder predicting the target-side sentence according to the predicted target-side sentence and the actual target-side sentence in the second parallel sentence pair; A34, determining the translation loss gradient according to the loss of the target-side decoder predicting the target-side sentence, and updating the model parameters of the neural machine translation model according to the translation loss gradient to complete one iteration of the model parameters of the neural machine translation model. By repeating steps A31 to A34, continuous iterative training of the neural machine translation model is realized until the iterative training of the neural machine translation model is ended when the second preset end condition is satisfied, and a target chapter-level translation model is obtained.

[0121] It should be noted that the context prediction model can be a skip - thought model based on the Transformer framework, and the neural machine translation model can be a translation model based on the Transformer framework. The neural machine translation model has been explained above and will not be elaborated here.

[0122] To further improve the overall translation effect of the target passage - level translation model, an improvement is made on the basis of Training Method 1 to obtain a training method of pre - training + joint training. Specifically:

[0123] Before executing step A1, use the passage - level monolingual corpus to train the pre - training model to obtain the pre - trained pre - training model. Among them, the pre - training model is used to learn the sentence representation containing context semantic information from the sentences at the source end of the passage - level monolingual corpus; initialize the neural machine translation model and the context prediction model based on the pre - trained pre - training model.

[0124] The model structure of the pre - training model and the process of training the pre - training model are the same as or similar to those in Training Method 2. You can refer to the relevant descriptions in Training Method 2 below. For the sake of simplicity of the specification, it will not be elaborated here.

[0125] Training Method 2: Pre - training + Fine - tuning Method

[0126] Since the joint training method requires a large amount of passage - level parallel corpus as training data. However, in practice, it is not easy to obtain passage - level parallel corpus with document boundaries. On the contrary, it is easy to obtain a large amount of passage - level monolingual corpus as training data. In the case where passage - level parallel corpus cannot be obtained, Training Method 2 can be adopted: in terms of the passage - level training corpus being a passage - level monolingual corpus, the pre - training and fine - tuning methods can be used to obtain the target passage - level translation model. The specific process of Training Method 2 includes the following steps B1 - B2:

[0127] Step B1: Pre - train the pre - training model based on the passage - level monolingual corpus; Step B2: Fine - tune the target combined model based on the pre - trained pre - training model to obtain the target passage translation model corresponding to the target combined model.

[0128] Specifically, in Step B1, use the passage - level monolingual corpus to pre - train the pre - training model, and use the pre - training model to predict the context sentences of the current source - end sentence, so as to obtain the source - end context encoder that can capture the dependencies between sentences through the pre - training of the pre - training model.

[0129] Specifically, the passage - level monolingual corpus contains a certain number of monolingual training sentence pairs, and the monolingual training sentence pairs can be expressed as (s i , s i-1 , s i+1 ), where si is the current source-side sentence, s i-1 is s i the previous source-side sentence of s i+1 is s i the previous source-side sentence of s

[0130] The pre-training process of the pre-trained model includes a multi-iteration training process of the pre-trained model. Among them, as shown in Figure 6 shown Figure 6 This is a schematic flowchart of the pre-training + fine-tuning method in the embodiments of the present invention. Any iteration training process of the pre-trained model includes the following steps B11 to B13:

[0131] Step B11: For the current source-side sentence of the monolingual training sentence pair in the passage-level monolingual corpus, predict the source-side context sentence of the current source-side sentence through the pre-trained model.

[0132] Specifically, the model structure of the pre-trained model can use any one of the following two model structures:

[0133] ① It consists of two source-side decoders (a front decoder and a rear decoder) and a source-side encoder shared by the front decoder and the rear decoder. The pre-trained model encodes the current source-side sentence s i based on the same source-side encoder. The two source-side decoders of the pre-trained model decode respectively according to the encoding results of the same source-side encoder to obtain the previous source-side sentence s i-1 and the next source-side sentence s i+1 .

[0134] ② The pre-trained model consists of Figure 3 the two encoder-decoder models shown, including: a front decoder, a rear decoder, a front encoder, and a rear encoder). After representing the current source-side sentence as a word vector, it is input into the two encoder-decoder models. Two independent source-side encoders (front encoder, rear encoder) in the pre-trained model are used to encode the current source-side sentence respectively. The front decoder is used to predict the previous source-side sentence s i of the current source-side sentence s i-1 , and the rear decoder is used to predict the next source-side sentence s i of the current source-side sentence s i+1 . It should be noted that, as shown in Figure 3 the two encoder-decoder models in the pre-trained model share the word vector of the source-side sentence, that is: after representing the current source-side sentence as a word vector, it is input into the two encoder-decoder models.

[0135] Step B12: Determine the context loss function jointly based on the source context sentence predicted by the pre-trained model and the actual source context sentence in the monolingual training sentence pair, and determine the gradient of the context loss function for this iteration according to the context loss function; Step B13: Adjust the model parameters of the pre-trained model according to the gradient of the context loss function for this iteration.

[0136] It should be noted that one or more monolingual training sentence pairs can be used in one iteration of the pre-trained model.

[0137] By repeating the above steps B11 - B13, the pre-trained model is continuously iteratively trained until convergence or the maximum number of iterations is reached, and the pre-trained pre-trained model is obtained. Among them, the calculation formula of the context loss function is as follows:

[0138] Loss = Loss_pre + Loss_next

[0139] Among them, Loss_pre is the loss of predicting the previous source sentence by an encoder-decoder model in the pre-trained model, and Loss_next is the loss of predicting the next source sentence by another encoder-decoder model in the pre-trained model.

[0140] Step B2 is specifically to initialize the target combined model according to the pre-trained pre-trained model. Specifically, referring to Figure 5 As shown, the target combined model is obtained by integrating a neural machine translation model and two source encoders (a pre-encoder and a next-encoder) in the pre-trained pre-trained model. Among them, the neural machine translation model is a translation model based on the Transformer framework, which has been explained above and will not be elaborated here.

[0141] Step B3: Fine-tune the initialized target combined model using the document-level monolingual corpus to obtain the target document-level translation model corresponding to the target combined model.

[0142] Specifically, referring to Figure 5 As shown, the sum of the output of the pre-encoder, the output of the next-encoder, and the word vector of the current source sentence in the target combined model is used as the input of the source encoder in the neural machine translation model to fine-tune the target combined model. During the fine-tuning process, the parameters of the word vector, the pre-encoder, and the next-encoder in the target combined model will be continuously optimized.

[0143] Through the above technical solution, since the target passage translation model captures the context semantic information of the sentence to be translated in the passage, it can eliminate translation ambiguities and keep the translation results of the same words more consistent in different sentences, thereby improving the translation effect for passage-level texts.

[0144] Next, experiments were conducted on the target passage translation models trained for various embodiments provided in the embodiments of the present invention in two translation tasks of Chinese-English and English-German to verify the effectiveness of the target passage translation model for passage-level texts, but it is not a limitation to the present invention:

[0145] In the Chinese-English translation task, LDC (Linguistic Data Consortium) corpus was used, which includes LDC2003E14, LDC2005T06, LDC2005T10, and a part of LDC2004T08 (conference records / law / news) corpus. The scale of these corpora is 2.8M parallel sentence pairs. 94K passages (a total of 900K sentence pairs) were selected from the 2.8M parallel sentence pair corpus. The NIST06 dataset in the NIS database was used as the development set, and NIST02 / NIST03 / NIST04 / NIST05 / NIST08 were used as the test sets. The development set and the test set together contain 588 passages and 5833 sentence pairs. Each document contains an average of 10 sentences. Passage-level monolingual corpora were collected for the experiment. The total number of passage-level monolingual corpora collected is 25M sentences and 700K documents, with an average of 35 sentences per document.

[0146] In the English-German translation task, the passage-level bilingual data of WMT19 was used as the training set (a total of 39K documents and 855K sentence pairs). In addition, passage-level monolingual corpora of 410K documents (including 10M sentences) were collected. newstest2019 was used as the development set, and newstest2017 and newstest2018 were used as the test sets. The development set contains 123 documents and 2998 sentences, and the test set contains 255 documents and 6002 sentences.

[0147] For the Chinese side, word segmentation was performed, and then byte pair encoding (BPE) was used to segment words into smaller-grained sub-words. On the English and German sides, byte pair encoding (BPE) was used to segment words into smaller-grained sub-words. The case-insensitive NISTBLEU (an evaluation metric for machine translation improved based on BLEU) score was used as the evaluation scale, and the "mteval-v11b.pl" script was used to calculate the BLEU score. Words out of the vocabulary were all marked with "UNK" as a replacement.

[0148] A neural machine translation model using the Transformer architecture is used as the baseline model. During training: the size of the hidden layer is set to 512, and the filter size is set to 2048. The number of layers of both the decoder and the encoder is set to 6 layers, and the number of attention heads is set to 8. The Adam algorithm is used to update the target passage translation model. The learning rate is set to 1.0, and the number of learning rate update steps (warm-steps) is set to 4000. At each iteration, 4096 words are used for one batch process. 4 TITAN XP GPUs are used for encoding during the encoding process, and 2 TITAN XP GPUs are used for decoding during the decoding process; the width of beam-search is set to 4 during the decoding process. Significance detection is performed on the translation results of the test set.

[0149] The experimental results of the joint training method are shown in Table 1. As shown in Table 1, for the joint training method, "Pre" / "Next" means only the previous encoder or the next encoder is used. According to the experimental results on the development set, μ = 0.5 and λ = 0.5 are set for the previous encoder or the next encoder respectively (on the development set, μ = 0.1, 0.5, 1.0 are verified respectively, and μ is set to 0.5 with the best effect). The BLEU scores of only using the previous encoder or the next encoder are similar, indicating that the influence of the previous and the next source sentences on the translation of the current source sentence is almost the same. When using the "pre+next" method to predict the previous source sentence and the next source sentence simultaneously (μ = 0.5, λ = 0.3 are set according to the results of the development set), it is better than the method of only using "pre" and "next", and it improves by +0.84 BLEU points compared with the baseline model. Indicates that the significance detection is better than the baseline model (P < 0.01)

[0150] Table 1 Comparison of BLEU scores for Chinese-English translation

[0151]

[0152] A pre-trained model was trained on a monolingual corpus of passages on the order of 25M, and then fine-tuned on two parallel corpora of different sizes: a 900K parallel corpus and a 2.8M parallel corpus. The sentences in the 900K corpus showed strong contextual relevance. However, in the 2.8M corpus, not all passages had clear document boundaries. When training the pre-trained model with monolingual passages, the sentences in the document were not randomly shuffled, and the original order of the sentences in each document was maintained. The results are shown in Tables 1 and 2. Similar to joint training, the pre-trained model was trained with a single encoder, which could be either a pre-encoder or a next-encoder, to obtain "pre" or "next" results. Of course, both encoders could be trained simultaneously, corresponding to the "Pre+Next" result. As shown in Tables 1 and 2, the "Pre" and "Next" of the pre-training and fine-tuning methods achieved significant improvements on the two baselines. Without using any context information, the improvements over the Transformer baseline model were +0.93 and +1.28 BLEU points respectively. In Table 2, Indicates that the significance detection is better than the baseline model (P<0.01)

[0153] Table 2 Comparison of BLEU scores of the pre-training + fine-tuning method in Chinese-English translation.

[0154]

[0155] It can be seen from the above experimental data that the target passage models obtained by both joint training and pre-training + fine-tuning methods can significantly improve the translation effect. Further, before training the joint model, the joint model was initialized through the pre-trained model. As can be seen from Table 1, the highest BLEU score was obtained, which was +1.14 higher than the Transformer baseline.

[0156] As shown in Table 3, it is superior to the baseline model by 0.81 and 0.92 BLEU points in English-German translation respectively.

[0157] Table 3 Comparison of BLEU scores in English-German translation

[0158]

[0159] It can be seen from the above experimental data that by predicting the source context sentences from the current source sentence in the passage through the target passage translation model, the translation performance of the model can be improved.

[0160] Experiments have shown that using two independent encoders to encode the current source sentence in the pre-trained model, which are used to predict the previous source sentence and the next source sentence respectively, can achieve better translation results. The BLEU data obtained from the experiments are shown in Table 4 below.

[0161] Table 4 Comparison between two encoders and a single encoder in the pre-trained model

[0162]

[0163] The encoding of the current source sentence using two independent encoders in the pre-trained model is compared with the encoding using a shared encoder. The results are shown in Table 4. Obviously, the non-shared encoder model is better than the shared encoder model. This indicates that two independent encoders are better at capturing the dependencies between the current sentence and the surrounding sentences. This is because the dependencies between the current sentence and the previous and next sentences are different.

[0164] As can be seen from the results shown in Table 5, if only the pre-trained word vectors are used as input, improvements of +0.34 and +0.66 BLEU points can be obtained on two parallel corpora of different scales. When the sum of the input word vectors, the output of the pre-encoder and the output of the next-encoder in the pre-trained model is used as the input of the encoder in the neural machine translation model, compared with using only word vectors as input, the BLEU is improved by +0.5 and +0.62.

[0165] Table 5 BLEU scores of the target combination model on two parallel corpora of different scales.

[0166]

[0167] In Table 5: "pre+next+word vectors" means using the outputs of the pre-encoder and the next-encoder, as well as the word vectors of the pre-trained model, as the input of the target combination model together. "Word vectors" means using only the word vectors of the pre-trained model as input.

[0168] The translation method based on the target text translation model provided by the embodiment of the present invention can improve the translation quality by eliminating translation ambiguity (such as Example 1 in Table 6) or making the translation effect more consistent (such as Example 2 in Table 6). In Example 1, "fragile" has two meanings, "weak" or "fragile". Since the target text translation model knows the exact contextual semantic information, it can correctly translate the word "fragile". In Example 2, the target text translation model translates "discovery" in the two sentences into the same translation "detected" because the "case" mentioned in the second sentence can be predicted from the meaning of "drugs" and "police" in the first sentence. In addition, 5 documents with a total of 48 sentences were randomly selected from the test set for translation analysis.

[0169] Table 6 Comparison of translation results between the baseline model and the target paragraph translation model

[0170]

[0171] In the second aspect, based on the same inventive concept, an embodiment of the present invention provides a translation model training method, comprising the following steps: jointly training a context prediction model and a neural machine translation model based on the paragraph-level parallel corpus to obtain a target paragraph translation model corresponding to the neural machine translation model, wherein the context prediction model is used to learn sentence representations containing contextual semantic information from source sentences of the paragraph-level parallel corpus.

[0172] In a specific implementation, the context prediction model and the neural machine translation model are jointly trained based on the paragraph-level parallel corpus to obtain a target paragraph translation model corresponding to the neural machine translation model, including: using the first paragraph-level parallel corpus obtained to jointly train the neural machine translation model and the context prediction model until a trained joint model is obtained, wherein the first parallel sentence pair in the first paragraph-level parallel corpus includes a current source sentence and a source context sentence and a target sentence for the current source sentence; extracting a trained neural machine translation model from the trained joint model;

[0173] In a specific implementation, the trained neural machine translation model is further trained using the acquired second chapter-level parallel corpus until a target chapter-level translation model corresponding to the neural machine translation model is obtained, wherein the second parallel sentence pair in the second chapter-level parallel corpus includes a current source sentence and a target sentence for the current source sentence.

[0174] In a specific embodiment, before jointly training the neural machine translation model and the context prediction model by using the obtained first chapter-level parallel corpus, the method further includes: training a pre-trained model by using a chapter-level monolingual corpus to obtain a pre-trained pre-trained model, where the pre-trained model is used to learn sentence representations containing context semantic information from the sentences at the source end of the chapter-level monolingual corpus; initializing the neural machine translation model and the context prediction model based on the pre-trained pre-trained model.

[0175] In a third aspect, based on the same inventive concept, an embodiment of the present invention provides a translation model training method, including the following steps: pre-training a pre-trained model based on a chapter-level monolingual corpus; fine-tuning a target combined model according to the pre-trained pre-trained model to obtain a target chapter translation model corresponding to the target combined model, where the pre-trained model is used to learn sentence representations containing context semantic information from the source-end sentences of the chapter-level monolingual corpus, and the target combined model includes a neural machine translation model and a source-end context encoder.

[0176] In a fourth aspect, based on the same inventive concept, an embodiment of the present invention provides a chapter translation device, as shown in Figure 7 including:

[0177] A sentence translation unit 701, configured to, for each sentence to be translated in a to-be-translated chapter, obtain a sentence representation containing context semantic information of the sentence to be translated through a target chapter translation model, and translate the sentence to be translated based on the sentence representation;

[0178] A translation result forming unit 702, configured to obtain a chapter translation result corresponding to the to-be-translated chapter according to the translation results of each sentence to be translated in the to-be-translated chapter.

[0179] In a specific embodiment, the device further includes:

[0180] A model training unit, configured to obtain the target chapter translation model by learning sentence representations containing context semantic information in a chapter-level training corpus, where the chapter-level training corpus is a chapter-level parallel corpus and / or a chapter-level monolingual corpus.

[0181] In a specific embodiment, the model training unit includes:

[0182] The first training unit is used to jointly train a context prediction model and a neural machine translation model for the passage-level parallel corpus based on the passage-level training corpus as the passage-level parallel corpus, so as to obtain a target passage translation model corresponding to the neural machine translation model, wherein the context prediction model is used to learn sentence representations containing context semantic information from the source-side sentences of the passage-level parallel corpus;

[0183] Or the model training unit includes:

[0184] The second training unit is used to pre-train a pre-trained model based on the passage-level training corpus as the passage-level monolingual corpus;

[0185] The model fine-tuning unit is used to fine-tune the target combination model according to the pre-trained pre-trained model to obtain a target passage translation model corresponding to the target combination model, wherein the pre-trained model is used to learn sentence representations containing context semantic information from the source-side sentences of the passage-level monolingual corpus, and the target combination model includes a neural machine translation model and a source-side context encoder.

[0186] In a specific embodiment, the first training unit includes:

[0187] The joint training subunit is used to jointly train the neural machine translation model and the context prediction model by using the obtained first passage-level parallel corpus until a trained joint model is obtained. The first parallel sentence pairs in the first passage-level parallel corpus include the current source-side sentence, the source-side context sentence, and the target-side sentence for the current source-side sentence;

[0188] The model extraction subunit is used to extract the trained neural machine translation model from the trained joint model;

[0189] The continuous training subunit is used to continue training the trained neural machine translation model by using the obtained second passage-level parallel corpus until a target passage-level translation model corresponding to the neural machine translation model is obtained, wherein the second parallel sentence pairs in the second passage-level parallel corpus include the current source-side sentence and the target-side sentence for the current source-side sentence.

[0190] In a specific embodiment, the neural machine translation model and the context prediction model share the same source-side encoder, the neural machine translation model further includes a target-side decoder, and the context prediction model further includes a source-side context decoder.

[0191] In a specific embodiment, the joint training subunit is used to perform multiple joint iterative trainings on the neural machine translation model and the context prediction model; in any joint iterative training, the joint training subunit is specifically used for:

[0192] Encoding the current source sentence of the first parallel sentence pair in the first passage-level parallel corpus through the source encoder;

[0193] Respectively decoding the encoding result of the same source encoder through the target decoder and the source context decoder to predict the target sentence and the source context sentence corresponding to the current source sentence;

[0194] Joint loss gradients are jointly determined based on the predicted source context sentence, the predicted target prediction sentence, and the first parallel sentence pair;

[0195] Based on the joint loss gradients, update the model parameters of the neural machine translation model and the model parameters of the context prediction model.

[0196] In a specific embodiment, the device further includes:

[0197] A first pre-training unit, configured to train a pre-training model using the passage-level monolingual corpus to obtain a pre-trained pre-training model, where the pre-training model is used to learn sentence representations containing context semantic information from the sentences at the source end of the passage-level monolingual corpus;

[0198] An initialization unit, configured to initialize the neural machine translation model and the context prediction model based on the pre-trained pre-training model.

[0199] In a specific embodiment, the model fine-tuning unit includes:

[0200] An initialization subunit, configured to initialize the target combined model based on the pre-trained pre-training model;

[0201] A fine-tuning training unit, configured to train the target combined model according to the passage-level monolingual corpus to obtain a target passage translation model corresponding to the target combined model.

[0202] In a fifth aspect, based on the same inventive concept, an embodiment of the present invention provides a translation model training device, including:

[0203] The first training unit is used to jointly train a context prediction model and a neural machine translation model based on the passage-level parallel corpus to obtain a target passage translation model corresponding to the neural machine translation model, where the context prediction model is used to learn sentence representations containing context semantic information from the source-side sentences of the passage-level parallel corpus.

[0204] In a specific embodiment, the first training unit includes:

[0205] The joint training subunit is used to jointly train the neural machine translation model and the context prediction model using the obtained first passage-level parallel corpus until a trained joint model is obtained. The first parallel sentence pairs in the first passage-level parallel corpus include the current source-side sentence, the source-side context sentence, and the target-side sentence for the current source-side sentence.

[0206] The model extraction subunit is used to extract the trained neural machine translation model from the trained joint model.

[0207] The continued training subunit is used to continue training the trained neural machine translation model using the obtained second passage-level parallel corpus until a target passage-level translation model corresponding to the neural machine translation model is obtained, where the second parallel sentence pairs in the second passage-level parallel corpus include the current source-side sentence and the target-side sentence for the current source-side sentence.

[0208] In a specific embodiment, the device further includes:

[0209] The first pre-training unit is used to train a pre-training model using the passage-level monolingual corpus to obtain a pre-trained pre-training model, where the pre-training model is used to learn sentence representations containing context semantic information from the sentences on the source side of the passage-level monolingual corpus.

[0210] The initialization unit is used to initialize the neural machine translation model and the context prediction model based on the pre-trained pre-training model.

[0211] In a sixth aspect, based on the same inventive concept, an embodiment of the present invention provides a translation model training device, including:

[0212] The initialization subunit is used to initialize the target combined model based on the pre-trained pre-training model;

[0213] The fine-tuning training unit is used to train the target combined model according to the passage-level monolingual corpus to obtain a target passage translation model corresponding to the target combined model.

[0214] Regarding the above device, the specific functions of each unit have been described in detail in the method embodiments of passage translation provided in one aspect of the present invention, and will not be elaborated here. The specific implementation process can refer to the method embodiments of passage translation provided in the first aspect above.

[0215] In a seventh aspect, based on the same inventive concept as the foregoing method embodiments of passage translation, embodiments of this specification further provide an electronic device. Based on the same inventive concept as the foregoing method embodiments of passage translation, embodiments of this specification further provide an electronic device. Figure 8 FIG. is a block diagram of an electronic device 800 shown according to an exemplary embodiment. For example, the device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0216] Referring to Figure 8 , the device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0217] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing element 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing unit 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0218] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0219] The power component 806 provides power to various components of the device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.

[0220] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0221] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0222] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0223] The sensor component 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the device 800. The sensor component 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0224] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0225] In an exemplary embodiment, the device 800 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0226] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0227] A non-transitory computer-readable storage medium is also provided. When the instructions in the storage medium are executed by a processor of a mobile terminal, the device 800 can execute a translation model training method or a passage translation method, and the method includes any one of the implementation manners in the foregoing translation model training method embodiments, and the method includes any one of the implementation manners in the foregoing passage translation method embodiments.

[0228] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0229] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims. The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for chapter translation, characterized in that, including: For each sentence to be translated in the passage to be translated, obtaining a sentence representation of the sentence to be translated containing context semantic information through a target passage translation model, and translating the sentence to be translated based on the sentence representation; Obtaining a passage translation result corresponding to the passage to be translated according to the translation results of each sentence to be translated in the passage to be translated; Obtaining the target passage translation model by learning sentence representations containing context semantic information in a passage-level training corpus, including: for the passage-level training corpus being a passage-level parallel corpus, jointly training a neural machine translation model and a context prediction model using the obtained first passage-level parallel corpus to obtain a trained joint model, where the first parallel sentence pair in the first passage-level parallel corpus includes a current source-side sentence, a source-side context sentence, and a target-side sentence for the current source-side sentence; extracting the trained neural machine translation model from the trained joint model; training the trained neural machine translation model using the obtained second passage-level parallel corpus to obtain a target passage-level translation model corresponding to the neural machine translation model, where the second parallel sentence pair in the second passage-level parallel corpus includes a current source-side sentence and a target-side sentence for the current source-side sentence; Wherein, the neural machine translation model and the context prediction model share the same source-side encoder, the neural machine translation model further includes a target-side decoder, and the context prediction model further includes a source-side context decoder; any joint iterative training of the neural machine translation model and the context prediction model includes: encoding the current source-side sentence of the first parallel sentence pair in the first passage-level parallel corpus through the source-side encoder; decoding the encoding result of the same source-side encoder through the target-side decoder and the source-side context decoder respectively to predict the target-side sentence and the source-side context sentence corresponding to the current source-side sentence; jointly determining a joint loss gradient according to the predicted source-side context sentence, the target-side prediction sentence, and the first parallel sentence pair; updating the model parameters of the neural machine translation model and the model parameters of the context prediction model based on the joint loss gradient.

2. The method according to claim 1, characterized in that The obtaining of the target passage translation model by learning sentence representations containing context semantic information in a passage-level training corpus further includes: For the passage-level training corpus being a passage-level monolingual corpus, pre-training a pre-trained model based on the passage-level monolingual corpus, and fine-tuning a target combination model according to the pre-trained pre-trained model to obtain a target passage translation model corresponding to the target combination model, where the pre-trained model is used to learn sentence representations containing context semantic information from the source-side sentences of the passage-level monolingual corpus, and the target combination model includes a neural machine translation model and a source-side context encoder.

3. A method for training a translation model, characterized in that, including: Training a context prediction model and a neural machine translation model jointly based on a passage-level parallel corpus to obtain a target passage translation model corresponding to the neural machine translation model, including: jointly training the neural machine translation model and the context prediction model by using the obtained first passage-level parallel corpus to obtain a trained joint model, where the first parallel sentence pairs in the first passage-level parallel corpus include a current source-side sentence, a source-side context sentence, and a target-side sentence for the current source-side sentence; extracting the trained neural machine translation model from the trained joint model; training the trained neural machine translation model by using the obtained second passage-level parallel corpus to obtain a target passage-level translation model corresponding to the neural machine translation model, where the second parallel sentence pairs in the second passage-level parallel corpus include a current source-side sentence and a target-side sentence for the current source-side sentence; Wherein, the neural machine translation model and the context prediction model share the same source-side encoder, the neural machine translation model further includes a target-side decoder, and the context prediction model further includes a source-side context decoder; any joint iterative training of the neural machine translation model and the context prediction model includes: encoding the current source-side sentence of the first parallel sentence pair in the first passage-level parallel corpus through the source-side encoder; decoding the encoding result of the same source-side encoder through the target-side decoder and the source-side context decoder respectively to predict the target-side sentence and the source-side context sentence corresponding to the current source-side sentence; jointly determining a joint loss gradient according to the predicted source-side context sentence, the target-side predicted sentence, and the first parallel sentence pair; updating the model parameters of the neural machine translation model and the model parameters of the context prediction model based on the joint loss gradient.

4. A passage translation device, characterized in that, Including: A sentence translation unit, configured to, for each sentence to be translated in a passage to be translated, obtain a sentence representation including context semantic information of the sentence to be translated through the target passage translation model, and translate the sentence to be translated based on the sentence representation; A model training unit, which is used to obtain the target passage translation model by learning sentence representations containing context semantic information in passage-level training corpora, includes: for the passage-level training corpus being a passage-level parallel corpus, jointly training a neural machine translation model and a context prediction model using the obtained first passage-level parallel corpus to obtain a trained joint model, where the first parallel sentence pairs in the first passage-level parallel corpus include the current source-side sentence, the source-side context sentence for the current source-side sentence, and the target-side sentence; extracting the trained neural machine translation model from the trained joint model; training the trained neural machine translation model using the obtained second passage-level parallel corpus to obtain a target passage-level translation model corresponding to the neural machine translation model, where the second parallel sentence pairs in the second passage-level parallel corpus include the current source-side sentence and the target-side sentence for the current source-side sentence; wherein, the neural machine translation model and the context prediction model share the same source-side encoder, the neural machine translation model further includes a target-side decoder, and the context prediction model further includes a source-side context decoder; any joint iterative training of the neural machine translation model and the context prediction model includes: encoding the current source-side sentence in the first parallel sentence pair in the first passage-level parallel corpus through the source-side encoder; decoding the encoding result of the same source-side encoder through the target-side decoder and the source-side context decoder respectively to predict the target-side sentence and the source-side context sentence corresponding to the current source-side sentence; jointly determining a joint loss gradient based on the predicted source-side context sentence, the target-side predicted sentence, and the first parallel sentence pair; updating the model parameters of the neural machine translation model and the model parameters of the context prediction model based on the joint loss gradient; A translation result forming unit, which is used to obtain the passage translation result corresponding to the passage to be translated according to the translation results of each sentence to be translated in the passage to be translated.

5. A translation model training device, characterized in that, including: A first training unit, which is used to jointly train a neural machine translation model and a context prediction model using the obtained first passage-level parallel corpus for the passage-level training corpus being a passage-level parallel corpus to obtain a trained joint model, where the first parallel sentence pairs in the first passage-level parallel corpus include the current source-side sentence, the source-side context sentence for the current source-side sentence, and the target-side sentence; extracting the trained neural machine translation model from the trained joint model; training the trained neural machine translation model using the obtained second passage-level parallel corpus to obtain a target passage-level translation model corresponding to the neural machine translation model, where the second parallel sentence pairs in the second passage-level parallel corpus include the current source-side sentence and the target-side sentence for the current source-side sentence; Among them, the neural machine translation model and the context prediction model share the same source encoder. The neural machine translation model further includes a target decoder, and the context prediction model further includes a source context decoder. Any joint iterative training of the neural machine translation model and the context prediction model includes: encoding the current source sentence of the first parallel sentence pair in the first passage-level parallel corpus through the source encoder; respectively decoding the encoding result of the same source encoder through the target decoder and the source context decoder to predict the target sentence and the source context sentence corresponding to the current source sentence; jointly determining a joint loss gradient according to the predicted source context sentence, the target prediction sentence, and the first parallel sentence pair; and updating the model parameters of the neural machine translation model and the model parameters of the context prediction model based on the joint loss gradient.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1 to 3.

7. An electronic device, characterized in that, It includes a memory and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors with operation instructions for performing the method according to any one of claims 1 to 3 included in the one or more programs.

Citation Information

Patent Citations

  • Text-level neural machine translation method based on context memory network

    CN111160050A

  • Encoder-decoder framework pre-training method for neural machine translation

    CN111382580A