Neural machine translation method and model training method, device and electronic equipment
By obtaining the highest similar sentence pairs from the translation memory and training the neural machine translation model, the problems of incorrect translation results and poor translation quality in the existing technology are solved, and higher translation accuracy and translation quality are achieved.
Patent Information
- Application Number
- CN202010646609.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-07-07
AI Technical Summary
Existing neural machine translation methods often have errors in translation results, and the translation quality is poor, making it difficult to effectively use sentence pair information in translation memory to improve translation quality.
By obtaining the translation memory sentence pairs with the highest similarity to each source-end sentence in the parallel corpus from the translation memory, and using them as a training sample set, the preset initial model is trained until the parameters of the source-end encoder layer, target-end encoder layer and decoder layer of the model's translation memory converge, a neural machine translation model is obtained.
By using similar sentences in the translation memory to match information, the accuracy of neural machine translation results is improved and the quality of translation is improved.
Smart Images

Figure CN113919373B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of machine translation, and more specifically, to a training method for a neural machine translation model, a neural machine translation method, a training device for a neural machine translation model, a neural machine translation device, and an electronic device. Background Art
[0002] Neural Machine Translation (NMT) is a method based on neural network technology to achieve contextual accurate translation. Common neural machine translation methods are usually implemented based on encoders and decoders. The encoder encodes the source language sentence into a set of hidden state representations, and the decoder obtains the hidden state representations generated by the encoder through the attention network, thereby generating the target sentence word by word. However, due to the performance limitations of existing neural machine translation, the translation results often contain many errors and the translation quality is poor. Therefore, it is necessary to propose a new neural machine translation method. Summary of the invention
[0003] One purpose of the embodiments of the present disclosure is to provide a new technical solution for training a neural machine translation model.
[0004] According to a first aspect of the present disclosure, a method for training a neural machine translation model is provided, the method comprising:
[0005] Obtaining from the translation memory a translation memory sentence pair with the highest similarity to each source sentence in the parallel corpus;
[0006] Using source language sentence pairs and translation memory sentence pairs as training sample sets to train a preset initial model until a first parameter of a translation memory source-end encoder layer of the preset initial model converges, a second parameter of a translation memory target-end encoder layer of the preset initial model converges, and a third parameter of a decoder layer containing translation memory information of the preset initial model converges, thereby obtaining a neural machine translation model;
[0007] The source-end sentence pair includes a source-end sentence and a corresponding target-end sentence, and the translation memory sentence pair includes a translation memory source-end sentence and a corresponding translation memory target-end sentence.
[0008] Optionally, the step of obtaining from the translation memory a translation memory sentence pair having the highest similarity to each source sentence in the parallel corpus comprises:
[0009] Calculating the similarity between the source sentence and each translation memory source sentence in the translation memory;
[0010] The similarities are sorted from large to small, and the translation memory source-end sentence with the greatest similarity and its corresponding target-end sentence are determined as the translation memory sentence pair with the highest similarity to the source-end sentence.
[0011] Optionally, the calculating the similarity between the source sentence and each translation memory source sentence in the translation memory includes:
[0012] The edit distance between the source sentence and each translation memory source sentence in the translation memory is calculated as the similarity.
[0013] Optionally, the step of using the source language sentence pairs and the translation memory sentence pairs as training sample sets to train a preset initial model until a first parameter of a translation memory source-end encoder layer of the preset initial model converges, a second parameter of a translation memory target-end encoder layer of the preset initial model converges, and a third parameter of a decoder layer containing translation memory information of the preset initial model converges, comprises:
[0014] Translate each sample in the training sample set based on the preset initial model to obtain a translation result;
[0015] Substituting the translation result into a preset loss function for calculation to obtain the loss of each sample;
[0016] The first parameter, the second parameter, and the third parameter are updated based on the loss until the first parameter, the second parameter, and the third parameter converge.
[0017] Optionally, the updating the first parameter, the second parameter, and the third parameter based on the loss until the first parameter, the second parameter, and the third parameter converge comprises:
[0018] Based on the loss and a preset back propagation algorithm, respectively calculating a first derivative of the first parameter, a second derivative of the second parameter, and a third derivative of the third parameter;
[0019] updating the first parameter based on the first derivative and a gradient descent algorithm, updating the second parameter based on the second derivative and a gradient descent algorithm, and updating the third parameter based on the third derivative and a gradient descent algorithm;
[0020] The first parameter, the second parameter and the third parameter are updated multiple times based on the losses of multiple samples in the training sample set until convergence, thereby obtaining the neural machine translation model.
[0021] According to a second aspect of the present disclosure, a neural machine translation method is provided, the method comprising:
[0022] Get the source sentence to be translated;
[0023] The source sentence to be translated is input into a neural machine translation model, and a target sentence is output as a translation result; wherein the neural machine translation model is trained according to the training method of a neural machine translation model as described in any one of the first aspects of the present disclosure.
[0024] According to a third aspect of the present disclosure, a training device for a neural machine translation model is provided, the device comprising:
[0025] An acquisition module is used to acquire, from the translation memory, the translation memory sentence pair with the highest similarity to each source sentence in the parallel corpus;
[0026] A training module is used to train a preset initial model using source language sentence pairs and translation memory sentence pairs as training sample sets until a first parameter of a translation memory source-end encoder layer of the preset initial model converges, a second parameter of a translation memory target-end encoder layer of the preset initial model converges, and a third parameter of a decoder layer containing translation memory information of the preset initial model converges, thereby obtaining a neural machine translation model; wherein the source-end sentence pair includes a source-end sentence and a corresponding target-end sentence, and the translation memory sentence pair includes a translation memory source-end sentence and a corresponding translation memory target-end sentence.
[0027] According to a fourth aspect of the present disclosure, a neural machine translation device is provided, the device comprising:
[0028] The acquisition module is used to obtain the source sentence to be translated;
[0029] A translation module is used to input the source sentence to be translated into a neural machine translation model and output a target sentence as a translation result; wherein the neural machine translation model is trained according to the training method of the neural machine translation model as described in any one of the first aspects of the present disclosure.
[0030] According to the fifth aspect of the embodiments of the present disclosure, there is also provided an electronic device, comprising a processor and a memory; the memory stores machine executable instructions that can be executed by the processor; the processor executes the machine executable instructions to implement the training method of the neural machine translation model described in any one of the first aspects of the embodiments of the present disclosure.
[0031] According to the sixth aspect of an embodiment of the present disclosure, there is also provided an electronic device, comprising a processor and a memory; the memory stores machine executable instructions that can be executed by the processor; the processor executes the machine executable instructions to implement the neural machine translation method described in the second aspect of an embodiment of the present disclosure.
[0032] According to an embodiment of the present disclosure, the accuracy of neural machine translation results can be improved and the quality of translation can be improved.
[0033] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0035] Figure 1 A schematic diagram of the composition structure of an electronic device to which the training method of a neural machine translation model according to an embodiment of the present disclosure can be applied;
[0036] Figure 2 is a flowchart of a training method for a neural machine translation model according to an embodiment of the present disclosure;
[0037] Figure 3 is a schematic diagram of the structure of an encoder according to an embodiment of the present disclosure;
[0038] Figure 4 is a schematic diagram of a decoder structure according to an embodiment of the present disclosure;
[0039] Figure 5 is a structural schematic diagram of a training device for a neural machine translation model according to an embodiment of the present disclosure;
[0040] Figure 6 A functional block diagram of an electronic device according to a first embodiment of the present disclosure;
[0041] Figure 7 is a flowchart of a neural machine translation method according to an embodiment of the present disclosure;
[0042] Figure 8 is a schematic diagram of the structure of a neural machine translation device according to an embodiment of the present disclosure;
[0043] Fig. 9 A functional block diagram of an electronic device according to a second embodiment of the present disclosure. DETAILED DESCRIPTION
[0044] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.
[0045] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0046] Technologies, methods and equipment known to persons of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods and equipment should be considered part of the specification.
[0047] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0048] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0049] <Hardware Configuration>
[0050] Current neural machine translation systems can generate sentences t' in other languages with the same meaning based on the input sentence s. However, due to the performance limitations of neural machine translation systems, the resulting translations often contain many errors.
[0051] The translation memory can store bilingual translation pairs (s m ,t m ), where the bilingual translation pairs can be manually translated or collected through other means. If there is a sentence in the translation memory that is very similar to sentence s and its translation We can use this sentence pair information to help the neural network translation system translate sentence s, but how to translate the sentence pairs in the translation memory Integrating it into the neural machine translation system to generate better quality translation for sentence s has become an urgent problem to be solved. Based on this, the present disclosure proposes a method for improving the translation quality of the neural machine translation system by using a translation memory, which can be used in a variety of neural machine translation structures, including but not limited to convolutional neural networks, recurrent neural networks, and self-attention networks.
[0052] Figure 1 The present invention is a schematic diagram of the composition structure of an electronic device to which the training method of the neural machine translation model according to an embodiment of the present disclosure can be applied.
[0053] like Figure 1 As shown, the electronic device 1000 of this embodiment may include a processor 1010, a memory 1020, an interface device 1030, a communication device 1040, a display device 1050, an input device 1060, a speaker 1070, a microphone 1080, and the like.
[0054] The processor 1010 may be a central processing unit (CPU), a microprocessor (MCU), etc. The memory 1020 may include, for example, a ROM (read-only memory), a RAM (random access memory), a non-volatile memory such as a hard disk, etc. The interface device 1030 may include, for example, a USB interface, a headphone interface, etc. The communication device 1040 may be capable of wired or wireless communication, for example. The display device 1050 may be, for example, a liquid crystal display, a touch display, etc. The input device 1060 may include, for example, a touch screen, a keyboard, etc.
[0055] The electronic device 1000 may output audio information through the speaker 1070. The electronic device 1000 may pick up voice information input by a user through the microphone 1080.
[0056] The electronic device 1000 may be a portable computer, a desktop computer or the like.
[0057] In this embodiment, the electronic device 1000 can obtain the translation memory sentence pairs with the highest similarity to each source sentence in the parallel corpus from the translation memory; use the source language sentence pairs and the translation memory sentence pairs as training sample sets to train a preset initial model until the first parameter of the translation memory source encoder layer of the preset initial model converges, the second parameter of the translation memory target encoder layer of the preset initial model converges, and the third parameter of the decoder layer containing translation memory information of the preset initial model converges, so as to obtain a neural machine translation model; wherein the source sentence pair includes a source sentence and a corresponding target sentence, and the translation memory sentence pair includes a translation memory source sentence and a corresponding translation memory target sentence.
[0058] In this embodiment, the memory 1020 of the electronic device 1000 is used to store instructions, which are used to control the processor 1010 to operate to support the training method of the neural machine translation model according to any embodiment of this specification.
[0059] It should be understood by those skilled in the art that although Figure 1 Multiple devices of the electronic device 1000 are shown in the figure, but the electronic device 1000 of the embodiment of this specification may only involve some of the devices, for example, only the processor 1010, the memory 1020, the display device 1050, the input device 1060, etc.
[0060] The technicians can design instructions according to the scheme disclosed in the present disclosure. How instructions control the processor to operate is well known in the art, so it will not be described in detail here.
[0061] <First Embodiment>
[0062] <Method>
[0063] This embodiment provides a method for training a neural machine translation model. The method can be implemented by an electronic device, for example, Figure 1 An electronic device 1000 is shown.
[0064] like Figure 2 As shown, the method includes the following steps 2100 and 2200:
[0065] Step 2100: Obtain from the translation memory a translation memory sentence pair having the highest similarity to each source sentence in the parallel corpus.
[0066] The parallel corpus refers to a corpus that searches and displays the source text and its target text. In this embodiment, (s, t) represents a parallel sentence pair. (s i ,t i ) is used to represent the i-th sentence pair in the parallel corpus, where s i represents the source sentence of the ith sentence pair in the parallel corpus, with t i represents the target sentence, t i Yes i The translation memory stores bilingual translation pairs (s m ,t m ), the bilingual translation sentence pair (s m ,t m ) can be manually translated or can be sentence pairs of mutual translation collected through other means. It is used to represent the i-th sentence pair in the translation memory. represents the source sentence of the ith sentence pair in the translation memory, represents the target sentence, yes Translation of .
[0067] In this step, when the electronic device 1000 obtains the translation memory sentence pair with the highest similarity, it can specifically calculate the similarity between the source sentence and each translation memory source sentence in the translation memory; and sort the similarities from large to small, and determine the translation memory source sentence with the greatest similarity and its corresponding target sentence as the translation memory sentence pair with the highest similarity to the source sentence.
[0068] For example, given a parallel sentence pair (s i ,t i ), we use translation memory (s m ,t m ) All calculated i and The similarity of iAfter calculation, e i Sort and select e i The highest pair As i ,t i ) to form a training sample set
[0069] For example, when calculating the similarity, s i and The edit distance is taken as the similarity.
[0070] Step 2200, using source language sentence pairs and translation memory sentence pairs as training sample sets to train a preset initial model until the first parameter of the translation memory source-end encoder layer of the preset initial model converges, the second parameter of the translation memory target-end encoder layer of the preset initial model converges, and the third parameter of the decoder layer containing translation memory information of the preset initial model converges, to obtain a neural machine translation model.
[0071] The source-end sentence pair includes the source-end sentence s i and the corresponding target sentence t i , the translation memory sentence pair includes the translation memory source sentence and the corresponding translation memory target sentence
[0072] In this embodiment, in order to Introduced into the neural machine translation model, based on the traditional encoder, two new encoder layers are added, one for and As input. Figure 3 As shown in Figure 2, when encoding, the retrieved sentences are first encoded using the self-attention network. Then the source sentence s, which is also encoded using the self-attention network, is i The source context representation s-context is different from the encoded Perform cross-attention encoding to obtain The source context represents sm-context.
[0073] Similarly, for the translation memory target sentence First use self-attention network encoding Then and The source context representation sm-context is encoded through the cross-attention network to obtain The target-side context is represented by tm-context.
[0074] At the same time, in this embodiment, based on the traditional 6-layer decoder, a new layer of decoder is added, such as Figure 4 As shown, in this embodiment, the new decoder layer is named as the decoder layer containing translation memory information, which has three inputs, namely, the output TM of the original 6-layer decoder, the source sentence s i The source context is represented by s-context, and the target sentence in the translation memory The target context is represented by tm-context. These three pieces of information are used together to predict the translation t i .
[0075] When training the preset initial model, we can first train the traditional neural machine translation model as a benchmark system. After training the benchmark system, use its parameters to initialize the parameters of the preset initial model of this embodiment, that is, the first parameter of the translation memory source encoder layer of the preset initial model, the second parameter of the translation memory target encoder layer of the preset initial model, and the third parameter of the decoder layer of the preset initial model containing translation memory information.
[0076] Specifically, during training, the electronic device 1000 can translate each sample in the training sample set based on the preset initial model to obtain a translation result; substitute the translation result into a preset loss function for calculation to obtain the loss of each sample; and update the first parameter, the second parameter, and the third parameter based on the loss until the first parameter, the second parameter, and the third parameter converge.
[0077] Among them, the step of updating the first parameter, the second parameter and the third parameter based on the loss until the first parameter, the second parameter and the third parameter converge can specifically be: based on the loss and a preset back propagation algorithm, respectively calculating the first derivative of the first parameter, the second derivative of the second parameter and the third derivative of the third parameter; updating the first parameter based on the first derivative and the gradient descent algorithm, updating the second parameter based on the second derivative and the gradient descent algorithm, and updating the third parameter based on the third derivative and the gradient descent algorithm; updating the first parameter, the second parameter and the third parameter multiple times based on the loss of multiple samples in the training sample set until convergence to obtain the neural machine translation model.
[0078] For example, the electronic device 1000 calculates the first derivative of the first parameter W1 based on the loss L and a preset back propagation algorithm. Calculate the second derivative of the second parameter W2 and the third derivative of the third parameter W3 Then based on the first derivative and the stochastic gradient descent algorithm for the first parameter Update based on the second derivative And the stochastic gradient descent algorithm for the second parameter Update and based on the third derivative And the stochastic gradient descent algorithm for the third parameter to update.
[0079] Finally, the electronic device 1000 can update the first parameter, the second parameter and the third parameter multiple times based on the losses of multiple samples in the training sample set until convergence to obtain the neural machine translation model.
[0080] The training method of the neural machine translation model of the present embodiment has been described above with reference to the accompanying drawings and examples. The method of the present embodiment obtains the translation memory sentence pairs with the highest similarity to each source sentence in the parallel corpus from the translation memory; uses the source language sentence pairs and the translation memory sentence pairs as training sample sets to train the preset initial model until the first parameter of the translation memory source encoder layer of the preset initial model converges, the second parameter of the translation memory target encoder layer of the preset initial model converges, and the third parameter of the decoder layer containing translation memory information of the preset initial model converges, thereby obtaining a neural machine translation model; wherein the source sentence pairs include source sentences and corresponding target sentences, and the translation memory sentence pairs include translation memory source sentences and corresponding translation memory target sentences. The method according to the present embodiment obtains the neural machine translation model by utilizing the source sentence pairs existing in the translation memory and the source sentence pairs. i Very similar translation memory pairs To help the neural network translation model translate the source sentence s i , which can improve the accuracy of neural machine translation results and improve the quality of translation.
[0081] <Device>
[0082] This embodiment provides a training device for a neural machine translation model, such as Figure 5 The training apparatus 5000 of the neural machine translation model is shown.
[0083] like Figure 5 As shown, the training device 5000 of the neural machine translation model may include an acquisition module 5100 and a training module 5200.
[0084] The acquisition module 5100 is used to acquire, from the translation memory, the translation memory sentence pairs with the highest similarity to each source sentence in the parallel corpus.
[0085] The training module 5200 is used to train a preset initial model using source language sentence pairs and translation memory sentence pairs as training sample sets until the first parameter of the translation memory source-end encoder layer of the preset initial model converges, the second parameter of the translation memory target-end encoder layer of the preset initial model converges, and the third parameter of the decoder layer containing translation memory information of the preset initial model converges, so as to obtain a neural machine translation model; wherein the source-end sentence pair includes a source-end sentence and a corresponding target-end sentence, and the translation memory sentence pair includes a translation memory source-end sentence and a corresponding translation memory target-end sentence.
[0086] Specifically, the acquisition module 5100 is used to calculate the similarity between the source sentence and each translation memory source sentence in the translation memory; sort the similarities from large to small, and determine the translation memory source sentence with the greatest similarity and its corresponding target sentence as the translation memory sentence pair with the highest similarity to the source sentence.
[0087] In an example, the acquisition module 5100 may calculate the edit distance between the source sentence and each translation memory source sentence in the translation memory as the similarity.
[0088] Specifically, the training module 5200 is used to translate each sample in the training sample set based on the preset initial model to obtain a translation result; substitute the translation result into a preset loss function for calculation to obtain the loss of each sample; and update the first parameter, the second parameter, and the third parameter based on the loss until the first parameter, the second parameter, and the third parameter converge.
[0089] Among them, the training module 5200 updates the first parameter, the second parameter and the third parameter based on the loss until the first parameter, the second parameter and the third parameter converge. Specifically, based on the loss and a preset back propagation algorithm, the first derivative of the first parameter, the second derivative of the second parameter and the third derivative of the third parameter can be calculated respectively; the first parameter is updated based on the first derivative and the gradient descent algorithm, the second parameter is updated based on the second derivative and the gradient descent algorithm, and the third parameter is updated based on the third derivative and the gradient descent algorithm; the first parameter, the second parameter and the third parameter are updated multiple times based on the losses of multiple samples in the training sample set until convergence to obtain the neural machine translation model.
[0090] The training device of the neural machine translation model of this embodiment can be used to execute the method and technical solution of this embodiment. Its implementation principle and technical effects are similar and will not be repeated here.
[0091] <Device>
[0092] In this embodiment, an electronic device is also provided, which includes the training device 5000 of the neural machine translation model described in the device embodiment of the present disclosure; or, the electronic device is Figure 6 The electronic device 6000 shown includes a processor 6200 and a memory 6100 .
[0093] The memory 6100 stores machine executable instructions that can be executed by the processor; the processor 6200 executes the machine executable instructions to implement the training method of the neural machine translation model as any one of the embodiments of the present invention.
[0094] <Computer Readable Storage Medium Embodiment>
[0095] This embodiment provides a computer-readable storage medium, in which executable commands are stored. When the executable commands are executed by a processor, the method described in any method embodiment of the present disclosure is executed.
[0096] <Second Embodiment>
[0097] <Method>
[0098] This embodiment provides a neural machine translation method, which translates a source sentence to be translated by applying the neural machine translation model trained by the above embodiment.
[0099] Specifically, Figure 7 As shown, the method includes the following steps 7100 to 7200:
[0100] Step 7100, obtaining the source sentence to be translated.
[0101] The source sentence to be translated may be in a certain language, such as Chinese, English or Japanese.
[0102] Step 7200, input the source sentence to be translated into the neural machine translation model, and output the target sentence as the translation result.
[0103] Among them, the neural machine translation model is obtained by obtaining the translation memory sentence pairs with the highest similarity to each source sentence in the parallel corpus from the translation memory; and the source language sentence pairs and the translation memory sentence pairs are used as training sample sets to train the preset initial model until the first parameter of the translation memory source encoder layer of the preset initial model converges, the second parameter of the translation memory target encoder layer of the preset initial model converges, and the third parameter of the decoder layer containing translation memory information of the preset initial model converges. Among them, the source sentence pair includes the source sentence and the corresponding target sentence, and the translation memory sentence pair includes the translation memory source sentence and the corresponding translation memory target sentence. The specific training process can be referred to in the first embodiment above, and will not be repeated here.
[0104] The neural machine translation method of this embodiment utilizes the source sentence s existing in the translation memory i Very similar translation memory pairs To help the neural network translation model translate the source sentence s i , which can improve the accuracy of neural machine translation results and improve the quality of translation.
[0105] <Device>
[0106] This embodiment provides a neural machine translation device, which is, for example, Figure 8 The neural machine translation device 8000 is shown.
[0107] like Figure 8 As shown, the neural machine translation device 8000 may include an acquisition module 8100 and a translation module 8200 .
[0108] The acquisition module 8100 is used to acquire source sentences to be translated.
[0109] The translation module 8200 is used to input the source sentence to be translated into the neural machine translation model and output the target sentence as the translation result; wherein the neural machine translation model is trained according to the training method of the neural machine translation model in the first embodiment mentioned above.
[0110] The neural machine translation device of this embodiment can be used to execute the method and technical solution of this embodiment. Its implementation principle and technical effects are similar and will not be repeated here.
[0111] <Device>
[0112] In this embodiment, an electronic device is also provided, which includes the neural machine translation device 8000 described in the device embodiment of the present disclosure; or, the electronic device is Fig. 9 The electronic device 9000 shown includes a processor 9200 and a memory 9100:
[0113] The memory 9100 stores machine executable instructions that can be executed by the processor; the processor 9200 executes the machine executable instructions to implement the neural machine translation method as described in any one of the embodiments.
[0114] <Computer Readable Storage Medium Embodiment>
[0115] This embodiment provides a computer-readable storage medium, in which executable commands are stored. When the executable commands are executed by a processor, the method described in any method embodiment of the present disclosure is executed.
[0116] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0117] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0118] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0119] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0120] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0121] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0122] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0123] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of an instruction, and a part of the module, program segment or instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented by a dedicated hardware-based system that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by a combination of software and hardware.
[0124] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the marketplace, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for training a neural machine translation model, the method comprising: Obtaining from the translation memory a translation memory sentence pair with the highest similarity to each source sentence in the parallel corpus; Using source language sentence pairs and translation memory sentence pairs as training sample sets to train a preset initial model until a first parameter of a translation memory source-end encoder layer of the preset initial model converges, a second parameter of a translation memory target-end encoder layer of the preset initial model converges, and a third parameter of a decoder layer containing translation memory information of the preset initial model converges, to obtain a neural machine translation model, wherein the preset initial model is constructed based on a feedforward neural network, a cross-attention network, a self-attention network, and an encoder-decoder attention network, the translation memory source-end encoder layer is used to determine a context representation of a translation memory source-end sentence, the translation memory target-end encoder layer is used to determine a context representation of a translation memory target-end sentence, and the translation memory information decoder layer is used to determine a target-end sentence; The source-end sentence pair includes a source-end sentence and a corresponding target-end sentence, and the translation memory sentence pair includes a translation memory source-end sentence and a corresponding translation memory target-end sentence.
2. The method according to claim 1, wherein: The step of obtaining the translation memory sentence pair with the highest similarity to each source sentence in the parallel corpus from the translation memory includes: Calculating the similarity between the source sentence and each of the translation memory source sentences in the translation memory; The similarities are sorted from large to small, and the translation memory source-end sentence with the greatest similarity and its corresponding target-end sentence are determined as the translation memory sentence pair with the highest similarity to the source-end sentence.
3. The method according to claim 2, wherein: The calculating the similarity between the source sentence and each source sentence in the translation memory includes: The edit distance between the source sentence and each translation memory source sentence in the translation memory is calculated as the similarity.
4. The method according to claim 1, wherein: The method of using the source language sentence pairs and the translation memory sentence pairs as training sample sets to train the preset initial model until a first parameter of a translation memory source-end encoder layer of the preset initial model converges, a second parameter of a translation memory target-end encoder layer of the preset initial model converges, and a third parameter of a decoder layer containing translation memory information of the preset initial model converges, comprises: Translate each sample in the training sample set based on the preset initial model to obtain a translation result; Substituting the translation result into a preset loss function for calculation to obtain the loss of each sample; The first parameter, the second parameter, and the third parameter are updated based on the loss until the first parameter, the second parameter, and the third parameter converge.
5. The method according to claim 4, wherein: The updating of the first parameter, the second parameter, and the third parameter based on the loss until the first parameter, the second parameter, and the third parameter converge, comprises: Based on the loss and a preset back propagation algorithm, respectively calculating a first derivative of the first parameter, a second derivative of the second parameter, and a third derivative of the third parameter; updating the first parameter based on the first derivative and a gradient descent algorithm, updating the second parameter based on the second derivative and a gradient descent algorithm, and updating the third parameter based on the third derivative and a gradient descent algorithm; The first parameter, the second parameter and the third parameter are updated multiple times based on the losses of multiple samples in the training sample set until convergence, thereby obtaining the neural machine translation model.
6. A neural machine translation method, the method comprising: Get the source sentence to be translated; The source sentence to be translated is input into a neural machine translation model, and a target sentence is output as a translation result; wherein the neural machine translation model is trained according to the training method of the neural machine translation model according to any one of claims 1 to 5.
7. A training device for a neural machine translation model, the device comprising: An acquisition module is used to acquire, from the translation memory, the translation memory sentence pair with the highest similarity to each source sentence in the parallel corpus; A training module is used to train a preset initial model using source language sentence pairs and translation memory sentence pairs as training sample sets until a first parameter of a translation memory source-end encoder layer of the preset initial model converges, a second parameter of a translation memory target-end encoder layer of the preset initial model converges, and a third parameter of a decoder layer containing translation memory information of the preset initial model converges, thereby obtaining a neural machine translation model, wherein the preset initial model is constructed based on a feedforward neural network, a spanning attention network, a self-attention network, and an encoder-decoder attention network, the translation memory source-end encoder layer is used to determine a context representation of a translation memory source-end sentence, the translation memory target-end encoder layer is used to determine a context representation of a translation memory target-end sentence, and the translation memory information decoder layer is used to determine a target-end sentence; wherein the source-end sentence pair includes a source-end sentence and a corresponding target-end sentence, and the translation memory sentence pair includes a translation memory source-end sentence and a corresponding translation memory target-end sentence.
8. A neural machine translation device, the device comprising: The acquisition module is used to obtain the source sentence to be translated; A translation module, used to input the source sentence to be translated into a neural machine translation model and output a target sentence as a translation result; wherein the neural machine translation model is trained according to the training method of the neural machine translation model according to any one of claims 1 to 5.
9. An electronic device, characterized in that: It comprises a processor and a memory; the memory stores machine executable instructions that can be executed by the processor; the processor executes the machine executable instructions to implement the training method of the neural machine translation model described in any one of claims 1-5.
10. An electronic device, characterized in that: It comprises a processor and a memory; the memory stores machine executable instructions that can be executed by the processor; the processor executes the machine executable instructions to implement the neural machine translation method according to claim 6.
Citation Information
Patent Citations
Fast incremental fuzzy matching method of cloud translation memory bank
CN107329961A
A method for integrating translation memory into neural machine translation through gating mechanism
CN109299479A