Non-autoregressive neural machine translation decoding method, device, equipment and storage medium

By adding a conditional random field model to the neural network model, the sentence decoding length of the target language is dynamically determined, which solves the problem that the length of the text to be translated cannot be dynamically determined in the non-autoregressive neural machine translation method, and improves the translation quality and accuracy.

CN114611505BActive Publication Date: 2025-06-03BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210224779.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-06-03
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

The current non-autoregressive neural machine translation method cannot dynamically determine the length of the target language text to be translated during the translation process, resulting in the problem of repeated translation or missing translation.

Method used

By adding a pre-trained conditional random field model to the top of the neural network model, the dependency between the target words in the target sentence is established, and the corresponding labels output at the output end are determined, thereby dynamically determining the sentence decoding length of the target language.

Benefits of technology

Repeated translations or omissions caused by incorrect predefined decoding lengths are avoided, and the translation quality and accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611505B_ABST
    Figure CN114611505B_ABST
Patent Text Reader

Abstract

The present application discloses a non-autoregressive neural machine translation decoding method, apparatus, device and storage medium. The method includes: obtaining a text to be translated in a source language and a word vector corresponding to a word to be translated in the text to be translated; preprocessing the text to be translated and performing vector encoding on the word vector corresponding to the word to be translated to obtain an encoded vector that focuses on context information; translating the text to be translated into a target sentence in a target language through a pre-trained neural network model according to the word vector corresponding to the word to be translated and the encoded vector; establishing a dependency relationship between target words in the target sentence through a pre-trained conditional random field model and outputting the target sentence. In the present application, the sentence decoding length of the target language is dynamically determined during the decoding process without pre-defining the decoding length, which can avoid the phenomenon of repeated translation or missed translation caused by incorrect pre-defined decoding length, and improve the translation quality and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a non-autoregressive neural machine translation decoding method, apparatus, device, and storage medium. Background Art

[0002] Currently, a commonly used autoregressive decoding method in neural machine translation decodes the target language sequentially from left to right according to the sentence. However, this decoding characteristic of autoregression causes that during the decoding process, words at different positions cannot be generated in parallel. To overcome this difficulty, a non-autoregressive neural machine translation method is adopted, which does not need to consider the timing of the target-side language generation process and can generate all target language vocabulary simultaneously during the decoding process, which can greatly improve the decoding speed of the model.

[0003] In the process of implementing the embodiments of the present disclosure, it is found that there are at least the following problems in the related art:

[0004] Although the current non-autoregressive neural machine translation method can generate all target language vocabulary at all times simultaneously and greatly improve the decoding speed, during the decoding process, it is necessary to determine the length of the target language text in advance according to the statistical model, and the length of the target language text to be translated cannot be dynamically determined during the translation process, resulting in problems of repeated translation or missing translation.

[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] In order to solve the technical problem that the length of the target language text to be translated cannot be dynamically determined during the translation process in the current non-autoregressive neural machine translation method, resulting in repeated translation or missing translation, the present application provides a non-autoregressive neural machine translation decoding method, apparatus, device, and storage medium.

[0007] In some embodiments, the present application provides a non-autoregressive neural machine translation decoding method, and the method includes the following steps:

[0008] Obtain the text to be translated in the source language, and the word vector corresponding to the word to be translated in the text to be translated;

[0009] Preprocess the text to be translated, and perform vector encoding on the word vector corresponding to the word to be translated to obtain an encoded vector that focuses on context information;

[0010] According to the word vector corresponding to the word to be translated and the encoded vector, translate the text to be translated into a target sentence in the target language through a pre-trained neural network model;

[0011] Establish the dependency relationship between target words in the target sentence through a pre-trained conditional random field model, and output the target sentence.

[0012] Optionally, the training process of the neural network model includes:

[0013] Preprocess the training sample data of the source language to obtain the sample word vector representation of the source language;

[0014] Input the sample word vector representation into the initial neural network model so that the initial neural network model outputs the predicted translation of the target language;

[0015] If the similarity between the predicted translation and the historical translation is greater than or equal to the set similarity threshold, the initial neural network model is successfully trained to obtain a trained neural network model;

[0016] If the similarity between the predicted translation and the historical translation is less than the set similarity threshold, adjust the parameters in the initial neural network model until the initial neural network model is successfully trained.

[0017] Optionally, preprocessing the training sample data of the source language to obtain the sample word vector representation of the source language includes:

[0018] Perform subword segmentation on the sample sentences in the training sample data to obtain a number of sample subword sequences;

[0019] Use the first label to pad the sample input sequence of the source language to a preset subword sequence length, and use the second label to pad the sample output sequence of the target language to a preset subword sequence length;

[0020] Randomly initialize the sample input sequence of the source language to obtain a word vector encoding representing the sample word vector corresponding to each sample subword in the sample input sequence.

[0021] Optionally, inputting the sample word vector representation into the initial neural network model so that the initial neural network model outputs the predicted translation of the target language includes:

[0022] Input the sample word vector representation into the encoder of the initial neural network model, perform vector encoding on the sample word vector representation, and obtain the position vector encoding of the sample input sequence;

[0023] According to the sum of the position vector encoding and the word vector encoding, obtain the input vector encoding of the source language;

[0024] Through the Transformer layer based on the self-attention mechanism, the input vector encoding is encoded by the encoder to obtain the top-level encoding;

[0025] Perform a linear transformation on the top-level encoding;

[0026] Based on the result of the linear transformation, set the output probability distribution at each moment through a pre-trained conditional random field model, and take the word corresponding to the maximum probability value as the generation result at the corresponding moment;

[0027] Decode sequentially, and take the position where the second label is output as the end position of the sentence to obtain the predicted translation.

[0028] Optionally, the top-level encoding obtained by encoding the input vector through the encoder is calculated by the following formula:

[0029]

[0030] V n = SelfAttn(E X , E X , E X );

[0031] Among them, E x represents the input vector encoding, V n represents the output of the Transformer layer based on the self-attention mechanism, represents the top-level encoding representation obtained by encoding through the encoder.

[0032] Optionally, based on the result of the linear transformation, set the output probability distribution at each moment through a pre-trained conditional random field model, which is calculated by the following formula:

[0033]

[0034] Among them, s represents the score of the target word y i predicted according to the Transformer layer, t represents the transition probability between words, z(x) represents the normalization factor; n represents the number of sub-words; x represents the word of the source language at the input end, y represents the word of the target language at the output end, obtained through linear transformation

[0035] The word corresponding to the maximum probability value is used as the generation result at moment i, and is calculated by the following formula:

[0036] y = Max(Prob y / x );

[0037] The generation results obtained by decoding sequentially are:

[0038] y = [y 1 ,..., y n; where, [y 1 ,..., y j represents the sample output sequence, and [y j+1 ,..., y n represents the second tag used to pad the sample output sequence to a preset sub-word sequence length; x = [x 1 ,..., x n represents the sub-word sequence of the source language, [x 1 ,..., x i represents the sample input sequence, and [x i+1 ,..., x n represents the first tag used to pad the sample input sequence to a preset sub-word sequence length.

[0039] Optionally, performing sub-word segmentation on the sample sentences in the training sample data includes:

[0040] Performing sub-word segmentation on the sample sentences in the training sample data by using the BPE tokenization algorithm.

[0041] This application also provides a non-autoregressive neural machine translation decoding device, and the device includes:

[0042] An acquisition unit, configured to acquire the text to be translated in the source language and the word vector corresponding to the word to be translated in the text to be translated;

[0043] A preprocessing and encoding unit, configured to preprocess the text to be translated and perform vector encoding on the word vector corresponding to the word to be translated to obtain an encoded vector that focuses on context information;

[0044] A translation unit, configured to translate the text to be translated into a target sentence in the target language through a pre-trained neural network model according to the word vector corresponding to the word to be translated and the encoded vector; and

[0045] An output unit, configured to establish a dependency relationship between the target words in the target sentence through a pre-trained conditional random field model and output the target sentence.

[0046] This application also provides an electronic device, and the electronic device includes: at least one processor, a memory, at least one network interface, and a user interface;

[0047] The at least one processor, the memory, the at least one network interface, and the user interface are coupled together through a bus system;

[0048] The processor is configured to execute the steps of the above non-autoregressive neural machine translation decoding method by calling the program or instruction stored in the memory.

[0049] The present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the non-autoregressive neural machine translation decoding method as described above are implemented.

[0050] The above technical solution provided by the embodiments of the present application has the following advantages compared with the prior art:

[0051] A non-autoregressive neural machine translation decoding method, device, equipment and storage medium provided by the embodiments of the present application can complete the non-autoregressive neural machine translation decoding task only by using a pre-trained neural network model compared with the traditional encoder-decoder model. At the same time, a pre-trained conditional random field model is added on the top of the neural network model to establish a context time series dependence relationship. During the decoding process, it is judged whether the end position of the target language sentence is reached by the corresponding tags output at the output end, so as to judge whether the decoding is ended. By dynamically determining the decoding length of the target language sentence during the decoding process without pre-defining the decoding length, the phenomenon of repeated translation or missing translation caused by incorrect pre-defined decoding length can be avoided, and the translation quality and accuracy can be improved. Description of the Drawings

[0052] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained according to these drawings without creative labor.

[0054] Figure 1 It is a schematic flowchart of a non-autoregressive neural machine translation decoding method provided by an embodiment of the present application;

[0055] Figure 2 It is a schematic flowchart of another non-autoregressive neural machine translation decoding method provided by an embodiment of the present application;

[0056] Figure 3 It is a schematic structural diagram of a neural network model provided by an embodiment of the present application;

[0057] Figure 4 It is a schematic flowchart of yet another non-autoregressive neural machine translation decoding method provided by an embodiment of the present application;

[0058] Figure 5Schematic diagram of the structure of a non-autoregressive neural machine translation decoding device provided by an embodiment of the present application;

[0059] Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0061] Figure 1 Flowchart of a non-autoregressive neural machine translation decoding method provided by an embodiment of the present application. The method specifically includes the following steps:

[0062] S101. Obtain the text to be translated in the source language and the word vectors corresponding to the words to be translated in the text to be translated.

[0063] For example, the source language can be languages such as Chinese and English, and the source language is translated into the corresponding target language. The target language can be languages such as English and Chinese, or can also be languages such as Japanese and Korean. The embodiments of the present application do not make any limitations thereto.

[0064] S102. Preprocess the text to be translated and perform vector encoding on the word vectors corresponding to the words to be translated to obtain encoded vectors that focus on context information.

[0065] Optionally, in order to reduce the impact of out-of-vocabulary words on the model generation performance, the BPE tokenization algorithm can be used to perform sub-word segmentation on the sentence to be translated in the text to be translated to generate a sub-word sequence, so that both the input unit of the encoder and the output unit of the decoder are sub-word sequences.

[0066] S103. According to the word vectors corresponding to the words to be translated and the encoded vectors, translate the text to be translated into the target sentence in the target language through a pre-trained neural network model.

[0067] S104. Through a pre-trained conditional random field model, establish the dependency relationship between the target words in the target sentence and output the target sentence.

[0068] Compared with the traditional encoder-decoder model, a non-autoregressive neural machine translation decoding method provided by this application can complete the non-autoregressive neural machine translation decoding task by only using a pre-trained neural network model. At the same time, by adding a pre-trained conditional random field model on top of the neural network model, a context temporal dependence relationship is established. During the decoding process, it is judged whether the end position of the target language sentence is reached through the corresponding tags output at the output end, so as to judge whether the decoding is completed. By dynamically determining the decoding length of the target language sentence during the decoding process without pre-defining the decoding length, the phenomenon of repeated translation or missing translation caused by incorrect pre-defined decoding length can be avoided, and the translation quality and accuracy can be improved.

[0069] In another embodiment of this application, the training process of the neural network model includes:

[0070] Preprocess the training sample data of the source language to obtain the sample word vector representation of the source language;

[0071] Input the sample word vector representation into the initial neural network model so that the initial neural network model outputs the predicted translation of the target language;

[0072] If the similarity between the predicted translation and the historical translation is greater than or equal to the set similarity threshold, the initial neural network model is successfully trained to obtain a trained neural network model;

[0073] If the similarity between the predicted translation and the historical translation is less than the set similarity threshold, adjust the parameters in the initial neural network model until the initial neural network model is successfully trained.

[0074] Optionally, as Figure 2 shown, in another embodiment of this application, preprocessing the training sample data of the source language to obtain the sample word vector representation of the source language includes the following steps:

[0075] S201. Sub-word segment the sample sentences in the training sample data to obtain a number of sample sub-word sequences.

[0076] Optionally, sub-word segment the sample sentences in the training sample data by the BPE tokenization algorithm. During the process of training the neural network model in this application, the BPE tokenization algorithm is used to sub-word segment the sample sentences in all the training sample data, which can make the input units of the encoder and the output units of the decoder both sub-word sequences, thereby reducing the impact of out-of-vocabulary words on the generation performance of the neural network model.

[0077] S202. Use the first tag to pad the sample input sequence in the source language to a preset sub-word sequence length, and use the second tag to pad the sample output sequence in the target language to a preset sub-word sequence length.

[0078] As Figure 3 shown, it is a schematic structural diagram of the neural network model of the present application. During the training process of the neural network model, a sub-word sequence length, that is, the maximum input length (for example, the input length is 128), is predefined. At the input end of the neural network model, use the first tag <pad>Pad the sentences with insufficient input lengths in the source language to the preset sub-word sequence length (i.e., the input length of 128); while at the output end of the neural network model, use the second label <eos>Pad the sentences with insufficient output lengths in the target language to the preset sub-word sequence length (i.e., the output length of 128). Therefore, during the decoding process, when the neural network model continuously outputs two second labels <eos>When it indicates the end of the output, this method does not require a predefined decoding length, but rather determines it during the decoding process through the second tag of the output <eos>Based on the dynamics of [], it is determined whether the output length of the sentence in the target language has reached the position of the end of the sentence in the target language, so as to determine whether the decoding is completed, without pre-defining the decoding length, which can avoid the phenomena of repeated translation or missed translation caused by incorrect pre-defined decoding length, and improve the translation quality and accuracy.

[0079] S203. Randomly initialize the sample input sequence in the source language to obtain a word vector encoding representing the sample word vectors corresponding to each sample sub-word in the sample input sequence.

[0080] Optionally, by defining x = [x 1 ,..., x n to represent the sub-word sequence in the source language, [x 1 ,..., x i to represent the sample input sequence, and [x i+1 ,..., x n to represent the first label for padding the sample input sequence to a preset sub-word sequence length. <pad>The word vector encoding TE is obtained by randomly initializing the sample input sequence in the source language. X =[V 1 ,...,V n , where V i represents the sample word vector corresponding to the i-th sample sub-word.

[0081] As Figure 4 shown, in another embodiment of the present application, inputting the sample word vector representation into the initial neural network model to enable the initial neural network model to output a predicted translation of the target language includes the following steps:

[0082] S401. Input the sample word vector representation into the encoder of the initial neural network model to perform vector encoding on the sample word vector representation and obtain the positional vector encoding of the sample input sequence.

[0083] S402. Obtain the input vector encoding of the source language according to the sum of the positional vector encoding and the word vector encoding.

[0084] The input vector encoding of the source language is calculated by the following formula: E X =TE X +PE X , where E X represents the input vector encoding of the source language, TE X represents the word vector encoding, and PE X represents the positional vector encoding.

[0085] S403. Through the Transformer layer based on the self-attention mechanism, the input vector encoding is encoded by the encoder to obtain the top-level encoding.

[0086] Optionally, the top-level encoding obtained by encoding the input vector encoding by the encoder is calculated by the following formula:

[0087]

[0088] V n =SelfAttn(E X ,E X ,E X );

[0089] where E x represents the input vector encoding, V n represents the output of the Transformer layer based on the self-attention mechanism, represents the top-level encoding representation obtained by encoding through the encoder.

[0090] S404. Perform a linear transformation on the top-level encoding;

[0091] Obtained through the linear transformation f()

[0092] S405. Based on the result of the linear transformation, set the output probability distribution at each moment through a pre-trained conditional random field model, and use the word corresponding to the maximum probability value as the generation result at the corresponding moment.

[0093] S406. Decode sequentially, and use the position where the second label is output as the end position of the sentence to obtain the predicted translation.

[0094] Optionally, in step S405, based on the result of the linear transformation, control the output probability distribution at each moment through the transition chain of the pre-trained conditional random field model, which is calculated by the following formula:

[0095]

[0096] where s represents the score of the target word y predicted according to the Transformer layer, t represents the transfer probability between words, z(x) represents the normalization factor; n represents the number of sub-words; x represents the word of the source language at the input end, and y represents the word of the target language at the output end. i Obtained through linear transformation Obtained through linear transformation

[0097] The word corresponding to the maximum probability value is used as the generation result at moment i, and is calculated by the following formula:

[0098] y = Max(Prob y / x );

[0099] The generation results obtained by decoding sequentially are:

[0100] y = [y 1 ,..., y n ; where [y 1 ,..., y j represents the sample output sequence, [y j+1 ,..., y n represents the second label used to pad the sample output sequence to the preset sub-word sequence length; x = [x 1 ,..., x n represents the sub-word sequence of the source language, [x 1 ,..., x i represents the sample input sequence, [x i+1 ,..., x n Indicates the first tag for padding a sample input sequence to a preset sub-word sequence length.

[0101] An embodiment of the present application proposes a dynamic decoding method for non-autoregressive neural machine translation. This method only uses a set of representer models (trained neural network models) to replace the existing encoder and decoder structures, and adds a pre-trained conditional random field model on top of the representer model to build context dependencies. During the model training process, the first tag is used <pad>and the second label <eos>Pad the input sequence and the sample output sequence respectively, so that during the decoding process, since the conditional random field model does not establish information about the second label <eos>Transfer relationship to target language characters, when the second label is output <eos>When this happens, it can be considered that the decoding is completed. This method does not require a predefined decoding length, but dynamically determines the length of the target language during the decoding process, thus alleviating the situation of repeated translation and missing translation that occur due to incorrect predefinition of the translation length in the decoding process of non-autoregressive neural machine translation, and improving the translation quality.

[0102] As Figure 5 shown, an embodiment of the present application provides a non-autoregressive neural machine translation decoding device, and the device includes:

[0103] An acquisition unit 51, configured to acquire a text to be translated in a source language, and a word vector corresponding to a word to be translated in the text to be translated;

[0104] A preprocessing and encoding unit 52, configured to preprocess the text to be translated, and perform vector encoding on the word vector corresponding to the word to be translated, so as to obtain an encoded vector that pays attention to context information;

[0105] A translation unit 53, configured to translate the text to be translated into a target sentence in a target language through a pre-trained neural network model according to the word vector corresponding to the word to be translated and the encoded vector; and

[0106] An output unit 54, configured to establish a dependency relationship between target words in the target sentence through a pre-trained conditional random field model, and output the target sentence.

[0107] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in each method embodiment are implemented, for example, including:

[0108] Acquire a text to be translated in a source language, and a word vector corresponding to a word to be translated in the text to be translated;

[0109] Preprocess the text to be translated, and perform vector encoding on the word vector corresponding to the word to be translated, so as to obtain an encoded vector that pays attention to context information;

[0110] Translate the text to be translated into a target sentence in a target language through a pre-trained neural network model according to the word vector corresponding to the word to be translated and the encoded vector;

[0111] Establish a dependency relationship between target words in the target sentence through a pre-trained conditional random field model, and output the target sentence.

[0112] Figure 6 It is a schematic structural diagram of an electronic device provided by another embodiment of the present invention. Figure 6 The electronic device 600 shown includes: at least one processor 601, a memory 602, at least one network interface 604, and other user interfaces 603. Each component in the electronic device 600 is coupled together through a bus system 605. It can be understood that the bus system 605 is used to implement the connection and communication between these components. In addition to a data bus, the bus system 605 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 6 all kinds of buses are labeled as the bus system 605.

[0113] Among them, the user interface 603 may include a display, a keyboard, or a pointing device (for example, a mouse, a trackball, a touchpad, or a touch screen, etc.).

[0114] It can be understood that the memory 602 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memory 602 described herein is intended to include but not be limited to these and any other suitable types of memory.

[0115] In some embodiments, the memory 602 stores the following elements, executable units, or data structures, or subsets thereof, or extended sets thereof: an operating system 6021 and an application program 6022.

[0116] Among them, the operating system 6021 includes various system programs, such as the framework layer, the core library layer, the driver layer, etc., which are used to implement various basic services and handle hardware-based tasks. The application program 6022 includes various application programs, such as the MediaPlayer, the Browser, etc., which are used to implement various application services. The program for implementing the method of the embodiment of the present invention may be included in the application program 6022.

[0117] In the embodiment of the present invention, by calling the program or instruction stored in the memory 602, specifically, it may be the program or instruction stored in the application program 6022, the processor 601 is used to execute the method steps provided by each method embodiment, for example, including:

[0118] Obtain the text to be translated in the source language, and the word vector corresponding to the word to be translated in the text to be translated;

[0119] Preprocess the text to be translated, and perform vector encoding on the word vector corresponding to the word to be translated to obtain an encoded vector that pays attention to context information;

[0120] According to the word vector corresponding to the word to be translated and the encoded vector, translate the text to be translated into the target sentence in the target language through a pre-trained neural network model;

[0121] Through a pre-trained conditional random field model, establish the dependency relationship between the target words in the target sentence, and output the target sentence.

[0122] The method disclosed in the embodiments of the present invention above can be applied to the processor 601 or implemented by the processor 601. The processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 601 or instructions in the form of software. The above-mentioned processor 601 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or completed by the combination of the hardware and software units in the decoding processor. The software unit may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 602, and the processor 601 reads the information in the memory 602 and combines its hardware to complete the steps of the above method.

[0123] It can be understood that these embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.

[0124] For software implementation, the techniques described herein can be implemented by units that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.

[0125] For the convenience of description, when describing the above device, various units are described separately according to their functions. Of course, when implementing the present invention, the functions of each unit can be implemented in one or more software and / or hardware.

[0126] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for device or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0127] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

[0128] The above are only specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.< / eos> < / eos> < / eos> < / pad> < / pad> < / eos> < / eos> < / eos> < / pad>

Claims

1. A non-autoregressive neural machine translation decoding method, characterized in that, the method comprises the following steps: obtain the text to be translated in the source language, and the word vector corresponding to the word to be translated in the text to be translated; perform preprocessing on the text to be translated, and perform vector encoding on the word vector corresponding to the word to be translated to obtain an encoded vector that pays attention to context information; translate the text to be translated into a target sentence in the target language through a pre-trained neural network model according to the word vector corresponding to the word to be translated and the encoded vector; establish the dependency relationship between target words in the target sentence through a pre-trained conditional random field model, and output the target sentence.

2. The method according to claim 1, characterized in that, the training process of the neural network model includes: perform preprocessing on the training sample data in the source language to obtain a sample word vector representation of the source language; input the sample word vector representation into the initial neural network model so that the initial neural network model outputs a predicted translation in the target language; if the similarity between the predicted translation and the historical translation is greater than or equal to a set similarity threshold, the initial neural network model is successfully trained to obtain a trained neural network model; if the similarity between the predicted translation and the historical translation is less than the set similarity threshold, adjust the parameters in the initial neural network model until the initial neural network model is successfully trained.

3. The method according to claim 2, characterized in that, performing preprocessing on the training sample data in the source language to obtain a sample word vector representation of the source language includes: perform sub-word segmentation on the sample sentences in the training sample data to obtain a number of sample sub-word sequences; use the first label to pad the sample input sequence in the source language to a preset sub-word sequence length, and use the second label to pad the sample output sequence in the target language to a preset sub-word sequence length; randomly initialize the sample input sequence in the source language to obtain a word vector encoding representing the sample word vector corresponding to each sample sub-word in the sample input sequence.

4. The method according to claim 3, characterized in that, inputting the sample word vector representation into the initial neural network model so that the initial neural network model outputs a predicted translation in the target language includes: input the sample word vector representation into the encoder of the initial neural network model, perform vector encoding on the sample word vector representation to obtain a position vector encoding of the sample input sequence; obtain the input vector encoding of the source language according to the sum of the position vector encoding and the word vector encoding; through the Transformer layer based on the self-attention mechanism, the input vector encoding is encoded by the encoder to obtain a top-level encoding; perform a linear transformation on the top-level encoding; based on the result of the linear transformation, set the output probability distribution at each moment through a pre-trained conditional random field model, and use the word corresponding to the maximum probability value as the generation result at the corresponding moment; decode sequentially, and take the position where the second label is output as the end position of the sentence to obtain the predicted translation.

5. The method according to claim 4, wherein, the top-level encoding obtained by encoding the input vector encoding through the encoder is calculated by the following formula: V n = SelfAttn(E X , E X , E X ); Among them, E x represents the input vector encoding, V n represents the output of the Transformer layer based on the self-attention mechanism, represents the top-level encoded representation obtained by encoding through the encoder.

6. The method according to claim 5, wherein, based on the result of the linear transformation, the output probability distribution at each moment is set through a pre-trained conditional random field model, and is calculated by the following formula: where s represents the score of the target word y predicted according to the Transformer layer, t represents the transfer probability between words, z(x) represents the normalization factor; n represents the number of subwords; x represents the word of the source language at the input end, and y represents the word of the target language at the output end i , t represents the transfer probability between words, z(x) represents the normalization factor; n represents the number of subwords; x represents the word of the source language at the input end, and y represents the word of the target language at the output end obtained through a linear transformation The word corresponding to the maximum probability value is used as the generation result at moment i, and is calculated by the following formula: y = Max(Prob y / x ); The generation results obtained by decoding in sequence are: y = [y 1 ,..., y n ; where, [y 1 ,..., y j represents the sample output sequence, and [y j+1 ,..., y n represents the second label used to pad the sample output sequence to a preset sub-word sequence length; x = [x 1 ,..., x n represents the sub-word sequence of the source language, [x 1 ,..., x i represents the sample input sequence, and [x i+1 ,..., x n represents the first label used to pad the sample input sequence to a preset sub-word sequence length.

7. The method according to claim 3, wherein, performing sub-word segmentation on the sample sentences in the training sample data, including: Performing sub-word segmentation on the sample sentences in the training sample data through the BPE word segmentation algorithm.

8. A non-autoregressive neural machine translation decoding device, wherein, the device includes: an acquisition unit configured to acquire the text to be translated in the source language and the word vector corresponding to the word to be translated in the text to be translated; a preprocessing and encoding unit configured to preprocess the text to be translated and perform vector encoding on the word vector corresponding to the word to be translated to obtain an encoding vector that focuses on context information; a translation unit configured to translate the text to be translated into a target sentence in the target language through a pre-trained neural network model according to the word vector corresponding to the word to be translated and the encoding vector; and an output unit configured to establish the dependency relationship between the target words in the target sentence through a pre-trained conditional random field model and output the target sentence.

9. An electronic device, wherein, the electronic device includes: at least one processor, a memory, at least one network interface, and a user interface; the at least one processor, the memory, the at least one network interface, and the user interface are coupled together through a bus system; the processor is configured to execute the steps of the non-autoregressive neural machine translation decoding method according to any one of claims 1 to 7 by calling the program or instruction stored in the memory.

10. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the non-autoregressive neural machine translation decoding method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Non-autoregressive neural machine translation method and device, computer device and medium

    CN110852116A

  • Translated text processing method and device, computer equipment and storage medium

    CN111368531A