Editing mechanism-based prototype neural machine translation method and system, electronic equipment and readable storage medium

By adopting an editing mechanism-based method in prototype neural machine translation, using the Informer encoder and the maximum internal product search algorithm for prototype sequence search, and combining the dual encoder structure and the Transformer model for noise reduction processing, the problem of prototype noise integration is solved and the translation quality is significantly improved.

CN120087376APending Publication Date: 2025-06-03YUNNAN MINZU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150141.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

During the process of prototype neural machine translation, the noise incorporation problem caused by the introduction of prototypes affects the quality of translation.

Method used

Using an editing mechanism-based method, the semantic features of source language sentences and prototype candidate sentences are extracted through the Informer encoder, the prototype sequence search is searched with the maximum internal product search algorithm, and the dual encoder structure and Transformer model are used for noise reduction processing and translation generation.

Benefits of technology

Effectively identify and process noise in the prototype sequence, improve translation accuracy and quality, and significantly improve the performance of prototype neural machine translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087376A_ABST
    Figure CN120087376A_ABST
Patent Text Reader

Abstract

The invention relates to a prototype neural machine translation method and system based on an editing mechanism, electronic equipment and a readable storage medium, and belongs to the technical field of natural language processing. The method comprises the steps that target sentences of training corpora serve as a prototype candidate sentence library; performing semantic feature extraction on the source language sentences and the prototype candidate sentences through an Informer encoder, and performing retrieval in a vector space based on a maximum inner product search algorithm to obtain a prototype sequence; performing noise reduction processing on the obtained prototype sequence to obtain a denoised prototype sequence; the prototype sequence and the source language sentence which are subjected to noise reduction jointly serve as input of a translation model; the double-encoder structure independently encodes the source language sentences and the prototype sequences after noise reduction and combines semantic features of the source language sentences and the prototype sequences after noise reduction to generate high-quality translations. The method can effectively identify the noise in the prototype sequence, improves the translation accuracy, and is suitable for a multi-language translation environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a prototype neural machine translation method, system, electronic device, and readable storage medium based on an editing mechanism, belonging to the technical field of natural language processing. Background Art

[0002] A prototype sequence, abbreviated as a prototype, represents a special translation memory and consists of target sentences that are semantically highly similar to the source sentence. A prototype is a target language sentence that is highly similar to the input source language sentence and can be used to guide the style and content of the translation. Prototype machine translation can guide the model to generate a translation close to the semantics of the prototype sequence by introducing prototypes during the translation process. The effectiveness of this method depends to a large extent on the quality of the prototype library. If the data in the prototype library is inaccurate or incomplete, it may affect the translation quality. Therefore, exploring how to use a limited prototype library to guide neural machine translation has very important research and application value.

[0003] The current mainstream approach is to use monolingual memory and perform learnable memory retrieval in a cross-lingual manner. On this basis, a neural machine framework of an encoder-decoder is used for translation, which effectively improves the performance of machine translation to a certain extent. However, in most cases, the entity words between the prototype and the source language do not match, or some words in the prototype are not aligned with the words in the source language, and the prototype often shows certain limitations. Although using prototypes provides valuable insights for translation, the irrelevant information embedded in the prototype sequence will also introduce noise, which will be amplified in the subsequent process, thereby reducing the translation performance. Therefore, the present invention proposes a prototype neural machine translation method based on an editing mechanism. Summary of the Invention

[0004] The technical problem solved by the present invention is that the present invention provides a prototype neural machine translation method, system, electronic device, and readable storage medium based on an editing mechanism to solve the problem of noise incorporation brought about by introducing prototypes during the prototype neural machine translation process, use the effective information in the prototype sequence to guide neural machine translation, and significantly improve the performance of prototype neural machine translation.

[0005] The technical solution of the present invention is: A prototype neural machine translation method based on an editing mechanism, the method comprising:

[0006] Step1, Corpus acquisition: Obtain parallel training corpora, validation corpora, and test corpora for multiple language pairs for model training, parameter tuning, and model effect testing; Use the target sentences of the training corpus as a prototype candidate sentence library;

[0007] Step 2. Construction of the prototype retrieval model: Adopt a cross - language retrieval mechanism and a monolingual memory retrieval prototype sequence; Use an Informer encoder to extract semantic features from the source - language sentence and the prototype candidate sentences, and then, based on the maximum inner - product search algorithm, perform retrieval in the vector space after semantic feature extraction to obtain the prototype sequence;

[0008] Step 3. Construction of the prototype noise - processing model: On the basis of Step 2, perform noise reduction processing on the obtained prototype sequence to obtain the denoised prototype sequence;

[0009] Step 4. Construction of the translation model: On the basis of Step 3, the denoised prototype sequence and the source - language sentence will jointly serve as the input of the translation model; The dual - encoder structure generates high - quality translations by independently encoding the source - language sentence and the denoised prototype sequence respectively and combining the semantic features of the source - language sentence and the denoised prototype sequence.

[0010] Furthermore, the said Step 2 includes:

[0011] Step 2.1. Mark the source - language sentences in the training corpus of Step 1 as x. Take the source - language sentence x as the encoder input, and use the Informer encoder to extract the semantic features of the source - language sentence x;

[0012] First, generate the word - embedding representation E of the input through the embedding layer x , and the word - embedding representation E x is expressed as:

[0013]

[0014] where d is the dimension of the embedding vector, e(x i ) represents the embedding representation of the i - th source - language word, and n represents the number of source - language word embeddings;

[0015] Use the Informer encoder to extract the context feature H x , and the context feature H x is expressed as:

[0016]

[0017] The core of the Informer encoder includes a self - attention mechanism and sparse factor optimization. Use ProbSparseSelf - Attention to calculate the self - attention score, and the self - attention score is expressed as:

[0018]

[0019] where Q = W q H x, K = W k H x , V = W v H x are the query matrix, key matrix, and value matrix respectively, and d k is the dimension of the key;

[0020] Step2.2. For the prototype candidate sentence set T = {t 1 , t 2 ,..., t i ,...t k}, input them one by one into the Informer encoder; the embedding and encoding representation of the i-th prototype candidate sentence t i = {t i1 , t i2 ,..., t im} is:

[0021]

[0022] where m is the number of words in the prototype candidate sentence t i , t i represents the i-th prototype candidate sentence, and t im refers to the m-th word in the prototype candidate sentence t i ; is the entire word embedding matrix of the prototype candidate sentence t i , and its shape is where d is the dimension of the word embedding vector; e(t im ) represents the embedding vector of the m-th word in the sentence t i ; represents the context feature representation of the prototype candidate sentence t i after being processed by the Informer encoder, and its shape is also

[0023] Step2.3. Based on Step2.1 and Step2.2, use the maximum inner product search technique to perform a fast retrieval in the prototype candidate sentence set; calculate the correlation score f between the source language sentence x and the i-th prototype candidate sentence t i for candidate, and the correlation score f(x, t i ) is expressed as:

[0024]

[0025] where t i represents the i-th candidate sentence in the prototype candidate sentence set T = {t 1 , t 2 ,..., t k}, and H xIt represents the context features of the source language sentence x after being processed by the Informer encoder;

[0026] According to the high and low relevance scores, N P prototype sentences that are most relevant to the source language sentence are selected, T1 = As the prototype sequence t, N P represents the number of prototype sentences selected from the candidate set, j represents the index of the sentence selected from the prototype candidate sentence set T, and the value range of j is 1 ≤ j ≤ N P .

[0027] Furthermore, the said Step3 includes:

[0028] Step3.1: Invoke the fast-align tool to perform word alignment on the source language sentence x in Step1 and the prototype sequence t retrieved in Step2, and generate an alignment mapping between the source language sentence x and the prototype sequence t;

[0029] Specifically, the fast-align tool obtains the alignment relationship between each word in the source language sentence x and the prototype sequence t by maximizing the alignment probability between word pairs; for each word in the prototype sequence use a binary indicator function to judge whether it is aligned with any word in the source language sentence; the binary indicator function is expressed as:

[0030]

[0031] Step3.2: On the basis of Step3.1, process the prototype sequence t; all words that are not found to be aligned in the source language sentence will be regarded as noise and replaced by a special placeholder; the denoised prototype sequence t′ is expressed as:

[0032]

[0033] Step3.3: On the basis of Step3.1, extract the unaligned words in the source language sentence, and denote the set of unaligned words in the source language sentence as z, that is:

[0034]

[0035] where represents the i-th word in the source language sentence x;

[0036] By inputting the unaligned words in the source language sentence and the denoised prototype sequence t′ into the Transformer model together, it is used to ensure that the key information of the source language is not missed in the final translation;

[0037] Step 3.4. Generate the final target sentence using the denoised prototype sequence t′ and the set z of unaligned words in the source language. Use the data of the prototype sequence t′ and the set z of unaligned words in the source language as inputs and feed them into the Transformer model; adopt a dual-input structure, and use the prototype sequence t′ and the set z of unaligned words in the source language as two different parts for input; the input result E inpu t is expressed as:

[0038] E input = Concat(Emmbed(t′), Embed(z));

[0039] Among them, Embed represents the word embedding layer, and Concat represents the concatenation operation.

[0040] Furthermore, Step 3.4 also includes: by adjusting the self-attention mechanism to avoid placeholder tokens from participating in the attention calculation, an attention mask matrix M is introduced;

[0041] The attention mask matrix M is negative infinity at the placeholder positions, and the weights at the negative infinity positions are reduced to zero during the calculation of the attention scores; the self-attention calculation formula is:

[0042]

[0043] where Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key, and M is the attention mask matrix.

[0044] Furthermore, Step 4 includes:

[0045] Step 4.1. The source language sentence x in Step 1 is fed into the source language encoder, and the source language encoder converts the source language sentence x into a dense vector representation; at the same time, the denoised prototype sequence t′ after Step 3 is also fed into the prototype encoder, and the prototype encoder converts each sentence t′i of the denoised prototype sequence into a group of word embeddings where k represents the k-th prototype sentence in the prototype candidate sentence set, and Li is the sentence length of the prototype sequence; semantic feature extraction is performed on the denoised prototype sequence through a neural network layer to obtain the context representation of the denoised prototype sequence;

[0046] Step 4.2. The semantic features respectively generated by the source language encoder and the prototype encoder need to be fused through a cross-attention mechanism; calculate the cross-attention:

[0047]

[0048] Among them, αi,j calculates the attention of the hidden state ht to the j-th sub-word in the i-th denoised prototype sequence sentence. Wm is a dimensional transformation matrix. ct is the context vector obtained by weighted summing each word of each sentence t′i in the denoised prototype sequence using αi,j. β is a learnable hyperparameter used to control the influence degree of the prototype sentence correlation score f(x, t i ) in the calculation of the attention weight; t′ i,j represents the j-th sub-word in the denoised prototype sequence t′ i ; h t is the decoder hidden state at the current time step t, used to generate the translation target word; The updated decoder hidden state, which incorporates the context information obtained from cross-attention; W c is a learnable matrix used to transform the context vector c t after weighted summation; N P represents the number of prototype sentences selected from the candidate set, and L i is the sentence length of the prototype sequence;

[0049] Step4.3. The decoder calculates the next sub-word probability value in an autoregressive manner:

[0050]

[0051] Among them, P v (y t ) is the probability that the decoder predicts the next target word y t based on its own context information. y t represents the target word to be predicted at the current time step t. 1 (·) is the indicator function, L i is the sentence length of the prototype sequence, and λ t is a gating unit composed of a feed-forward network, used to balance the information ratio between the source language sentence and the prototype sequence sentence, and prevent the prototype encoder from overly interfering with the decoding process.

[0052] The present invention also provides a prototype neural machine translation system based on an editing mechanism. The system includes: a module for executing the above-mentioned prototype neural machine translation method based on an editing mechanism.

[0053] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned prototype neural machine translation method based on an editing mechanism.

[0054] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned prototype neural machine translation method based on an editing mechanism is implemented.

[0055] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned prototype neural machine translation method based on an editing mechanism is implemented.

[0056] The beneficial effects of the present invention are as follows:

[0057] 1. The method proposed by the present invention can effectively identify the noise in the prototype sequence and use the edited prototype to guide the translation process, effectively avoiding the interference of irrelevant information on the translation process, thereby improving the translation accuracy.

[0058] 2. When editing the prototype in the present invention, an improved Transformer model is adopted. The input end of the Transformer model is a dual input of cross-domain languages. One part is the sentence that deletes the unaligned words in the prototype, and the other part is the unaligned words in the source sentence. These two parts work together to guide the model to rewrite the prototype sequence and generate a better-aligned prototype sequence, narrowing the semantic gap between the source end and the target end.

[0059] 3. This project adopts a dual-encoder structure, including a source language encoder and a prototype encoder, which respectively encode the prototype sequence and the source language sentence. The prototype encoder can effectively capture the unique features and pattern information in the prototype sequence and integrate them into the translation process, thereby guiding the model to generate a more accurate and fluent target text. The source language encoder is responsible for the overall semantic encoding of the source sentence and provides global information support for the translation process. The two are deeply integrated to realize the comprehensive utilization of the information of the prototype sequence and the source language sentence, providing more accurate guidance for the translation process.

[0060] 4. The present invention can realize the change of the end-to-end neural network structure. The changed network structure can better solve the problem of integrating the noise introduced by the prototype in the process of prototype neural machine translation, and can significantly improve the performance of prototype neural machine translation. The method is robust and applicable to various language translation environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is the data flow diagram of the present invention;

[0062] Figure 2 is the algorithm flow diagram of the present invention;

[0063] Figure 3 is the principle block diagram of the present invention;

[0064] Figure 4It is the model structure diagram proposed by the present invention;

[0065] Figure 5 It is the encoder for the retrieval work of the present invention;

[0066] Figure 6 It is the flow chart of unaligned words in the editing prototype proposed by the present invention;

[0067] Figure 7 It is the cross - language dual - input Transformer model diagram proposed by the present invention. Detailed implementation manners

[0068] Example 1: As Figures 1-7 shown, a prototype neural machine translation method based on an editing mechanism, the method includes:

[0069] Step1, Corpus acquisition: Obtain parallel training corpora, validation corpora, and test corpora for multiple language pairs for model training, parameter tuning, and model effect testing; Use the target sentences of the training corpus as the prototype candidate sentence library;

[0070] The processed parallel corpus is divided into four categories according to the translation direction: Spanish→English, English→Spanish, German→English, English→German; Apply the method of the present invention to different parallel corpora to verify the assumption that the proposed method reduces noise in the prototype; Table 1 shows the experimental corpus information;

[0071] Table 1 Experimental corpus data

[0072]

[0073] Step2, Prototype retrieval model construction: Adopt a cross - language retrieval mechanism and a monolingual memory to retrieve the prototype sequence; Use the Informer encoder to extract semantic features from the source - language sentence and the prototype candidate sentence, and then based on the maximum inner product search (MIPS) algorithm, perform retrieval in the vector space after semantic feature extraction to obtain the prototype sequence t; The prototype sequence does not include the target sentence corresponding to the source - language sentence; The Step2 includes:

[0074] Step2.1, Mark the source - language sentences in the training corpus of Step1 as x, take the source - language sentence x of Step1 as the encoder input, and use the Informer encoder to extract the semantic features of the source - language sentence x;

[0075] First, generate the word - embedding representation E of the input through the embedding layer x , the word - embedding representation E x is expressed as:

[0076]

[0077] Among them, d is the dimension of the embedding vector, and e(x i ) represents the embedding representation of the i-th source language word, and n represents the number of embeddings of source language words;

[0078] Use the Informer encoder to extract the context feature H x , and the context feature H x is expressed as:

[0079]

[0080] The core of the Informer encoder includes the self-attention mechanism and Sparse Factorization. The self-attention score is calculated using ProbSparse Self-Attention and is expressed as:

[0081]

[0082] Among them, Q = W q H x , K = W k H x , V = W v H x are the query matrix, key matrix, and value matrix respectively, and d k is the dimension of the key; ProbSparse(Q, K) filters out highly correlated keys, reduces the computational complexity, and enables the Informer to process long sequences more efficiently;

[0083] Step 2.2. For the set of prototype candidate sentences T = {t 1 , t 2 ,..., t i ,... t k}, input them one by one into the Informer encoder; the embedding and encoding representation of the i-th prototype candidate sentence t i = {t i1 , t i2 ,..., t im} is expressed as:

[0084]

[0085] Among them, m is the number of words in the prototype candidate sentence t i , t i represents the i-th prototype candidate sentence, t im refers to the m-th word in the prototype candidate sentence t i , is the entire word embedding matrix of the prototype candidate sentence t i with the shape of Among them, d is the dimension of the word embedding vector; e(t im ) represents the embedding vector of the m-th word in sentence t i . represents the context feature representation of the prototype candidate sentence t i after being processed by the Informer encoder, and its shape is also

[0086] Step 2.3. Based on Step 2.1 and Step 2.2, use the maximum inner product search technique (MIPS) to quickly retrieve in the set of prototype candidate sentences; MIPS essentially calculates the feature representation H x of the source language sentence and the feature representations i of all target language candidate sentences t to select the target sentence most relevant to the source language sentence; calculate the correlation score f between the source language sentence x and the i-th prototype candidate sentence t i for candidate. The correlation score f(x, t i ) is expressed as:

[0087]

[0088] Among them, t i represents the i-th candidate sentence in the set of prototype candidate sentences T = {t 1 , t 2 ,..., t k}, H x represents the context feature of the source language sentence x after being processed by the Informer encoder;

[0089] According to the high and low of the correlation score, screen out N P prototype sentences most relevant to the source language sentence as the prototype sequence t, N P represents the number of prototype sentences screened from the candidate set, j represents the sentence index screened from the set of prototype candidate sentences T, and the value range of j is 1 ≤ j ≤ N P .

[0090] Step 3. Construction of the prototype noise processing model: Based on Step 2, perform noise reduction processing on the obtained prototype sequence to obtain the denoised prototype sequence. Since the target language sentences in the candidate sentence library may contain some noise information that does not exactly match or is irrelevant to the source language sentence x, if this noise information is directly used for translation, it may affect the translation quality. Therefore, for the noise part in the prototype sequence t, the present invention designs a noise processing model based on a cross-lingual dual-input Transformer model to effectively reduce the noise of the prototype sequence, thereby improving the accuracy and robustness of the prototype neural machine translation model. The denoised prototype is marked as t'; the said Step 3 includes:

[0091] Step 3.1. Call the fast-align tool to perform word alignment on the source language sentence x in Step 1 and the prototype sequence t retrieved in Step 2, and generate an alignment mapping between the source language sentence x and the prototype sequence t;

[0092] Specifically, the fast-align tool obtains the alignment relationship between each word in the source language sentence x and the prototype sequence t by maximizing the alignment probability between word pairs; for each word in the prototype sequence Use a binary indicator function to determine whether it is aligned with any word in the source language sentence; the binary indicator function is expressed as:

[0093]

[0094] Step 3.2. On the basis of Step 3.1, process the prototype sequence t; all words that are not found to be aligned in the source language sentence will be regarded as noise, and through a special placeholder (such as <del>) is replaced; the denoised prototype sequence t′ is represented as:

[0095]

[0096] Step3.3. On the basis of Step3.1, extract the unaligned words in the source language sentence. These words contain important semantic information that may be missing in the prototype sentence. Denote the set of unaligned words in the source language sentence as z, that is:

[0097]

[0098] where represents the i-th word in the source language sentence x;

[0099] By inputting the unaligned words in the source language sentence and the denoised prototype sequence t′ into the Transformer model together, it is used to ensure that the key information of the source language is not missed in the final translation;

[0100] Step3.4. Use the denoised prototype sequence t′ and the set z of unaligned words in the source language to generate the final target sentence Take the data of the prototype sequence t′ and the set z of unaligned words in the source language as inputs and feed them into the Transformer model; different from the traditional Transformer model, the present invention adopts a dual-input structure, taking the prototype sequence t′ and the set z of unaligned words in the source language as two different parts for input; the input result E input is represented as:

[0101] E input = Concat(Embed(t′), Embed(z));

[0102] where Embed represents the word embedding layer and Concat represents the concatenation operation.

[0103] The said Step3.4 further includes: in order to ensure that the Transformer correctly pays attention to the unaligned part of the source language and the masked prototype sentence when processing these two parts of information, by adjusting the self-attention mechanism, avoid letting the placeholder <del>The tag participates in attention calculation. In the self-attention mechanism of the model, an attention mask matrix M is introduced, which ensures that <del>The position information of the tokens is not concerned by the Transformer;

[0104] The attention mask matrix M is at the placeholder <del>The position is negative infinity, and the weight at the negative infinity position is reduced to zero when calculating the attention score; the self-attention calculation formula is:

[0105]

[0106] where Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key, and M is the attention mask matrix; with such settings, <del>The markers do not affect the calculation of the attention mechanism, while the unaligned word z in the source language can fully participate in the attention calculation, thus contributing to the translation generation process.

[0107] Step4, Translation model construction: Based on Step3, the denoised prototype sequence t′ and the source language sentence x will jointly serve as the input of the translation model; in order to better handle the relationship between the source language and the target language, especially in the case of introducing the prototype sequence, the present invention adopts a dual-encoder structure based on Transformer to enhance the information interaction between the source language and the prototype sequence; the dual-encoder structure independently encodes the source language sentence and the denoised prototype sequence respectively and combines the semantic features of the source language sentence and the denoised prototype sequence, thereby improving the translation quality and ensuring that the prototype neural machine translation can still generate high-quality translations when facing noise. The said Step4 includes:

[0108] Step4.1, The source language sentence x of Step1 is fed into the source language encoder, and the source language encoder converts the source language sentence x into a dense vector representation; at the same time, the denoised prototype sequence t′ after Step3 is also fed into the prototype encoder, and the prototype encoder converts each sentence t′ of the denoised prototype sequence i into a set of word embeddings where k represents the kth prototype sentence in the prototype candidate sentence set, and L i is the sentence length of the prototype sequence; semantic feature extraction is performed on the denoised prototype sequence through a neural network layer to obtain the context representation of the denoised prototype sequence.

[0109] Step4.2, The semantic features respectively generated by the source language encoder and the prototype encoder need to be fused through a cross-attention mechanism; the core goal of this stage is to enable the model to weightedly transmit information according to the relevance between the source language and the prototype sequence during the translation process, thereby improving the accuracy and fluency of the translation. Calculate the cross-attention:

[0110]

[0111] where, α i,j calculates the attention of the hidden state h t to the jth sub-word in the ith sentence of the denoised prototype sequence, W m is a dimension transformation matrix, c t is the context vector obtained by weighted summing each word of each sentence t′ of the denoised prototype sequence using α i,j , β is a learnable hyperparameter used to control the influence degree of the prototype sentence correlation score f(x, t i ) in the calculation of the attention weights; t′ i ) i,j Denote the prototype sequence \(t'\) after noise reduction i The \(j\)-th subword in; \(h\) t The decoder hidden state at the current time step \(t\), used to generate the translation target word; The updated decoder hidden state, which incorporates the context information obtained from cross-attention; \(W\) c The learnable matrix used to transform the context vector \(c\) after weighted summation t ; \(N\) P Denote the number of prototype sentences selected from the candidate set, \(L\) i Is the sentence length of the prototype sequence;

[0112] Step 4.3. The decoder calculates the next subword probability value in an autoregressive manner:

[0113]

[0114] Where, \(P\) v (y t ) is the probability that the decoder predicts the next target word \(y\) based on its own context information t , \(y\) t Denotes the target word to be predicted at the current time step \(t\), \(1\) (·) Is the indicator function, \(L\) i Is the sentence length of the prototype sequence, \(\lambda\) t Is a gating unit composed of a feed-forward network, used to balance the information ratio between the source language sentence and the prototype sequence sentence, and prevent the prototype encoder from overly interfering with the decoding process.

[0115] The present invention analyzes the source language sentence through a deep learning model, identifies the irrelevant words that are not aligned in the source language sentence, and ensures that the noise information in the prototype sentence can be accurately located; the present invention performs a masking operation on the identified irrelevant words, shields them with special marking symbols, and uses an end-to-end model based on the self-attention mechanism to generate a complete prototype sequence, thereby ensuring the accuracy and fluency of the translation result.

[0116] The present invention also provides a prototype neural machine translation system based on an editing mechanism, and the system includes:

[0117] A corpus acquisition module, used to obtain parallel training corpora, validation corpora, and test corpora in multiple language pairs for model training, parameter tuning, and model effect testing; randomly select some target sentences from the training corpus as the prototype candidate sentence library;

[0118] A retrieval module, which is used to adopt a cross - language retrieval mechanism and a monolingual memory retrieval prototype sequence; perform semantic feature extraction on the source - language sentence and the prototype candidate sentence through an Informer encoder, and then perform retrieval in the vector space after semantic feature extraction based on the maximum inner - product search algorithm to obtain the prototype sequence;

[0119] A prototype processing module, which is used to construct a prototype noise - processing model: perform noise reduction processing on the obtained prototype sequence to obtain a denoised prototype sequence;

[0120] A translation module (generation module), which is used to use the denoised prototype sequence and the source - language sentence together as the input of the translation model; the dual - encoder structure generates high - quality translations by independently encoding the source - language sentence and the denoised prototype sequence and combining the semantic features of the source - language sentence and the denoised prototype sequence.

[0121] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above - mentioned prototype neural machine translation method based on the editing mechanism is implemented.

[0122] The present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above - mentioned prototype neural machine translation method based on the editing mechanism is implemented.

[0123] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the above - mentioned prototype neural machine translation method based on the editing mechanism is implemented.

[0124] To illustrate the translation effect of the present invention, the translations produced by a baseline system and the present invention are compared. Table 2 shows the improvement in translation quality brought by the model; Table 3 shows the improvement results on different corpora.

[0125] Table 2 shows the translation effect

[0126]

[0127] Table 3 shows the improvement in BLEU values on different corpora

[0128]

[0129] From the above results, it can be seen that the method proposed by the present invention, by processing the noise in the prototype, enables the prototype to better guide the translation process, and the translation quality has been greatly improved. The experimental results on different translation - direction corpora show that the method proposed by the present invention has a greater improvement in translation performance (measured by the BLEU value). Therefore, it is an effective translation method for solving the prototype noise problem.

[0130] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.< / del> < / del> < / del> < / del> < / del>

Claims

1. A prototype neural machine translation method based on editing mechanism, characterized by: The method comprises: Step 1, Corpus acquisition: Obtain parallel training corpus, verification corpus and test corpus of multiple language pairs for model training, parameter tuning and model effect testing; use the target sentences of the training corpus as the prototype candidate sentence library; Step 2, prototype retrieval model construction: adopt cross-language retrieval mechanism and monolingual memory to retrieve prototype sequence; extract semantic features of source language sentences and prototype candidate sentences through Informer encoder, and then search in the vector space after semantic feature extraction based on the maximum inner product search algorithm to obtain the prototype sequence; Step 3, prototype noise processing model construction: Based on Step 2, the obtained prototype sequence is subjected to noise reduction processing to obtain the prototype sequence after noise reduction; Step 4, translation model construction: Based on Step 3, the denoised prototype sequence and the source language sentence will be used as the input of the translation model together; the dual encoder structure generates high-quality translation by independently encoding the source language sentence and the denoised prototype sequence and combining the semantic features of the source language sentence and the denoised prototype sequence.

2. The method for prototypical neural machine translation based on editing mechanism according to claim 1, characterized in that: The Step 2 includes: Step 2.

1. The source language sentence in the training corpus of Step 1 is marked as x. The source language sentence x is used as the encoder input. The semantic features of the source language sentence x are extracted by using the Informer encoder. First, the input word embedding representation E is generated through the embedding layer x , word embedding representation E x It is expressed as: Among them, d is the dimension of the embedding vector, e(x i ) represents the embedding representation of the i-th source language word, and n represents the number of embeddings of source language words; Use Informer encoder to extract context features H x , context feature H x It is expressed as: The core of the Informer encoder includes the self-attention mechanism and sparse factor optimization. ProbSparseSelf-Attention is used to calculate the self-attention score, which is expressed as: Where Q = W q H x , K=W k H x 、V=W v H x are query matrix, key matrix and value matrix respectively, d k is the dimension of the key; Step 2.2, for the prototype candidate sentence set T = {t1, t2, ..., t i ,...t k }, one by one into the Informer encoder; the i-th prototype candidate sentence t i ={t i1 ,t i2 ,…,t im The embedding and encoding of} are expressed as: Where m is the prototype candidate sentence t i The number of words in t i represents the i-th prototype candidate sentence, t im Refers to the prototype candidate sentence t i The mth word in is the prototype candidate sentence t i The entire word embedding matrix has the shape of Where d is the dimension of the word embedding vector; e(t im ) represents sentence t i The embedding vector of the mth word in , Represents the prototype candidate sentence t i The context feature representation after the Informer encoder processing has the same shape Step 2.3, based on Step 2.1 and Step 2.2, use the maximum inner product search technology to quickly search in the prototype candidate sentence set; calculate the source language sentence x and the i-th prototype candidate sentence t i The correlation score f between them, the correlation score f(x,t i ) is expressed as: Among them, t i represents the prototype candidate sentence set T = {t1, t2, ..., t k }, H x Represents the contextual features of the source language sentence x after being processed by the Informer encoder; According to the correlation score, N P The prototype sentences that are most relevant to the source language sentences As the prototype sequence t, N P represents the number of prototype sentences selected from the candidate set, j represents the index of the sentence selected from the prototype candidate sentence set T, and the value range of j is 1≤j≤N P .

3. The prototype neural machine translation method based on editing mechanism according to claim 1, characterized in that: The Step 3 includes: Step 3.1, use the fast-align tool to align the source language sentence x in Step 1 and the prototype sequence t retrieved in Step 2, and generate an alignment mapping between the source language sentence x and the prototype sequence t; Specifically, the fast-align tool obtains the alignment relationship between the source language sentence x and each word in the prototype sequence t by maximizing the alignment probability between word pairs; for each word in the prototype sequence A binary indicator function is used to determine whether it is aligned with any word in the source language sentence; the binary indicator function is expressed as: Step 3.2: Based on Step 3.1, the prototype sequence t is processed; all words that are not aligned in the source language sentence are regarded as noise and replaced by a special placeholder; the prototype sequence t' after denoising is expressed as: Step 3.3, based on Step 3.1, extract the unaligned words in the source language sentence, and record the unaligned word set in the source language sentence as z, that is: in, represents the i-th word in the source language sentence x; By inputting the unaligned words in the source language sentence together with the denoised prototype sequence t' into the Transformer model, it is used to ensure that the key information of the source language is not missed in the final translation; Step 3.4: Generate the final target sentence using the denoised prototype sequence t' and the unaligned word set z in the source language The prototype sequence t' and the data of the unaligned word set z in the source language are input into the Transformer model; a dual input structure is adopted, and the prototype sequence t' and the unaligned word set z in the source language are input as two different parts; the input result E input It is expressed as: E input =Concat(Embed(t'),Embed(z)); Among them, Embed represents the word embedding layer, and Concat represents the concatenation operation.

4. The method of prototypical neural machine translation based on editing mechanism according to claim 3, characterized in that: The Step 3.4 also includes: by adjusting the self-attention mechanism to avoid the placeholder marker from participating in the attention calculation, an attention mask matrix M is introduced; The attention mask matrix M is negative infinity at the placeholder position, and the weight at the negative infinity position is reduced to zero when calculating the attention score; the self-attention calculation formula is: Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key and M is the attention mask matrix.

5. The method of prototypical neural machine translation based on editing mechanism according to claim 1, characterized in that: The Step 4 includes: Step 4.1: The source language sentence x in Step 1 is sent to the source language encoder, which converts the source language sentence x into a dense vector representation. At the same time, the prototype sequence t' after denoising in Step 3 is also sent to the prototype encoder, which converts each sentence t' of the prototype sequence into a dense vector representation. i Convert to a set of word embeddings Where k represents the kth prototype sentence in the prototype candidate sentence set, L i is the sentence length of the prototype sequence; the semantic features of the denoised prototype sequence are extracted through the neural network layer to obtain the context representation of the denoised prototype sequence; Step 4.2: The semantic features generated by the source language encoder and the prototype encoder need to be fused through the cross attention mechanism; calculate the cross attention: Among them, α i,j The hidden state h is calculated t The attention of the jth subword in the ith denoised prototype sequence sentence, W m is a dimension transformation matrix, c t Is to use α i,j For each sentence t' of the denoised prototype sequence i The context vector of each word in the weighted summation, β is a learnable hyperparameter used to control the prototype sentence relevance score f(x,t i ) in the calculation of attention weights; t' i,j Represents the prototype sequence t' after denoising i The jth subword in h t The decoder hidden state at the current time step t is used to generate the translation target word; The updated decoder hidden state incorporates the contextual information obtained by cross attention; W c Used to transform the weighted summed context vector c t The learnable matrix N P represents the number of prototype sentences selected from the candidate set, L i is the sentence length of the prototype sequence; Step 4.3, the decoder calculates the next subword probability value in an autoregressive manner: Among them, P v (y t ) is the decoder predicting the next target word y based on its own context information t The probability of y t Indicates the target word to be predicted at the current time step t, 1 (·) is the indicator function, L i is the sentence length of the prototype sequence, λ t It is a gating unit composed of a feedforward network, which is used to balance the information ratio between the source language sentence and the prototype sequence sentence, and to prevent the prototype encoder from excessively intervening in the decoding process.

6. A prototype neural machine translation system based on editing mechanism, characterized by: The system comprises: a module for executing the prototype neural machine translation method based on editing mechanism as claimed in any one of claims 1 to 5.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the prototype neural machine translation method based on the editing mechanism as described in any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the prototype neural machine translation method based on the editing mechanism as claimed in any one of claims 1 to 5 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the prototype neural machine translation method based on the editing mechanism as claimed in any one of claims 1 to 5 is implemented.