A translation method based on an omni-directional attention mechanism

Optimizing the non-autoregressive translation model through omnidirectional attention mechanism and knowledge distillation technology, multi-modal problems are solved, translation quality and efficiency are improved, and the translation performance of the model in complex scenarios is enhanced.

CN119129611BActive Publication Date: 2025-07-11ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411264320.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-07-11
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

The existing non-autoregressive translation model has multimodal problems when generating translation results, resulting in poor translation quality and efficiency and inability to effectively utilize context information.

Method used

The omnidirectional attention mechanism and knowledge distillation technology are adopted to collect and process parallel corpus data, use autoregressive models to perform knowledge distillation, generate distillation data sets, train non-autoregressive models and introduce omnidirectional attention reasoning modules, gradually adjust the teacher model to the student model, and optimize model performance using comprehensive loss function and course learning strategies.

Benefits of technology

It improves the translation accuracy and fluency of the non-autoregressive translation model, enhances the model's understanding of context, achieves more natural and accurate translation output, and improves the model's learning efficiency and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119129611B_ABST
    Figure CN119129611B_ABST
Patent Text Reader

Abstract

The present invention discloses a translation method based on an omni-directional attention mechanism, which relates to the technical field of natural language processing. The method includes collecting and processing parallel corpus data, generating a distilled dataset through knowledge distillation, training an autoregressive translation model using the distilled dataset and solving the multi-modal problem, converting the autoregressive model into a non-autoregressive model and training it until convergence. By introducing the omni-directional attention mechanism and the curriculum learning strategy, the present invention effectively solves the multi-modal problem in the non-autoregressive translation model, significantly improves the translation quality and training efficiency, and thus achieves a more accurate translation output effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a translation method based on an omni-directional attention mechanism. Background Art

[0002] With the development of information technology, natural language processing, as an important branch in the field of artificial intelligence, has been increasingly studied and applied. Traditional statistical machine translation models are inadequate due to their reliance on complex feature engineering, while neural machine translation based on deep learning has achieved a leap in translation performance through an end-to-end learning approach.

[0003] In the prior art, knowledge distillation is a commonly used method to improve the translation quality of non-autoregressive models by transferring the knowledge of autoregressive models to non-autoregressive models. During the knowledge distillation process, the target sequences generated by the autoregressive model are usually regarded as "soft labels", while the target sequences generated by the non-autoregressive model are regarded as "hard labels", which results in the non-autoregressive model being unable to generate high-quality translation results in actual situations. Therefore, how to use an omni-directional attention inference module to adjust the teacher model through a curriculum learning strategy to solve the multi-modal problem has become the key to improving the performance of non-autoregressive translation models. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a translation method based on an omni-directional attention mechanism, which solves the multi-modal problem of autoregressive translation models and improves the translation accuracy and speed of non-autoregressive models.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, an embodiment of the present invention provides a translation method based on an omni-directional attention mechanism, which includes collecting and processing parallel corpus data;

[0008] Performing knowledge distillation using the processed parallel corpus data and an autoregressive translation model with a parameter scale comparable to that of the non-autoregressive translation model to generate a distilled dataset;

[0009] Training the autoregressive translation model using the distilled dataset to identify existing multi-modal problems;

[0010] Proposing improvement ideas and solutions based on the existing multi-modal problems;

[0011] After proposing the improvement ideas and solutions, converting the target language generated by the autoregressive translation model into the target language generated by the non-autoregressive translation model;

[0012] After completing the conversion from the autoregressive translation model to the non-autoregressive translation model, use the distilled dataset to train the non-autoregressive translation model until convergence;

[0013] After the non-autoregressive translation model converges, evaluate and optimize the non-autoregressive translation model.

[0014] As a preferred solution of the translation method based on the omni-directional attention mechanism of the present invention, wherein: collecting and processing parallel corpus data includes the following steps,

[0015] Collect parallel corpus of the source language and the target language that have been correctly translated and strictly aligned;

[0016] Clean the collected parallel corpus, unify the case of sentences and standardize the punctuation marks, remove unnecessary tags and special characters, and at the same time filter out too long and too short sentences and sentences with serious errors;

[0017] Construct a vocabulary based on the cleaned parallel corpus and tokenize the sentences in the vocabulary.

[0018] As a preferred solution of the translation method based on the omni-directional attention mechanism of the present invention, wherein: use the processed parallel corpus data and an autoregressive translation model with a parameter scale equivalent to that of the non-autoregressive translation model for knowledge distillation to generate a distilled dataset, including the following steps,

[0019] Select an autoregressive model with excellent performance as the teacher model;

[0020] Use the teacher model to perform forward propagation and knowledge distillation on the processed parallel corpus data to generate a character sequence of the target sentence, and obtain a pure distilled dataset.

[0021] As a preferred solution of the translation method based on the omni-directional attention mechanism of the present invention, wherein: use the distilled dataset to train the autoregressive translation model to find the existing multi-modal problems, including the following steps,

[0022] After generating the distilled dataset, use the generated distilled dataset to train the student model and learn the hard labels of the original training data at the same time, so as to find the multi-modal problems;

[0023] Define a comprehensive loss function according to the multi-modal problems. One part is the prediction error of the student model for the original label, and the other part is the gap between the output of the student model and the soft label of the teacher model. The expression of the comprehensive loss function of the student model is:

[0024] ;

[0025] Wherein, represents the comprehensive loss function, represents the loss term on the original training dataset, and respectively represent the and weight coefficients for balancing the impacts of different loss terms, represents the loss term on the distilled dataset, represents the loss term during the conversion from the autoregressive model to the non - autoregressive model.

[0026] As a preferred solution of the translation method based on the omnidirectional attention mechanism according to the present invention, wherein: according to the existing multi - mode problem, improvement ideas and solutions are proposed, including the following steps,

[0027] According to the existing multi - mode problem, an omnidirectional attention inference module is introduced to enable the character sequence generation to observe the information at all positions;

[0028] The teacher model is gradually adjusted through the curriculum learning strategy to make the teacher model smoothly transition to the student model.

[0029] As a preferred solution of the translation method based on the omnidirectional attention mechanism according to the present invention, wherein: after the improvement ideas and solutions are proposed, the target language generated by the autoregressive translation model is converted into the target language generated by the non - autoregressive translation model, including the following steps,

[0030] At the beginning of the conversion, the information of the target language generated by the autoregressive translation model should be consistent with the information of the target language of the previously generated character sequence;

[0031] In each teacher model training stage, the substitution rate is calculated to evaluate the performance of the non - autoregressive translation model;

[0032] The expression for calculating the substitution rate of the teacher model training is:

[0033] ;

[0034] Wherein, represents the substitution rate, represents the total number of training rounds, represents the number of sentences participating in the calculation of entropy and hash function values, represents the sentence length of the dataset, represents the th sentence length for information entropy calculation, represents the hash function value corresponding to the sentence length, represents the evaluation value of the sentence length by the normal distribution density function with the mean and variance of the sentence length as parameters, represents the total number of sentences in the distilled dataset, Denotes the length of the th sentence in the distilled dataset, Denotes the function value of the length of the th distilled data sentence length ; Denotes the logarithmic function value of the sentence length, Denotes the small change in the length of each sentence in the distilled dataset.

[0035] As a preferred solution of the translation method based on the omni-directional attention mechanism described in the present invention, wherein: after completing the conversion from the autoregressive translation model to the non-autoregressive translation model, the non-autoregressive translation model is trained using the distilled dataset until convergence, including the following steps,

[0036] Before training the non-autoregressive translation model using the distilled dataset, select the neural network architecture for the non-autoregressive task, and perform random initialization and pre-training initialization on the non-autoregressive model using the neural network architecture for the non-autoregressive task;

[0037] During the training process, the non-autoregressive model uses cross-entropy loss to minimize the difference between the predicted target sequence and the target sequence provided by the teacher model in the distilled dataset;

[0038] When generating the final output, adopt the decoding strategy of the post-processing technique to remove the difference between the predicted target sequence and the target sequence provided by the teacher model in the non-autoregressive model, improve the translation quality, and at the same time, an auxiliary loss of consistency loss needs to be added to improve the performance of the non-autoregressive model, determine that for different noisy inputs, the non-autoregressive model can also produce the same output, and then complete the convergence of the autoregressive translation model.

[0039] As a preferred solution of the translation method based on the omni-directional attention mechanism described in the present invention, wherein: evaluating and optimizing the non-autoregressive translation model includes the following steps,

[0040] Evaluate and test the performance of the non-autoregressive translation model on the distilled dataset;

[0041] Adjust the hyperparameters of the learning rate and batch size of the non-autoregressive translation model according to the evaluation performance of the distilled dataset, and use model fusion with different initializations;

[0042] After improving the performance of the non-autoregressive translation model, export the model into an easily deployable format, and use quantization and pruning techniques to further improve the inference speed and reduce the memory occupation to optimize the model;

[0043] After optimizing the non-autoregressive translation model, a real-time monitoring system is established to monitor the performance metrics of the response time and throughput of the non-autoregressive translation model in the production environment. Meanwhile, an anomaly detection mechanism is set up to promptly detect and handle specific problems encountered by the model in practical applications.

[0044] According to the feedback in practical applications, the non-autoregressive translation model is retrained regularly and irregularly to adapt to the constantly changing data distribution.

[0045] In a second aspect, an embodiment of the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the translation method based on the omni-directional attention mechanism as described in the first aspect of the present invention is implemented.

[0046] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by the processor, any step of the translation method based on the omni-directional attention mechanism as described in the first aspect of the present invention is implemented.

[0047] The beneficial effects of the present invention are as follows: By collecting and processing parallel corpus data, the preparation of high-quality training data is achieved, improving the accuracy and fluency of translation. Using the autoregressive model to perform knowledge distillation on the processed corpus ensures that the non-autoregressive model can obtain high-quality knowledge transfer, thereby improving the learning efficiency and generalization ability of the translation model. Training the autoregressive model with the distilled dataset and identifying multi-modal problems realizes problem localization and improvement during the model training process, ensuring that the model can maintain high translation quality in complex scenarios. Introducing the omni-directional attention mechanism and gradually adjusting the teacher model to the student model solves the multi-modal problem, enhances the model's ability to understand context during translation, and achieves a more natural and accurate translation output effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a flowchart of the translation method based on the omni-directional attention mechanism in Embodiment 1.

[0050] Figure 2 It is a decision diagram of the translation method based on the omni-directional attention mechanism in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.

[0052] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0053] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that mutually excludes other embodiments.

[0054] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides a translation method based on an omni-directional attention mechanism, including the following steps:

[0055] S1. Collect and process parallel corpus data, including the following steps.

[0056] Collect parallel corpus of the source language and the target language that has been correctly translated and strictly aligned.

[0057] Parallel corpus refers to a dataset containing corresponding texts in two or more languages. The corresponding texts are matched with each other in different languages. One sentence in an article has an accurate translation in different language versions. Since parallel corpus is carefully translated and proofread by professional translators, the translation quality is relatively high. Usually, parallel corpus is used as the benchmark for training a high-quality machine translation system.

[0058] Clean the collected parallel corpus, unify the case of sentences and standardize the punctuation marks, remove unnecessary tags and special characters, and at the same time filter out too long and too short sentences and sentences with serious errors.

[0059] Construct a vocabulary based on the cleaned parallel corpus and perform word segmentation on the sentences in the vocabulary.

[0060] A vocabulary refers to an ordered set of words used to map words in natural language to numerical representations so that a machine learning model can process the words in natural language and convert the words in the text into corresponding index numbers. For example, the source language sentence "Hello, how are you?" can be encoded as [2, 6, 3, 4, 5, 7].

[0061] S2. Use the processed parallel corpus data to perform knowledge distillation with an autoregressive translation model having a parameter scale comparable to that of the non-autoregressive translation model to generate a distilled dataset, including the following steps.

[0062] Select an autoregressive model with excellent performance as the teacher model.

[0063] Use the teacher model to perform forward propagation and knowledge distillation on the processed parallel corpus data to generate the character sequence of the target sentence and obtain a pure distilled dataset.

[0064] Forward propagation is a basic operation process in a neural network, which describes how the input data passes through the various layers of the network and finally produces the output. In a neural network, the output of each layer is a function of the output of the previous layer, and this process continues until the final output layer produces the final prediction result. It is also the basis for making predictions using a trained model.

[0065] Knowledge distillation technology is a model compression method, the purpose of which is to transfer the knowledge of the teacher model to the student model. This technology is particularly popular in the field of deep learning because it can significantly reduce the size and computational cost of the model while maintaining the model performance. In the machine translation task, a complex autoregressive model is used as the teacher model, and a non-autoregressive model is trained as the student model through knowledge distillation technology to improve the translation speed while maintaining the translation quality. In the image recognition task, a large convolutional neural network is used as the teacher model, and a MobileNet network is trained through knowledge distillation technology to achieve higher computational efficiency.

[0066] S3. Use the distilled dataset to train the autoregressive translation model and find the existing multi-modal problems, including the following steps.

[0067] After generating the distilled dataset, use the generated distilled dataset to train the student model and simultaneously learn the hard labels of the original training data to find the multi-modal problems.

[0068] The multi-modal problem refers to the situation in the machine translation task where, given a source language sentence, there are multiple different target language translations. For example, translating "Thank you" into "Danke Dank" is unreasonable in actual language. This problem is particularly prominent in non-autoregressive translation models because non-autoregressive translation models lack context dependence when generating the target sequence, resulting in poor diversity and coherence of the generated results. For example, it may translate "I like to eat apples" into "Ich mag Äpfel essen zu" instead of the more reasonable "Ich mag Äpfel essen".

[0069] Define a comprehensive loss function based on the multi-modal problem. One part is the prediction error of the student model for the original labels, and the other part is the gap between the output of the student model and the soft labels of the teacher model. The expression of the comprehensive loss function of the student model is as follows:

[0070] ;

[0071] Among them, represents the comprehensive loss function, represents the loss term on the original training dataset, and respectively represent the and weight coefficients that balance the influence of different loss terms, represents the loss term on the distillation dataset, represents the loss term during the conversion from the autoregressive model to the non-autoregressive model.

[0072] S4. According to the existing multi-modal problem, propose improvement ideas and solutions, including the following steps.

[0073] According to the existing multi-modal problem, introduce an omni-directional attention inference module to enable the character sequence generation to observe the information at all positions;

[0074] The omni-directional attention inference module is an improved attention mechanism. To solve the multi-modal problem and the insufficient context dependence faced by the non-autoregressive model during sequence generation, the omni-directional attention inference module enables the non-autoregressive model to access the information at all positions when generating each target word, thereby better utilizing the context information and improving the coherence and accuracy of the generated sequence;

[0075] Gradually adjust the teacher model through the curriculum learning strategy to enable the teacher model to smoothly transition to the student model;

[0076] Curriculum learning is a training strategy that gradually adjusts the difficulty or conditions during the training process, enabling the model to better learn and adapt to the task. During the knowledge distillation process, the curriculum learning strategy is used to gradually adjust the guiding method of the teacher model for the student model, thereby helping the student model better learn the knowledge of the teacher model. At each adjustment stage, regularly evaluate the performance of the student model to ensure that the student model can gradually learn to independently generate high-quality outputs. If the performance of the student model is poor at a certain stage, the decay rate of the substitution rate and the final substitution rate value can be appropriately adjusted.

[0077] After proposing the improvement ideas and solutions, convert the target language generated by the autoregressive translation model into the target language generated by the non-autoregressive translation model, including the following steps.

[0078] When starting the conversion, the information in the target language generated by the autoregressive translation model should be consistent with the information in the target language of the previously generated character sequence;

[0079] In each teacher model training stage, the substitution rate is calculated to evaluate the performance of the non-autoregressive translation model;

[0080] The substitution rate is a concept used in the curriculum learning strategy, which refers to the proportion of using the output of the teacher model to replace the output generated by the student model itself when training the student model. The specific value ranges from 0 to 1 to control the degree of dependence of the student model on the teacher model during training. When the substitution rate is 1, it means that the student model completely depends on the output of the teacher model; when the substitution rate is 0, it means that the student model generates the output completely independently;

[0081] The expression for calculating the substitution rate of teacher model training is:

[0082] ;

[0083] where, represents the substitution rate, represents the total number of training rounds, represents the number of sentences participating in the calculation of entropy and hash function values, represents the sentence length of the dataset, represents the calculation of the information entropy of the th sentence length , represents the hash function value corresponding to the sentence length, represents the evaluation value of the sentence length by the normal distribution density function with the mean and variance of the sentence length as parameters, represents the total number of sentences in the distilled dataset, represents the th sentence length in the distilled dataset, represents the calculation of the th distilled data sentence length , represents the logarithmic function value of the sentence length, represents the small change amount of each sentence length in the distilled dataset.

[0084] S6. After completing the conversion from the autoregressive translation model to the non-autoregressive translation model, use the distilled dataset to train the non-autoregressive translation model until convergence, including the following steps,

[0085] Before training a non - autoregressive translation model using a distilled dataset, select a neural network architecture for the non - autoregressive task and perform random initialization and pre - training initialization on the non - autoregressive model using the neural network architecture for the non - autoregressive task;

[0086] Random initialization and pre - training initialization are two commonly used initialization methods in the training process of deep learning models. They are respectively applicable to different scenarios and have their own advantages and disadvantages;

[0087] Random initialization means that before the start of model training, the weights and biases of the model are randomly assigned. The advantage of random initialization is that it is very flexible, applicable to any model architecture, and does not require additional data and pre - trained models. The disadvantage is that it leads to a slower convergence rate in the initial stage of training in deep networks;

[0088] Pre - training initialization means that before the start of model training, the weights of the model are initialized to the weights of another pre - trained model. The advantage of pre - training initialization can help the model converge faster, has strong generalization ability, and helps to reduce the over - fitting phenomenon;

[0089] During the training process, the non - autoregressive model uses cross - entropy loss to minimize the difference between the predicted target sequence and the target sequence provided by the teacher model in the distilled dataset;

[0090] Cross - entropy loss refers to measuring the difference between two probability distributions in a multi - classification problem, enabling the model to better fit the training data. When the probability distribution predicted by the model is consistent with the true label distribution, the cross - entropy loss is minimized. During the training process, by minimizing the cross - entropy loss, the model can learn a prediction probability distribution closer to the true label. When the model makes a wrong prediction, the cross - entropy loss will increase significantly, thus prompting the model to adjust its parameters to reduce this error.

[0091] When generating the final output, adopt a decoding strategy of post - processing technology to remove the difference between the minimized predicted target sequence and the target sequence provided by the teacher model in the distilled dataset in the non - autoregressive model, improve the translation quality. At the same time, an auxiliary loss of consistency loss needs to be added to enhance the performance of the non - autoregressive model, ensuring that for different noisy inputs, the non - autoregressive model can also produce the same output, and then complete the convergence of the autoregressive translation model.

[0092] The decoding strategy of post - processing technology refers to the process of further optimizing the generated result through a series of techniques and methods after generating the output of the machine translation model. By correcting the errors and unreasonable parts in the generated result, ensuring that the generated sentence is more coherent and natural in grammar and semantics.

[0093] S7. Evaluate and optimize the non - autoregressive translation model, including the following steps,

[0094] Evaluate and test the performance of the non-autoregressive translation model on the distilled dataset;

[0095] Adjust the hyperparameters of the learning rate and batch size of the non-autoregressive translation model according to the evaluation performance of the distilled dataset, and use model fusion with different initializations;

[0096] Hyperparameters are parameters preset in machine learning and deep learning models, manually set by developers and researchers. Selecting appropriate hyperparameters can significantly improve the model's performance, while poor hyperparameter settings can lead to poor model performance;

[0097] The hyperparameter of the learning rate determines the magnitude of the update of the model parameters in each iteration. A higher learning rate will make the training process faster, but will lead to unstable convergence of the model. A lower learning rate will make the model converge more stably, but the training process is slower. Common learning rate adjustment strategies include fixed learning rate, exponential decay, and adaptive learning rate in the Adam optimizer;

[0098] The hyperparameter of the batch size determines the number of samples used when updating the model parameters. A larger batch size will make each gradient update closer to the true gradient, but requires more memory. A smaller batch size will make the training process more flexible, but will lead to fluctuations in gradient estimation. Common batch sizes are 32, 64, and 128;

[0099] After improving the performance of the non-autoregressive translation model, export the model into an easily deployable format, and use quantization and pruning techniques to further improve the inference speed and reduce memory occupancy to optimize the model;

[0100] Quantization refers to the process of converting the weights and activation values in the model from 32-bit floating-point numbers to 8-bit integers. Through quantization, the storage space requirements and computational overhead of the model can be significantly reduced;

[0101] Pruning refers to the process of reducing the number of parameters and computational complexity of the model by removing unimportant weights or channels in the model. Through pruning, the memory bandwidth required by the model during operation can be reduced, thereby improving computational efficiency;

[0102] After optimizing the non-autoregressive translation model, establish a real-time monitoring system to monitor the performance metrics of the response time and throughput of the non-autoregressive translation model in the production environment, and set up an anomaly detection mechanism to timely discover and handle specific problems encountered by the model in actual applications;

[0103] According to the feedback in actual applications, retrain the non-autoregressive translation model regularly and irregularly to adapt to the changing data distribution.

[0104] This embodiment also provides a computer device, which is applicable to the situation of the translation method based on the omnidirectional attention mechanism, and includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the translation method based on the omnidirectional attention mechanism proposed in the above embodiment.

[0105] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0106] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the translation method based on the omnidirectional attention mechanism proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0107] In summary, the present invention realizes the preparation of high-quality training data by collecting and processing parallel corpus data, improves the accuracy and fluency of translation, uses an autoregressive model to perform knowledge distillation on the processed corpus to ensure that the non-autoregressive model can obtain high-quality knowledge transfer, thereby improving the learning efficiency and generalization ability of the translation model, trains the autoregressive model using the distilled dataset and identifies multi-modal problems, realizes problem positioning and improvement during the model training process, ensures that the model can maintain high translation quality in complex scenarios, introduces an omni-directional attention mechanism and gradually adjusts the teacher model to the student model, solves the multi-modal problem, enhances the model's ability to understand context during translation, and realizes a more natural and accurate translation output effect.

[0108] Example 2. Referring to Table 1, this is the second example of the present invention. To further verify the technical solution of the present invention, experimental simulation data of a translation method based on an omni-directional attention mechanism is given.

[0109] Experimental Preparation and Implementation Process

[0110] To verify the effectiveness of the translation method based on the omni-directional attention mechanism, the following experimental steps were carried out.

[0111] Data Collection and Processing:

[0112] Collect a parallel corpus of English-Chinese with 100,000 correctly translated and strictly aligned sentences.

[0113] Clean the collected parallel corpus, including unifying case, normalizing punctuation, removing unnecessary tags and special characters, and filtering out sentences that are too long (more than 100 words) and too short (less than 5 words) as well as sentences with serious errors.

[0114] Construct a vocabulary and tokenize the sentences in the vocabulary.

[0115] Knowledge Distillation Preparation:

[0116] Select an autoregressive model that performs well in English-Chinese translation tasks as the teacher model.

[0117] Use the teacher model to perform forward propagation and knowledge distillation on the processed parallel corpus data to generate the character sequence of the target sentence and obtain a pure distilled dataset.

[0118] Multi-modal Problem Identification and Improvement:

[0119] Train a student model using the distilled dataset, and the parameter scale of this model is comparable to that of the non-autoregressive translation model.

[0120] During the training process, a comprehensive loss function is defined, which includes a loss term on the original training dataset, a loss term on the distilled dataset, and a loss term during the conversion from the autoregressive model to the non-autoregressive model.

[0121] By introducing an omni-directional attention inference module, all position information can be observed when generating character sequences.

[0122] The teacher model is gradually adjusted through a curriculum learning strategy, enabling a smooth transition from the teacher model to the student model.

[0123] Conversion and Training:

[0124] During each teacher model training stage, a substitution rate is calculated to evaluate the performance of the non-autoregressive translation model.

[0125] After completing the conversion from the autoregressive translation model to the non-autoregressive translation model, the non-autoregressive translation model is trained on the distilled dataset until convergence.

[0126] Select the neural network architecture for the non-autoregressive task, and randomly initialize and pre-train initialize the non-autoregressive model using the neural network architecture for the non-autoregressive task.

[0127] During the training process, the non-autoregressive model uses cross-entropy loss to minimize the difference between the predicted target sequence and the target sequence provided by the teacher model in the distilled dataset.

[0128] Adopt a decoding strategy with post-processing techniques to remove the difference between the minimized predicted target sequence and the target sequence provided by the teacher model in the distilled dataset in the non-autoregressive model, improve the translation quality, and add an auxiliary loss of consistency loss to enhance the performance of the non-autoregressive model.

[0129] Evaluation and Optimization:

[0130] Evaluate and test the performance of the non-autoregressive translation model on the distilled dataset.

[0131] Adjust the hyperparameters of the learning rate and batch size of the non-autoregressive translation model according to the evaluation performance, and use model fusion with different initializations.

[0132] After improving the performance of the non-autoregressive translation model, export the model into an easily deployable format, and use quantization and pruning techniques to further improve the inference speed and reduce the memory footprint.

[0133] Establish a real-time monitoring system to monitor the performance metrics of the response time and throughput of the non-autoregressive translation model in the production environment, and set up an anomaly detection mechanism to timely detect and handle specific problems encountered by the model in actual applications.

[0134] According to the feedback in actual applications, the non-autoregressive translation model is retrained regularly and irregularly to adapt to the changing data distribution.

[0135] Specifically, it is shown in Table 1 below:

[0136] Table 1 Performance Comparison Table of Translation Methods Based on Omnidirectional Attention Mechanism

[0137]

[0138] From the data in the above table, it can be seen that the method of the present invention has significant advantages compared with the traditional neural machine translation method:

[0139] BLEU Score (Bilingual Evaluation Understudy): The BLEU score of the method of the present invention before optimization is 28.5, showing a significant improvement compared with 25.3 of the traditional NMT. The BLEU score after optimization reaches 32.1, indicating that the method has made significant improvements in translation quality.

[0140] Perplexity: The perplexity of the method of the present invention before optimization is 9.8, lower than 11.4 of the traditional NMT. The perplexity after optimization is further reduced to 8.5, indicating that the model is more accurate in predicting the next word and the overall translation quality is higher.

[0141] Training Time: The training time of the method of the present invention before optimization is 9 hours, which is shorter than 10 hours of the traditional NMT. The training time after optimization is further shortened to 7 hours, indicating that the training efficiency of the model has been significantly improved.

[0142] Memory Usage: The memory usage of the method of the present invention before optimization is 1400MB, lower than 1500MB of the traditional NMT. The memory usage after optimization is further reduced to 1000MB, indicating that the model is more resource-saving and more suitable for deployment in resource-constrained environments.

[0143] In summary, by introducing the omnidirectional attention mechanism, knowledge distillation technology and curriculum learning strategy, the method of the present invention effectively solves the multi-modal problem existing in the traditional non-autoregressive translation model, improves the translation quality and training efficiency, and makes the non-autoregressive translation model more efficient and reliable in actual deployment.

[0144] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A translation method based on an omni-directional attention mechanism, characterized in that: including, collecting and processing parallel corpus data; using the processed parallel corpus data and an autoregressive translation model with a parameter scale comparable to that of the non-autoregressive translation model as a teacher model for knowledge distillation to generate a distilled dataset; using the distilled dataset to train the non-autoregressive translation model as a student model to find existing multi-modal problems; according to the existing multi-modal problems, introducing an omni-directional attention inference module into the student model of the non-autoregressive translation model so that all position information can be observed when generating the character sequence; gradually adjusting the guidance method of the teacher model to the student model through a curriculum learning strategy to help the student model learn the knowledge of the teacher model and complete the conversion from the teacher model of the autoregressive translation model to the student model of the non-autoregressive translation model, specifically including the following steps at the beginning of the conversion, the information of the target language generated by the teacher model of the autoregressive translation model is consistent with the information of the target language of the character sequence generated during knowledge distillation; in each conversion training stage using the teacher model, calculate the substitution rate to evaluate the performance of the student model of the non-autoregressive translation model. The substitution rate refers to the proportion of using the output of the teacher model to replace the output generated by the student model itself when training the student model. The specific value ranges from 0 to 1 to control the degree of dependence of the student model on the teacher model during training. When the substitution rate is 1, it means that the student model completely depends on the output of the teacher model; when the substitution rate is 0, it means that the student model generates the output completely independently; after completing the conversion from the teacher model of the autoregressive translation model to the student model of the non-autoregressive translation model, use the distilled dataset to train the student model of the non-autoregressive translation model until convergence; evaluate and optimize the student model of the non-autoregressive translation model; wherein, using the distilled dataset to train the non-autoregressive translation model as a student model to find existing multi-modal problems, including the following steps after generating the distilled dataset, use the generated distilled dataset to train the student model while learning the hard labels of the original training data to find multi-modal problems; define a comprehensive loss function according to the multi-modal problems. One part is the prediction error of the student model for the original labels, and the other part is the gap between the output of the student model and the soft labels of the teacher model. The expression of the comprehensive loss function of the student model is: ; Among them, represents the comprehensive loss function, represents the loss term on the original training dataset, and respectively represent the and weight coefficients for balancing the influence of different loss terms, represents the loss term on the distilled dataset, represents the loss term during the conversion from the autoregressive translation model to the non - autoregressive translation model.

2. The translation method based on the omnidirectional attention mechanism according to claim 1, characterized in that collecting and processing parallel corpus data including the following steps collecting parallel corpus of the source language and the target language that have been correctly translated and aligned; cleaning the collected parallel corpus, unifying the case of sentences and standardizing the punctuation, removing unnecessary tags and special characters, and filtering sentences at the same time; constructing a vocabulary based on the cleaned parallel corpus and tokenizing the sentences in the vocabulary.

3. The translation method based on the omnidirectional attention mechanism according to claim 2, wherein, using the processed parallel corpus data and an autoregressive translation model with a parameter scale comparable to that of the non-autoregressive translation model as a teacher model for knowledge distillation to generate a distilled dataset including the following steps using the teacher model to perform forward propagation and knowledge distillation on the processed parallel corpus data to generate the character sequence of the target sentence and obtain a pure distilled dataset.

4. The translation method based on the omnidirectional attention mechanism according to claim 3, wherein, After completing the conversion from the teacher model of the autoregressive translation model to the student model of the non-autoregressive translation model, the student model of the non-autoregressive translation model is trained using the distilled dataset until convergence, including the following steps, During the training process, the student model of the non-autoregressive translation model uses cross-entropy loss to minimize the difference between the predicted target sequence and the target sequence provided by the teacher model in the distilled dataset.

5. The translation method based on the omnidirectional attention mechanism according to claim 4, characterized in that Evaluate and optimize the student model of the non-autoregressive translation model, including the following steps, Evaluate and test the performance of the student model of the non-autoregressive translation model on the distilled dataset; Adjust the hyperparameters of the learning rate and batch size of the student model of the non-autoregressive translation model according to the evaluation performance on the distilled dataset; After improving the performance of the student model of the non-autoregressive translation model, export the model and use quantization and pruning techniques to further improve the inference speed and reduce the memory footprint to optimize the model; After completing the optimization of the student model of the non-autoregressive translation model, establish a real-time monitoring system to monitor the performance metrics of the response time and throughput of the student model of the non-autoregressive translation model in the production environment, and at the same time set up an anomaly detection mechanism to timely detect and handle specific problems encountered by the model in practical applications; According to the feedback in practical applications, retrain the student model of the non-autoregressive translation model regularly and irregularly to adapt to the changing data distribution.

6. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the translation method based on the omnidirectional attention mechanism according to any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the translation method based on the omnidirectional attention mechanism according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Non-autoregressive Mongolian-Chinese machine translation method based on round-robin decoding and vocabulary attention

    CN112417901A

  • Network learning resource analysis and personalized recommendation method based on knowledge graph

    CN114861069A