A bilingual mutual translation method and system based on deep learning and neural network technology
Patent Information
- Application Number
- CN202610955250.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-10-09
AI Technical Summary
[0005]然而,尽管这些技术取得了显著进步,现有模型在双语互译任务上仍面临重大挑战
[0034]1.本发明通过BL-LMNMT模型实现了双向语言翻译功能,无需为不同语言方向分别训练和部署独立模型,大幅简化了系统架构,降低了维护复杂度。
Smart Images

Figure CN122886629A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine translation technology in natural language processing (NLP), and more particularly to a bilingual translation method and system based on deep learning and neural network technology. Background Technology
[0002] Machine translation, as an important subfield of Natural Language Processing (NLP), has long been dedicated to achieving automatic text conversion between different languages. Before the rise of deep learning technology, machine translation mainly relied on rule-based methods and statistical machine translation (SMT) models. However, these traditional methods have significant limitations when dealing with complex language structures and long-range dependencies.
[0003] With the successful application of neural network technology in machine translation, neural machine translation (NMT) has begun to emerge and is gradually becoming a mainstream technology. Through end-to-end training, NMT can better capture the semantic and syntactic features of language, significantly improving translation quality. In particular, Google's Transformer model, which introduces a multi-head self-attention mechanism, has completely changed the way sequence modeling is done, greatly improving the efficiency and effectiveness of machine translation.
[0004] Subsequently, the rise of large-scale pre-trained models revolutionized the field of NLP. Models such as BERT, GPT, ALBERT, and ROBERTa, through pre-training on massive amounts of text, learned rich linguistic knowledge, providing a strong foundation for downstream tasks. The UNILM model proposed by Microsoft Research Asia further integrates multiple pre-training tasks (including GPT, BERT, and sequence-to-sequence tasks), providing excellent fine-tuning capabilities for various NLP tasks, including machine translation.
[0005] However, despite these significant advancements, existing models still face major challenges in bilingual translation tasks. Most pre-trained models are inherently biased towards monolingual processing, performing exceptionally well on single-language tasks but struggling to adapt to cross-language scenarios. Particularly in bidirectional translation tasks, traditional methods typically require training separate models for each language direction (e.g., English-to-Chinese and Chinese-to-English models), leading to high resource consumption, high maintenance costs, and difficulties in ensuring model consistency.
[0006] More importantly, existing technologies struggle to effectively address the problem of automatic source language recognition. In practical applications, systems often need to first determine the language of the input text before calling the corresponding translation model. This process not only increases system complexity but also introduces additional error propagation risks. Furthermore, existing models perform poorly when handling multilingual input and cannot efficiently complete multi-directional translation tasks within a unified architecture.
[0007] Therefore, there is an urgent need for a unified model architecture that can automatically identify the source language, support bidirectional translation, and has superior performance, in order to overcome the limitations of existing technologies and improve the efficiency and practicality of machine translation systems. Summary of the Invention
[0008] This invention addresses the shortcomings of existing technologies by providing a bilingual translation method and system based on deep learning and neural network technologies.
[0009] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows:
[0010] A bilingual translation method based on deep learning and neural network technology includes the following steps:
[0011] A bilingual translation model BL-LMNMT is constructed. This model adopts a joint encoder-decoder structure based on the UNILM sequence-to-sequence model. The input is a combination of source language sequence and target language sequence connected by the special symbols [CLS] and [SEP].
[0012] The BL-LMNMT model is designed with a multi-task learning mechanism, including a source language category recognition subtask and a machine translation subtask. The source language category recognition subtask is implemented using BERT's NSP task, and the machine translation subtask is completed using a restricted multi-head self-attention network.
[0013] The BL-LMNMT model is trained by updating the model parameters by minimizing the total loss function, which is a weighted sum of the loss functions of the two sub-tasks: total loss = w1×l1 + w2×l2, where w1 and w2 are weight coefficients, l1 is the loss function of the source language category recognition sub-task, and l2 is the loss function of the machine translation sub-task.
[0014] The trained BL-LMNMT model is used for bilingual translation, automatically identifying the source language and outputting the corresponding target language.
[0015] Furthermore, in the data preprocessing stage, the handling of uncertainties in the source language includes:
[0016] Construct a dataset from source language to target language: For each sample, process it using the format "[CLS]Source Language Identifier[SEP]Target Language Identifier[SEP]", and randomly shuffle the order of all samples;
[0017] Construct a target language to source language dataset: For each sample, process it using the format "[CLS]Target Language Identifier[SEP]Source Language Identifier[SEP]", and randomly shuffle the order of all samples;
[0018] The two subsets are merged, and the overall sample order is randomly shuffled again.
[0019] Furthermore, the weight coefficient combination in the total loss function is set to w1:w2=1:100 to balance the difference in importance between the source language category recognition subtask and the machine translation subtask.
[0020] Furthermore, in the source language category recognition subtask, the model performs binary classification through the output at the [CLS] position, and the output probability distribution [p1, p2] satisfies p1 + p2 = 1, where p1 represents the source language probability and p2 represents the target language probability; by adjusting the model parameters, the KL divergence between the output probability distribution and the true probability distribution [1,0] or [0,1] is minimized.
[0021] Furthermore, in the inference stage, a beam search algorithm is used to generate translation results. The beam size is set to 4. By calculating the probability product of all possible target sequences, the sequence with the largest probability product is selected as the final translation result.
[0022] Furthermore, during model training, a minimum loss threshold is set to 1e-5. When the model's loss value on the validation set is less than or equal to this threshold, training is stopped and the current model parameters are saved.
[0023] Furthermore, the LAMB optimizer is used for model training, supporting large-batch training with a batch size of up to 32,000 without loss of accuracy, thereby accelerating the training speed.
[0024] Furthermore, a multilingual BERT pre-trained model is used as the initialization basis for the BL-LMNMT model, transferring the pre-trained knowledge to the bilingual translation task to improve model performance.
[0025] Furthermore, parameter sharing and matrix factorization techniques are employed in the embedding layer to decompose the one-hot input matrix with a dimension equal to the dictionary size into a matrix with a dimension equal to the word embeddings, and then decompose the word embedding matrix into a matrix with a dimension equal to the model dimension, in order to reduce the number of training parameters for the model.
[0026] This invention also discloses a bilingual translation system for performing the above-described bilingual translation method, characterized in that the system comprises:
[0027] The data preprocessing module is used to construct a bilingual training dataset, connect the source language sequences and the target language sequences using the special symbols [CLS] and [SEP], and perform data formatting.
[0028] The model building module is used to build the BL-LMNMT bilingual translation model, which adopts a joint encoder-decoder structure based on the UNILM sequence-to-sequence model.
[0029] The multi-task learning module is used to simultaneously perform the source language category recognition subtask and the machine translation subtask. The source language category recognition subtask performs binary classification using the output at the [CLS] position, and the machine translation subtask is completed using a restricted multi-head self-attention network.
[0030] The model training module updates the model parameters by minimizing the total loss function;
[0031] The inference and verification module is used to perform bilingual translation using the trained BL-LMNMT model. It uses a beam search algorithm to generate translation results, with the beam size set to 4.
[0032] The results output module is used to output the generated target language translation results to the user in text form and to visualize the translation quality.
[0033] Compared with the prior art, the advantages of the present invention are as follows:
[0034] 1. This invention achieves bidirectional language translation functionality through the BL-LMNMT model, eliminating the need to train and deploy independent models for different language directions, thus greatly simplifying the system architecture and reducing maintenance complexity.
[0035] 2. By incorporating source language category recognition as a built-in subtask of the model, the system can automatically identify the language of the input text without the need for an additional language detection module, reducing the risk of error propagation and improving the robustness of the overall translation process.
[0036] 3. It integrates natural language understanding and generation capabilities. Through the collaborative training of two sub-tasks, source language recognition and machine translation, the model enhances its understanding of language features while completing the translation task, thereby improving the translation quality.
[0037] 4. By using parameter sharing mechanisms and matrix factorization techniques, the number of model parameters is significantly reduced; at the same time, a single model supports bidirectional translation, halving the number of models required and greatly reducing hardware resource consumption and storage costs.
[0038] 5. Employing the LAMB optimizer and mixed-precision training strategy, it supports training on large batches of data, significantly accelerating model convergence. The adaptive training stopping mechanism ensures that the model terminates training in its optimal state, avoiding overfitting and wasting computational resources.
[0039] 6. By combining the bundle search algorithm and the multi-head self-attention mechanism, the model can consider multiple possible translation results, effectively capture long-distance dependencies, and improve the accuracy and fluency of the translation results at the lexical, syntactic, grammatical and semantic levels.
[0040] 7. Through multilingual pre-training knowledge transfer and diverse training data processing, the model demonstrates excellent generalization ability, capable of handling sentences of different lengths and structures, and adapting to various practical application scenarios. Attached Figure Description
[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart of a bilingual translation method based on deep learning and neural network technology in an embodiment of the present invention;
[0043] Figure 2 This is the architecture diagram of the bilingual translation model BL-LMNMT in this embodiment of the invention;
[0044] Figure 3 This is a sequence-to-sequence model diagram of UNILM based on the Transformer module in this embodiment of the invention;
[0045] Figure 4 This is a BERT model diagram based on the Transformer encoder module in an embodiment of the present invention;
[0046] Figure 5 This is a structural diagram of the joint encoder-decoder in an embodiment of the present invention;
[0047] Figure 6 This is a diagram of the Transformer model architecture. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] Please refer to Figure 1 This invention provides a bilingual translation method based on deep learning and neural network technology, which may include the following steps:
[0050] Step S1: Detailed construction of the preprocessing dataset. For bilingual translation, a complete sequence pair dataset has been established, including but not limited to training set, validation set and test set.
[0051] To ensure the diversity and breadth of the dataset, this embodiment selected ten different bilingual pairings from IWSLT2017, covering DE-EN, DE-IT, DE-NL, DE-RO, EN-IT, EN-NL, EN-RO, IT-NL, IT-RO, and NL-RO. These pairings not only represent multiple mainstream European languages but also ensure diverse translation from source to target language. For the 20 bilingual translation tasks in the IWSLT2017 dataset, this embodiment selected over 200,000 sequence pairs as the primary data source for the training set. To ensure the robustness and generalization ability of the model, this embodiment selected dev2010 and test2010 as test sets, each containing 1,000 sequence pairs.
[0052] Preprocessing is a crucial step in NLP tasks, ensuring that the input data received by the model is cleaned and formatted uniformly. First, sequence tokenization and lowercase conversion are performed to make sentences more structured and reduce lexical complexity. This embodiment also specifically inserts a special classification marker "CLS" at the beginning of each source language sequence and a "SEP" marker at the end of each source and target language sequence. This not only helps distinguish different sequences but also allows the model to know when to end the decoding, avoiding redundancy in the output.
[0053] To ensure data accuracy and consistency, this embodiment uses Moses software to filter sequences exceeding 80 words, avoiding training difficulties caused by excessively long sequences. This embodiment also uses a bilingual true-case model to adjust the capitalization of the sequences, ensuring each word is in its most natural form. Furthermore, to enhance the model's generalization ability, this embodiment employs BERT's masking method, replacing some tokens with mask tokens. This allows the model to better learn to handle unknown or omitted words.
[0054] Considering that many words are used at different frequencies in different languages, this embodiment also performs BPE encoding on all sequences. This method can effectively reduce the occurrence of out-of-vocabulary words and rare words, further improving the processing efficiency and accuracy of the model.
[0055] Finally, considering the characteristics of bilingual translation, this embodiment creates a reverse version for each unidirectional dataset. For example, the DE-EN dataset corresponds to both DE-EN and EN-DE versions. This method ensures the bidirectionality and completeness of the dataset. After completing these steps, this embodiment randomly shuffles the sequence pairs in the entire bilingual dataset, which helps avoid overfitting during model training.
[0056] Step S2: Carefully configure hyperparameters, select appropriate deep learning tool libraries, and determine the required hardware resource configuration to lay a solid foundation for subsequent model training.
[0057] First, this embodiment carefully sets a series of hyperparameters to suit the model's needs. These hyperparameters are designed to ensure both model performance and training stability. For example, the learning rate is designed to use a warm-up strategy, which often accelerates model convergence and improves final performance in deep learning. Throughout training, this embodiment selects 0.1 as the dropout value to prevent overfitting and improve the model's generalization ability.
[0058] Specifically, regarding the model structure, this embodiment selects the Transformer structure (…). Figure 6 Some key parameters in the model are specified. The dimension of the linear transformation hidden state of each Transformer module is set to 3,072, which ensures that the model has sufficient capacity to learn complex language patterns. In addition, this embodiment also determines that the model dimension is 768, and the number of attention layers and heads of each Transformer module is 12, providing the model with sufficient depth and breadth to capture information at various levels.
[0059] For this bilingual translation task, this embodiment selected the UNILM sequence-to-sequence model as the benchmark. As an evaluation of model performance, this embodiment used the BLEU metric, which is currently a very popular and widely used evaluation metric in neural machine translation.
[0060] To facilitate model training and evaluation, this embodiment selects two popular and powerful deep learning toolkits: TensorFlow and Keras. Their excellent interoperability simplifies the model building, training, and testing processes. Of course, to ensure training efficiency, hardware resources are also considered. Using GPUs for training not only significantly accelerates computation but also effectively utilizes substantial video memory resources, which is crucial for handling large-scale neural network models and datasets.
[0061] Step S3: Construct the neural network model architecture BL-LMNMT and set the loss function. Build a network model for the bilingual translation task, determine the model input and output, and thus decide the model parameters to be trained in the network model; construct an appropriate loss function to solve the bilingual translation problem, including two tasks: identifying the source language category and machine translation; and learn the weight coefficients w1 and w2 of the losses l1 and l2 for the two tasks.
[0062] Figure 2It is the network architecture of the BL-LMNMT model. Although this model is similar to... Figure 3 The UNILM sequence-to-sequence model has a similar architecture, but it addresses different tasks. The UNILM sequence-to-sequence model solves unidirectional tasks such as automatic summarization, question answering, and machine translation, and its "[CLS]" symbol has no definitive meaning.
[0063] The BL-LMNMT model borrows from the BERT model's NSP and MLM training tasks and the UNILM model architecture. The BERT model's NSP task is used to determine the contextual relationship between two sentences; MLM understands the syntactic and semantic information of sentences in a self-supervised manner; and the BERT model uses bidirectional multi-head self-attention relationships to learn the words and sentence features of a sentence. The UNILM model employs a joint encoder and decoder structure (…). Figure 5 This structure employs a restricted multi-head attention network. The encoder part still uses bidirectional multi-head attention relationships, meaning that there are semantic relationships between the words in the source language sequence. The decoder part uses masked multi-head attention relationships, meaning that the words in the target language sequence have linguistic relationships with all the words in the source language sequence and the generated words in the target language sequence, but have no relationship with the ungenerated target words.
[0064] The common features of the three models BL-LMNMT, BERT and UNILM are that the input of the model is the concatenation of two sequences, they all use the special symbols [CLS] and [SEP], and they all use the same data preprocessing method.
[0065] The difference between BL-LMNMT and BERT lies in their implementations: BERT is based on a Transformer encoder architecture, where the two sequences in the training samples are of the same language. The NSP task uses the [CLS] notation to determine the contextual relationship between the two sequences, while the MLM task is used to learn the parameters of the Transformer encoder, such as the parameters of the multi-head self-attention layer. BL-LMNMT, on the other hand, is based on a Transformer encoder and decoder architecture, using the [CLS] notation to determine the source language category. Its mask is used to learn the parameters of the Transformer joint encoder and decoder. However, both models use [CLS] to perform a binary classification task, and the mask annotation method is used to learn the parameters of the Transformer block.
[0066] The difference between BL-LMNMT and UNILM is that UNILM's [CLS] has no practical function, while BL-LMNMT's [CLS] is used to determine the category of the source language. Like BERT's [CLS], both need to complete a binary classification task.
[0067] The model architecture of BL-LMNMT is as applied for. Figure 2 As shown. The input sample is “[CLS]word1[MASK]word3[SEP]word4[MASK][SEP]”, where “word1[MASK]word3” is the masked source language sequence, and “word4[MASK]” is the masked target language sequence. The masking method is that 12% of the words are replaced with the “[MASK]” label. The loss function for machine translation is l2, used to output the correct words that are “[MASK]”. Each output loss calculation corresponds to a multi-class classification task. The training process is to learn the application. Figure 4 The joint encoder-decoder parameters of multiple Transformer blocks. The loss function for source language category discrimination is l1, which is accomplished with the help of [CLS] labeling to complete the binary classification task.
[0068] In the BL-LMNMT model architecture, three advanced mechanisms are introduced to improve translation performance and training speed: parameter sharing mechanism, absolute position encoding, and LAMB optimizer.
[0069] To reduce the total number of parameters required for training, the BL-LMNMT model employs a parameter-sharing mechanism in two modules: the Transformer module and the embedding layer. On one hand, adopting the idea of ALBERT pre-training, the model shares parameters across all Transformer modules. On the other hand, to reduce the number of training parameters, the embedding layer utilizes ALBERT matrix factorization. In the embedding layer, the one-hot input matrix, with dimensions equal to the dictionary size, is first decomposed into matrices with dimensions equal to the word embeddings. Then, the word embedding matrix is further decomposed into matrices with dimensions equal to the model dimension. If the word embedding dimension is smaller than the model dimension, the number of training parameters can be significantly reduced.
[0070] In the embedding layer before the sequence of words enters the model, the original Transformer used trigonometric functions of absolute position for encoding, without considering word order and assuming that each word has the same status in the sentence. Therefore, the BL-LMNMT model introduces a relative position module, where neighboring words have a significant influence on each word, while words that are far away have a weaker influence.
[0071] During the backpropagation optimization process, the BL-LMNMT model employs the LAMB optimizer. While a larger batch size can improve task performance when the learning rate is stable, there is an implicit upper limit to batch size; exceeding this limit can make the loss function difficult to converge. The LAMB optimizer supports adaptive element-wise updates and accurate layer-by-layer corrections, allowing the model to maintain gradient accuracy during large-batch training and preventing the loss function from failing to converge due to abnormal gradient updates. Using the LAMB optimizer, the batch size can be increased to 32,000 without loss of accuracy, thus enabling the use of very large batch sizes to accelerate training.
[0072] Step S4: Model training and saving.
[0073] To improve model training speed and translation performance, this application uses a multilingual BERT pre-trained model as the basis for training the BL-LMNMT model. Compared to automatic summarization and question answering tasks, machine translation is a multilingual task (at least two languages, such as EN-DE). Typically, generation tasks use BERT pre-trained models instead of embedding layers, transferring them to downstream tasks to improve performance. However, in bilingual translation tasks, two languages are involved in training, while in many-to-one translation tasks, multiple languages are involved. Therefore, the model in this section uses a multilingual BERT pre-trained model. Furthermore, the BL-LMNMT model is a multi-task learning model, where the first task is to determine the source language type, and the second task is to generate the target language sequence. During training, the first task is a classification task, and the second task includes multiple multi-classification tasks. Clearly, the weights of the loss functions for these two tasks should not be equal. Compared to the second task, the objective of the first task is simpler, so its weight value should be lower. That is, the weight values of both should satisfy w1 < w1 < w2. <w2。
[0074] In training the BL-LMNMT model, the model parameters are first initialized. The network architecture parameters are initialized using Kaiming, with w1 set to 0.1 and w2 set to 10.
[0075] During training, the BL-LMNMT model is implemented using a teacher-driven approach. This model can be applied to two translation tasks with similar loss functions and two objectives: identifying the language type of the source sequence and learning the parameters of the joint encoder-decoder structure. Similar to the BERT model, due to its inherent advantages, this model can learn latent features of the source language category as a supervised classification task. Therefore, cross-entropy loss is chosen as the first loss l1. For the second objective, a good joint model can be trained using mask tags, which is essentially a multi-class classification task (m is the number of words in the mask). The goal of this task is to predict mask tags in a self-supervised learning manner. To achieve the second objective, the model adopts maximum likelihood probability as the training strategy; therefore, cross-entropy loss is used as the second loss l2. The overall loss function is a linear combination of the two loss functions, where w1 and w2 represent the learned coefficients used to balance the importance of loss functions l1 and l2. The training objective is to minimize the loss function l using the LAMB optimizer. For bilingual translation tasks, l1 is used as the loss function to solve both binary and multi-class classification problems.
[0076] This application employs a mixed-precision training method, which can accelerate training speed and save GPU memory resources. Common training methods use FP32 format weights to calculate gradients and update model parameters, while mixed-precision training first uses FP16 format weights to calculate gradients, then converts the gradients to FP32 format, and finally updates the FP32 format weights.
[0077] During training, a minimum threshold for the loss function is first set, such as 1e-5. This threshold is kept as small as possible to ensure that the parameters converge and the model reaches a stable and optimal state. Simultaneously, a maximum length for each sample is set, such as 128. Since the bilingual translation task involves short sentences, the maximum length of all samples does not exceed 300. This application tests two methods for calculating the average sample length: first, the total sample length of the training set is calculated, and the average is used as the maximum sample length; second, the maximum and minimum lengths of all samples are calculated, and the average of the two is used as the maximum sample length.
[0078] In addition, the batch size required for each training session needs to be set. It should not be too large or too small. Too large a batch will exceed the GPU's storage capacity, while too small a batch will affect the training speed. This application uses a Tesla V100 GPU, and the batch size is set to 32.
[0079] In each training round, the gradient of the parameter weights based on the loss function is calculated once, and the gradient descent algorithm is used to optimize the parameter weights, as well as the values of w1 and w2.
[0080] After obtaining the parameter weights, a temporary model can be saved. Based on this model, a temporary loss value is obtained using the forward propagation algorithm. If this value is greater than a threshold, training continues on the next batch of training sets, continuously optimizing the model until the temporary loss function value is less than or equal to the threshold, thus obtaining the optimal bilingual translation model. This model is then saved for inference verification and application deployment.
[0081] Step S5: Model Inference and Validation. Validate the model on multiple validation sets and evaluate the model using the BLUE metric.
[0082] During the decoding process, the source language sequence is known, and the target language sequence words are generated in an autoregressive manner. Decoding terminates when the "EOS" token is predicted, and the predicted sequence at this point represents the decoding result of the source language sequence. In essence, the probability distribution of the target word at each time step is obtained at this point. Then, a beam search algorithm is used to obtain the target language sequence with the highest probability; the beam size in this application is set to 4.
[0083] Step S6: Model Application. Apply the model to multiple test sets and perform model analysis using multiple test cases.
[0084] On five EN-DE and DE-EN bilingual tasks, the UNILM sequence-to-sequence model and the BL-LMNMT bilingual translation model were used respectively to obtain 10 translation results. The results were manually analyzed from four levels: lexical, syntactic, grammatical and semantic.
[0085] The five EN sentences and five DE sentences sampled are not in a one-to-one correspondence; there is no single correct answer for translating each sentence. This facilitates manual analysis and verification of the corresponding results. To verify the model's generalization ability, these sentences are diverse. First, the sentences vary in length: the shortest is one sentence containing two words, the longest is one sentence containing 200 words, and the remaining three sentences are between 80 and 100 words in length. Second, in terms of sentence structure, there is one simple sentence with a subject-verb structure, one compound sentence connected by "and" (English) or "Und" (German), and three complex sentences connected by the subordinating conjunction whether (English) or ob (German), the interrogative pronoun who (English) or who (German), and the interrogative adverb when (English) or wanna (German), respectively.
[0086] Five students were used to cross-validate the translation results on 10 pairs to ensure fairness in assessing the reliability of the model.
[0087] In another embodiment, a bilingual translation system is provided, which corresponds one-to-one with the bilingual translation methods in the above embodiments. It includes the following functional modules:
[0088] The data preprocessing module is used to construct a bilingual training dataset, connect the source language sequences and the target language sequences using the special symbols [CLS] and [SEP], and perform data formatting.
[0089] The model building module is used to build the BL-LMNMT bilingual translation model, which adopts a joint encoder-decoder structure based on the UNILM sequence-to-sequence model.
[0090] The multi-task learning module is used to simultaneously perform the source language category recognition subtask and the machine translation subtask. The source language category recognition subtask performs binary classification using the output at the [CLS] position, and the machine translation subtask is completed using a restricted multi-head self-attention network.
[0091] The model training module updates the model parameters by minimizing the total loss function;
[0092] The inference and verification module is used to perform bilingual translation using the trained BL-LMNMT model. It uses a beam search algorithm to generate translation results, with the beam size set to 4.
[0093] The results output module is used to output the generated target language translation results to the user in text form and to visualize the translation quality.
[0094] In another embodiment of the present invention, a terminal device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of a bilingual translation method.
[0095] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). This computer-readable storage medium is a memory device in a terminal device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and extended storage media supported by the terminal device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device.
[0096] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the bilingual translation method in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by a processor.
[0097] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0099] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A bilingual translation method based on deep learning and neural network technology, characterized in that: The method includes the following steps: A bilingual translation model BL-LMNMT is constructed. This model adopts a joint encoder-decoder structure based on the UNILM sequence-to-sequence model. The input is a combination of source language sequence and target language sequence connected by the special symbols [CLS] and [SEP]. The BL-LMNMT model is designed with a multi-task learning mechanism, including a source language category recognition subtask and a machine translation subtask. The source language category recognition subtask is implemented using BERT's NSP task, and the machine translation subtask is completed using a restricted multi-head self-attention network. The BL-LMNMT model is trained by updating the model parameters by minimizing the total loss function, which is a weighted sum of the loss functions of the two sub-tasks: total loss = w1×l1 + w2×l2, where w1 and w2 are weight coefficients, l1 is the loss function of the source language category recognition sub-task, and l2 is the loss function of the machine translation sub-task. The trained BL-LMNMT model is used for bilingual translation, automatically identifying the source language and outputting the corresponding target language.
2. The bilingual translation method according to claim 1, characterized in that: In the data preprocessing stage, handling uncertainties in the source language includes: Construct a dataset from source language to target language: For each sample, process it using the format "[CLS]Source Language Identifier[SEP]Target Language Identifier[SEP]", and randomly shuffle the order of all samples; Construct a target language to source language dataset: For each sample, process it using the format "[CLS]Target Language Identifier[SEP]Source Language Identifier[SEP]", and randomly shuffle the order of all samples; The datasets from the source language to the target language and from the target language to the source language are merged, and the overall sample order is randomly shuffled again.
3. The bilingual translation method according to claim 1, characterized in that: The weight coefficient combination in the total loss function is set to w1:w2=1:100 to balance the difference in importance between the source language category recognition subtask and the machine translation subtask.
4. The bilingual translation method according to claim 1, characterized in that: In the source language category recognition subtask, the model performs binary classification through the output at the [CLS] position, and the output probability distribution [p1, p2] satisfies p1 + p2 = 1, where p1 represents the source language probability and p2 represents the target language probability; the KL divergence between the output probability distribution and the true probability distribution [1,0] or [0,1] is minimized by adjusting the model parameters.
5. The bilingual translation method according to claim 1, characterized in that: During the inference phase, a beam search algorithm is used to generate translation results. The beam size is set to 4. By calculating the probability product of all possible target sequences, the sequence with the largest probability product is selected as the final translation result.
6. The bilingual translation method according to claim 1, characterized in that: During model training, a minimum loss threshold of 1e-5 is set. When the model's loss value on the validation set is less than or equal to this threshold, training is stopped and the current model parameters are saved.
7. The bilingual translation method according to claim 1, characterized in that: The LAMB optimizer is used for model training, which supports large-scale training.
8. The bilingual translation method according to claim 1, characterized in that: Using a multilingual BERT pre-trained model as the initialization basis for the BL-LMNMT model, the pre-trained knowledge is transferred to the bilingual translation task to improve model performance.
9. The bilingual translation method according to claim 1, characterized in that: In the embedding layer, parameter sharing and matrix factorization techniques are used to decompose the one-hot input matrix with a dimension equal to the size of the dictionary into a matrix with a dimension equal to the word embedding, and then decompose the word embedding matrix into a matrix with a dimension equal to the model dimension, so as to reduce the number of training parameters of the model.
10. A bilingual translation system for performing the bilingual translation method according to any one of claims 1 to 9, characterized in that, The system includes: The data preprocessing module is used to construct a bilingual training dataset, connect the source language sequences and the target language sequences using the special symbols [CLS] and [SEP], and perform data formatting. The model building module is used to build the BL-LMNMT bilingual translation model, which adopts a joint encoder-decoder structure based on the UNILM sequence-to-sequence model. The multi-task learning module is used to simultaneously perform the source language category recognition subtask and the machine translation subtask. The source language category recognition subtask performs binary classification using the output at the [CLS] position, and the machine translation subtask is completed using a restricted multi-head self-attention network. The model training module updates the model parameters by minimizing the total loss function; The inference and verification module is used to perform bilingual translation using the trained BL-LMNMT model. It uses a beam search algorithm to generate translation results, with the beam size set to 4. The results output module is used to output the generated target language translation results to the user in text form and to visualize the translation quality.