A model training, machine translation method, device, equipment and storage medium
By using the method of exchanging parallel corpus and bidirectional training in machine translation model training, the problem of low machine translation accuracy in the prior art is solved, and better machine translation performance is achieved, especially in low resource scenarios.
Patent Information
- Application Number
- CN202210686002.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-16
AI Technical Summary
The training effect of existing machine translation models is poor, resulting in low machine translation accuracy.
By obtaining the original parallel corpus, the source side data is used as the target side data, and the target side data is exchanged as the source side data, forming an exchange parallel corpus, and bidirectionally training the original translation model based on multiple sets of original and exchange parallel corpus to obtain the intermediate translation model, and then forward training the intermediate translation model based on the original parallel corpus to obtain the machine translation model.
It improves the model training effect and improves machine translation performance, especially in low-resource scenarios, effectively solves the problem of insufficient training samples.
Smart Images

Figure CN115130481B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of natural language processing technology, and in particular to a model training, machine translation method, apparatus, device, and storage medium. Background Art
[0002] Machine translation is a key research area in natural language processing and artificial intelligence. It aims to use computers to automatically translate between natural languages. With the advent of the deep learning era, machine translation technology has achieved breakthroughs.
[0003] In the process of realizing the present invention, the inventors discovered that the following technical problems exist in the prior art: due to the poor training effect of the existing machine translation model, the current machine translation accuracy needs to be improved. Summary of the Invention
[0004] Embodiments of the present invention provide a model training, machine translation method, apparatus, device and storage medium, which solve the problem of low machine translation accuracy caused by poor training effect of the machine translation model.
[0005] According to one aspect of the present invention, a model training method is provided, which may include:
[0006] Obtaining original parallel corpora including original source data and original target data;
[0007] The original source data is used as the exchange target data, and the original target data is used as the exchange source data to obtain exchange parallel corpora, and the original translation model is trained based on the multiple sets of original parallel corpora and the multiple sets of exchange parallel corpora to obtain an intermediate translation model;
[0008] The intermediate translation model is trained based on multiple sets of original parallel corpora to obtain a machine translation model.
[0009] According to another aspect of the present invention, a machine translation method is provided, which may include:
[0010] Obtaining source data to be translated and a machine translation model trained according to the model training method provided in any embodiment of the present invention, wherein the source data to be translated is in the same language as the original source data in the model training method;
[0011] The source data to be translated is input into the machine translation model, and the translated target data is obtained based on the output of the machine translation model.
[0012] According to another aspect of the present invention, a model training device is provided, which may include:
[0013] A corpus acquisition module is used to acquire original parallel corpora including original source data and original target data;
[0014] A bidirectional training module is used to use the original source data as the exchange target data and the original target data as the exchange source data to obtain exchange parallel corpora, and to train the original translation model based on multiple sets of original parallel corpora and multiple sets of exchange parallel corpora to obtain an intermediate translation model;
[0015] The forward training module is used to train the intermediate translation model based on multiple sets of original parallel corpora to obtain a machine translation model.
[0016] According to another aspect of the present invention, a machine translation apparatus is provided, which may include:
[0017] A model acquisition module, configured to acquire source data to be translated and a machine translation model trained according to the model training method provided in any embodiment of the present invention, wherein the source data to be translated is in the same language as the original source data in the model training method;
[0018] The machine translation module is used to input the source data to be translated into the machine translation model and obtain the translated target data based on the output results of the machine translation model.
[0019] According to another aspect of the present invention, there is provided an electronic device, which may include:
[0020] at least one processor; and
[0021] a memory communicatively connected to at least one processor; wherein,
[0022] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by at least one processor so that the at least one processor implements the model training method or machine translation method provided by any embodiment of the present invention when executing the computer program.
[0023] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer instructions are stored. The computer instructions are used to enable a processor to implement the model training method or machine translation method provided by any embodiment of the present invention when executed.
[0024] The technical solution of the embodiment of the present invention obtains original parallel corpora including original source data and original target data; uses the original source data as exchange target data, and uses the original target data as exchange source data to obtain exchange parallel corpora, and trains the original translation model based on multiple sets of original parallel corpora and multiple sets of exchange parallel corpora to obtain an intermediate translation model; then, trains the intermediate translation model based on multiple sets of original parallel corpora to obtain a machine translation model. The above technical solution solves the problem of insufficient training samples in low-resource scenarios by exchanging the original source data and the original target data and adding the exchange results to the training samples. On this basis, by adding bidirectional training to the forward training, all information in the bilingual data can be fully learned. The two work together to improve the model training effect and obtain a machine translation model with better machine translation performance.
[0025] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 is a flowchart of a model training method provided according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of bidirectional data in a model training method provided by an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of bidirectional training in a model training method provided according to an embodiment of the present invention;
[0030] Figure 4 is a schematic diagram of forward training in a model training method provided according to an embodiment of the present invention;
[0031] Figure 5a is a framework diagram of an optional example of a model training method provided according to an embodiment of the present invention;
[0032] Figure 5b is a flowchart of an optional example of a model training method provided according to an embodiment of the present invention;
[0033] Figure 6 is a flowchart of another model training method provided according to an embodiment of the present invention;
[0034] Figure 7 is a flowchart of a machine translation method provided according to an embodiment of the present invention;
[0035] Figure 8 is a structural block diagram of a model training device provided according to an embodiment of the present invention;
[0036] Figure 9 is a structural block diagram of a machine translation device provided according to an embodiment of the present invention;
[0037] Figure 10 It is a structural diagram of an electronic device for implementing the model training method or machine translation method of an embodiment of the present invention. DETAILED DESCRIPTION
[0038] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0039] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The situations of "target", "original", etc. are similar and will not be repeated here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.
[0040] Before introducing the embodiments of the present invention, an exemplary description of the application scenarios of the embodiments of the present invention is given first: Bilingual data is a very important part of machine translation. In real scenarios, bilingual data is extremely scarce, and this low-resource scenario directly affects the training effect of the machine translation model. Moreover, currently, when using bilingual data, usually only one-way language information is utilized, and all the information in the bilingual data is not utilized. However, through the study of the human learning behavior pattern, it is found that two-way language learning can better learn language information. Therefore, the current data utilization method will also affect the training effect of the machine translation model. To solve the above problems, the inventors propose the model training methods in the following embodiments. Specifically,
[0041] Figure 1 FIG. is a flowchart of a model training method provided in an embodiment of the present invention. This embodiment is applicable to the situation of training a machine translation model in a low-resource scenario. This method can be executed by a model training device provided in an embodiment of the present invention. The device can be implemented in a software and / or hardware manner, and the device can be integrated on an electronic device, and the device can be various user terminals or servers.
[0042] See Figure 1 , the method of the embodiment of the present invention specifically includes the following steps:
[0043] S110. Obtain an original parallel corpus including original source-side data and original target-side data.
[0044] Among them, the original parallel corpus can be an unprocessed parallel corpus directly obtained, which can include original source-side data and its parallel corresponding original target-side data. In terms of natural language, the original target-side data can be considered as the translation of the original source-side data. For example, "Hello" is the translation of "你好". In practical applications, optionally, the above original parallel corpus can be a corpus at any of the following levels: chapter level, paragraph level, sentence level, phrase level, word level, etc., and no specific limitation is made here.
[0045] S120. Take the original source-side data as the exchanged target-side data, and take the original target-side data as the exchanged source-side data to obtain an exchanged parallel corpus, and train an original translation model based on multiple groups of original parallel corpora and multiple groups of exchanged parallel corpora to obtain an intermediate translation model.
[0046] Among them, the original source end data is used as the exchange target end data, and the original target end data is used as the exchange source end data, that is, the original source end data and the original target end data in the original parallel corpus are exchanged to obtain the exchange parallel corpus. Then, the original parallel corpus and the exchange parallel corpus are both used as training samples for model training, that is, the exchange parallel corpus is added to the original training sample (that is, the training sample containing only the original parallel corpus), thereby achieving the effect of doubling the number of samples. For example, see Figure 2 , assuming that the training samples consisting of multiple sets of original parallel corpora are defined as:
[0047]
[0048] Where N is the number of samples, x i Represents the i-th original source data, y i Represents the i-th original target data. Since the training direction is from x to y, the training sample (i.e., bilingual data) is recorded as Exchange the original source data and the original target data and add them to The new training samples (i.e., bidirectional data) obtained in this way can be expressed as follows:
[0049]
[0050] Furthermore, the original translation model is trained based on multiple groups of original parallel corpora and multiple groups of exchange parallel corpora to obtain an intermediate translation model. The original translation model may be a machine learning model to be trained for implementing machine translation, such as a statistical machine translation model or a neural network machine translation model, and the neural network machine translation model may be a neural network model to be trained for implementing machine translation, such as a recurrent neural network (RNN) model, a neural network model composed of an encoder-decoder framework (Transformer) based on a self-attention neural network, etc., which are not specifically limited here. In practical applications, optionally, the concept of batch size (batchsize) is involved in the model training process. The original parallel corpora and exchange parallel corpora under a batchsize are not necessarily one-to-one corresponding, that is, a certain original parallel corpus and the exchange parallel corpus corresponding to the original parallel corpus do not necessarily exist under the same batchsize, which is not specifically limited here. Exemplarily, the model training process of this step can be understood as follows Figure 3 The bidirectional training process of the original translation model based on bidirectional data is shown.
[0051] Regarding this step, it should be noted that, on the one hand, by exchanging the original source data and the original target data in the original parallel corpus, the problem of insufficient training samples in low-resource scenarios is solved. On the other hand, the information of forward translation and reverse translation in machine translation is taken into account at the same time. By adding a two-way training method that fits human learning behavior, the encoder and decoder are strengthened to understand the source information and target information, thereby improving the alignment quality at both ends and improving the model training effect. On this basis, because there is no need to rely on external tools (such as word alignment or monolingual knowledge) and no need to involve complex model structure improvements, it can be applied to a wider range of language scenarios and model architectures, thereby achieving the effect of model translation in multiple application scenarios, with good versatility.
[0052] S130. Train an intermediate translation model based on multiple sets of original parallel corpora to obtain a machine translation model.
[0053] The intermediate translation model obtained through the above-mentioned two-way training process has fully learned the source and target information on a large number of training samples. Therefore, the intermediate translation model can be trained on multiple sets of original parallel corpora to obtain the final machine translation model. For example, the model training process in this step can be understood as follows: Figure 4 The forward training process of the intermediate translation model based on bilingual data (i.e., original parallel corpus) is shown.
[0054] The technical solution of the embodiment of the present invention obtains original parallel corpora including original source data and original target data; uses the original source data as exchange target data, and uses the original target data as exchange source data to obtain exchange parallel corpora, and trains the original translation model based on multiple sets of original parallel corpora and multiple sets of exchange parallel corpora to obtain an intermediate translation model; then, trains the intermediate translation model based on multiple sets of original parallel corpora to obtain a machine translation model. The above technical solution solves the problem of insufficient training samples in low-resource scenarios by exchanging the original source data and the original target data and adding the exchange results to the training samples. On this basis, by adding bidirectional training to the forward training, all information in the bilingual data can be fully learned. The two work together to improve the model training effect and obtain a machine translation model with better machine translation performance.
[0055] On this basis, an optional technical solution is that after obtaining the original parallel corpus including the original source-side data and the original target-side data, the above model training method may further include: performing data augmentation on the original source-side data and / or the original target-side data in the original parallel corpus to obtain an augmented parallel corpus, and using both the augmented parallel corpus and the original parallel corpus as the original parallel corpus. Herein, the data augmentation of the original source-side data and / or the original target-side data can be understood as data augmentation achieved by utilizing the correspondence relationship between bilingual data, which can be implemented in various ways in practical applications, such as Curriculum learning (CL), Back Translation (BT), Knowledge Distillation (KD), Data Diversification (DD), etc., and no specific limitation is made herein. Exemplarily, taking the bilingual data as Hello→你好, after performing addition, deletion, or modification on Hello to obtain Hella, on the premise of maintaining the correspondence relationship unchanged, an augmented parallel corpus composed of Hella→你好 can thus be obtained. Through the above technical solution, by performing data augmentation on the original source-side data and / or the original target-side data in the original parallel corpus and then using the obtained augmented parallel corpus as the original parallel corpus, the problem of insufficient training samples in low-resource scenarios is solved. On this basis, in cooperation with the data exchange and bidirectional training solutions in the embodiments of the present invention, the model training effect is further improved.
[0056] Another optional technical solution, after obtaining the original parallel corpus including the original source-side data and the original target-side data, the above-mentioned model training method may also include: obtaining monolingual source-side data of the same language as the original source-side data, and a preliminary translation model that has been preliminarily trained, inputting the monolingual source-side data or the original source-side data into the preliminary translation model to obtain pseudo-target-side data, and using the pseudo-parallel corpus composed of the monolingual source-side data and the pseudo-target-side data, or the original source-side data and the pseudo-target-side data, and the original parallel corpus as the original parallel corpus. As mentioned above, in real scenarios, bilingual data is very scarce, but monolingual data exists in large quantities. Therefore, on the basis of the original parallel corpus, data enhancement can be performed by introducing additional monolingual data. Specifically, language is the abbreviation of language type, such as Chinese, English, French, German or Japanese. The monolingual source data can be data of the same language as the original source data and for which there is no parallel translation corresponding thereto. The preliminary translation model can be a model obtained through preliminary training that can realize the machine translation function. In practical applications, it can optionally be a model obtained after training the original translation model based on multiple groups of original parallel corpora. The monolingual source data or the original source data is input into the preliminary translation model to obtain pseudo-target data, thereby obtaining a pseudo-parallel corpus consisting of the monolingual source data and the pseudo-target data or a pseudo-parallel corpus consisting of the original source data and the pseudo-target data. Furthermore, the pseudo-parallel corpus and the directly obtained original parallel corpus are both used as the original parallel corpus to perform subsequent steps, thereby solving the problem of insufficient training samples in low-resource scenarios. On this basis, it is combined with the data exchange and two-way training scheme in the embodiment of the present invention to further improve the model training effect.
[0057] Another optional technical solution, the above-mentioned model training method may further include: obtaining a preset total number of training steps and a ratio of training steps for bidirectional training; obtaining the bidirectional training steps of bidirectional training and the forward training steps of forward training based on the total number of training steps and the ratio of training steps; accordingly, training the original translation model based on multiple sets of original parallel corpora and multiple sets of exchanged parallel corpora to obtain an intermediate translation model may include: training the original translation model based on multiple sets of original parallel corpora for the number of bidirectional training steps to obtain an intermediate translation model; training the intermediate translation model based on multiple sets of original parallel corpora to obtain a machine translation model may include: training the intermediate translation model based on multiple sets of original parallel corpora for the number of forward training steps to obtain a machine translation model. Among them, compared with the single forward training process, in order to avoid the increase in the time cost of model training due to the addition of the bidirectional training process, that is, in order to avoid the situation where the accuracy performance of model training is improved by sacrificing the time performance of model training, the total number of training steps required for the single forward training process and the ratio of the number of training steps of the bidirectional training process in the total training steps can be obtained, and then the bidirectional training steps of the bidirectional training and the forward training steps of the forward training can be obtained based on the two, and bidirectional training is performed according to the bidirectional training steps and forward training is performed according to the forward training steps, thereby achieving the effect of improving the accuracy of model training while keeping the total number of training steps unchanged.
[0058] In order to better understand the above technical solutions as a whole, the following is an exemplary description of them with reference to specific examples. Figure 5a and Figure 5b , swap the original source data and the original target data in the original parallel corpus to obtain the swapped parallel corpus, then mix the original parallel corpus and the swapped parallel corpus to obtain bidirectional data. Assuming the training step ratio is 1 / 3, 1 / 3 of the total training steps can be used for bidirectional training on the bidirectional data to obtain an intermediate translation model. Further, 2 / 3 (1-1 / 3) of the total training steps can be used for forward training on the original parallel corpus (i.e., bilingual data) to obtain a machine translation model. Then, machine translation can be performed based on this machine translation model.
[0059] Figure 6It is a flow chart of another model training method provided in an embodiment of the present invention. This embodiment is optimized based on the above-mentioned technical solutions. In this embodiment, optionally, before the original translation model is trained based on multiple groups of original parallel corpora and multiple groups of exchange parallel corpora to obtain the intermediate translation model, the above-mentioned model training method may also include: respectively performing word segmentation on the original source-end data and the original target-end data in the original parallel corpora, and performing subword segmentation on the obtained word segmentation results to obtain the original subword representation; updating the original parallel corpora based on the original subword representation corresponding to the original source-end data and the original subword representation corresponding to the original target-end data; respectively performing word segmentation on the exchange source-end data and the exchange target-end data in the exchange parallel corpora, and performing subword segmentation on the obtained word segmentation results to obtain the exchange subword representation; updating the exchange parallel corpora based on the exchange subword representation corresponding to the exchange source-end data and the exchange subword representation corresponding to the exchange target-end data. Among them, the explanations of the terms that are the same as or corresponding to the above-mentioned embodiments are not repeated here.
[0060] See also Figure 6 The method of this embodiment may specifically include the following steps:
[0061] S210: Obtain original parallel corpus including original source-end data and original target-end data.
[0062] S220 , performing word segmentation on the original source-end data and the original target-end data in the original parallel corpus, and performing subword segmentation on the obtained word segmentation results to obtain original subword representations.
[0063] Among them, after obtaining multiple original parallel corpora, since there may be a large amount of different original source data and original target data, these large amounts of different original source data and original target data will not only take up a lot of data storage space, but also affect the number of network parameters in the original translation model during the model training process, thereby affecting the accuracy and timeliness of model training. Therefore, they can be compressed first, and then the model training can be performed based on the data compression results. Specifically,
[0064] The original source data and the original target data in the original parallel corpus are segmented separately to obtain segmentation results. Since these segmentation results may contain some frequently occurring original subword representations, such as "app" in "apple," "application," "app," and "approach," we can extract "app" as a separate original subword representation. This can then be combined with the remaining segmentation results to extract additional original subword representations. This allows a large number of segmentation results to be represented using a limited number of original subword representations, thus achieving data compression.
[0065] S230: Update the original parallel corpus based on the original subword representation corresponding to the original source-end data and the original subword representation corresponding to the original target-end data.
[0066] The original source data is updated based on the original subword representations corresponding to the original source data, and the original target data is updated based on the original subword representations corresponding to the original target data, thereby achieving the effect of updating the original parallel corpus. Furthermore, bidirectional training and forward training can be combined with subsequent steps. For forward training, for example, the original subword representations under a corresponding relationship can be input into the intermediate training model in pairs for model training, thereby further improving the model training effect through data compression.
[0067] S240: Use the original source-end data as exchange target-end data, and use the original target-end data as exchange source-end data to obtain exchange parallel corpus.
[0068] S250 , performing word segmentation on the exchange source-end data and the exchange target-end data in the exchange parallel corpus, and performing subword segmentation on the obtained word segmentation results to obtain exchange subword representations.
[0069] The implementation process of this step is similar to that of S220 and will not be described in detail here. It should be noted that the original subword representation and the exchanged subword representation are essentially subword representations. The different names here are only used to distinguish the objects of subword segmentation, and are not a specific limitation on their actual meanings.
[0070] S260: Update the exchange parallel corpus based on the exchange subword representation corresponding to the exchange source data and the exchange subword representation corresponding to the exchange target data.
[0071] The implementation process of this step is similar to that of S230 and will not be repeated here.
[0072] S270: Training the original translation model based on the multiple sets of original parallel corpora and the multiple sets of exchanged parallel corpora to obtain an intermediate translation model.
[0073] S280: Train an intermediate translation model based on multiple sets of original parallel corpora to obtain a machine translation model.
[0074] The technical solution of the embodiment of the present invention compresses the original parallel corpus and the exchanged parallel corpus through the technical means of word segmentation and subword segmentation, thereby further improving the model training effect.
[0075] On the basis of any of the above technical solutions, optionally, the bidirectional loss function matched with the original translation model includes a forward loss function and a reverse loss function; the original translation model is trained based on multiple groups of original parallel corpora and multiple groups of exchanged parallel corpora to obtain an intermediate translation model, which may include: for each group of original parallel corpora, the original source-end data in the original parallel corpora are input into the original translation model to obtain forward target-end data; for each group of exchanged parallel corpora, the exchanged source-end data in the exchanged parallel corpora are input into the original translation model to obtain reverse target-end data; combining the forward loss function, the forward loss is obtained based on the forward target-end data and the original target-end data, and combining the reverse loss function, the reverse loss is obtained based on the reverse target-end data and the exchanged target-end data; the bidirectional loss is obtained according to the forward loss and the reverse loss, and the network parameters in the original translation model are adjusted according to the bidirectional loss to obtain the intermediate translation model. For example, in order to understand the above-mentioned bidirectional loss function more vividly, it is illustrated below with reference to specific examples. For example, see the following formula:
[0076]
[0077] Among them L BiT (θ) represents the bidirectional loss function, arg max θ logp(y|y';θ) represents the forward loss function, argmax θ logp(x|x';θ) represents the reverse loss function, x represents the exchanged target data, x' represents the reverse target data, y represents the original target data, y' represents the forward target data, θ represents the network parameters, and p represents the probability.
[0078] Optionally, after obtaining the machine translation model, the above-mentioned model training method may further include: obtaining verification parallel corpus, and inputting the verification source data in the verification parallel corpus into the machine translation model, obtaining translation target data according to the output result of the machine translation model; matching the translation target data with the verification target data in the verification parallel corpus to obtain the machine translation accuracy. In order to verify the effectiveness of the above-mentioned model training method, after obtaining the machine translation model, verification parallel corpus may be obtained, which may include verification source data and verification target data. Then, the verification source data is input into the machine translation model to obtain translation target data, and then the translation target data and the verification target data may be matched to obtain the machine translation accuracy, so as to determine whether the machine translation model obtained by the above-mentioned training meets the standards based on the machine translation accuracy.
[0079] In practical applications, optionally, the above-mentioned machine translation accuracy can be represented by a BLEU score. Optionally, the above-mentioned matching process can be implemented by the following steps: for each single word in the translation target end data, it is matched with each single word in the verification target end data in sequence to obtain a first matching result; for each two adjacent single words in the translation target end data, it is matched with each two adjacent single words in the verification target end data in sequence to obtain a second matching result; and then, the machine translation accuracy is obtained based on the first matching result and the second matching result. For example, assuming that the translation target end data is ABCD and the verification target end data is BCEF, when matching based on 1 single word, A is matched with B, C, E and F respectively, and the matching degrees are all 0; then B is matched with B, C, E and F respectively, and the matching degrees are 1, 0, 0 and 0 respectively; and so on for the processing of C and D, thereby obtaining the first matching result. Furthermore, when matching based on two words, AB is matched with BC, CE, and EF, respectively; BC is matched with BC, CE, and EF, respectively; and so on for CD, thereby obtaining a second matching result. Alternatively, matching can be performed based on three or four words, which is not specifically limited here. The machine translation accuracy can then be obtained based on each matching result.
[0080] To verify the effectiveness of the model training method proposed in this embodiment of the present invention, the following experiments were conducted on the IWSLT2014 English-German & German-English datasets, the WMT2016 English-Romanian & Romanian-English datasets, the IWSLT2021 English-Swahili & Swahili-English datasets, and the WMT2014 and 2019 English-German & German-English datasets. The results are shown in Table 1 (where 160K indicates 160,000 sentences in the data source, 0.6M indicates 600,000 sentences in the data source, and 20M indicates 20 million sentences in the data source):
[0081] Table 1 Experimental results under different data scales
[0082]
[0083] The last two rows of data in Table 1 are both in percentages. The second-to-last row is the BLEU score of the machine translation model obtained through forward training, and the last row is the BLEU score of the machine translation model trained using the model training method of the embodiment of the present invention (BiT represents bidirectional training, and +BiT represents adding bidirectional training to forward training). As can be seen from Table 1, the model training method of the embodiment of the present invention achieved a significant increase under p<0.01 on 7 / 10 tasks (through ), and achieved significant improvements at p < 0.05 on another 3 / 10 tasks (via denoted tasks), achieving an average significant improvement of +1.1, demonstrating the effectiveness and versatility of this model training method. Notably, this model training method can save one-third of the training cost for reverse training. For example, a bidirectional update model pre-trained for English-German translation can be used for the reverse direction, German-English translation. This advantage demonstrates that this model training method is well suited for multilingual scenarios, such as multilingual pre-training and translation. Pre-training can be understood as training performed before forward training, such as the bidirectional training in this embodiment of the present invention.
[0084] In addition, we selected two languages with significant differences in language families (Chinese: Sino-Tibetan, English: Indo-European, and Japanese: Japanese-Ryukyuan): WMT2017 Chinese-English & English-Chinese, and WAT2017 Japanese-English to verify the performance of the model training method in these situations. The experimental results are shown in Table 2. It can be seen that even in the case of large language differences, the model training method still achieved an average significance improvement of +0.9.
[0085] Table 2 Experimental results in languages with large language differences
[0086]
[0087] In addition, the complementarity of this model training method with existing work is also verified: the complementary results with three typical data augmentation works are listed here, including: BT, KD and DD. The experimental results are shown in Table 3. It can be seen that this model training method can be combined with existing data augmentation work to achieve further improvement.
[0088] Table 3 Complementarity verification with classic data augmentation work
[0089]
[0090] In addition to the above experiments, other experiments and analyses were conducted, and the following conclusions were obtained:
[0091] 1) The aforementioned bidirectional training strategy is a better and simpler bilingual code-switcher. Related work has shown that using code-switching for pre-training can effectively improve downstream multilingual translation performance. However, it relies on unsupervised word alignment tools from three parties to extract alignment information, which is then used to perform code-switching replacements of fragments at different granularities. Experimental analysis demonstrates that this bidirectional training strategy is a sentence-level code-switching method with a replacement probability of 0.5. For example, consider the English-Chinese sentence {"A held a talk with B" -> "A and B held a meeting"}. During pre-training, the reconstructed pre-training data includes both the forward pair {"A held a talk with B" -> "A and B held a meeting"} and the reverse pair {"A and B held a meeting" -> "A held a talk with B"}. In this case, the reverse pair can be considered a sentence-level switch with a probability of 0.5. To verify the above statement, we compared two classic code-switch pre-training methods. The experimental results are shown in Table 4. We found that this bidirectional training strategy is indeed an excellent alternative to code-switch in bilingual scenarios. Code-switch can be understood as the process of replacing some source words with aligned words on the target side.
[0092] Table 4 Comparison with code-switch pre-training method
[0093]
[0094] 2) The aforementioned bidirectional training strategy can improve alignment quality. This bidirectional training strategy encourages the self-attention mechanism to learn better bilingual relationships, and therefore has great potential for obtaining a better bilingual attention matrix, i.e., alignment information. To verify this, experiments were conducted on the Gold Alignment dataset with alignment labels, and evaluated based on alignment error rate (AER), precision (P), and recall (R). The experimental results are shown in Table 5. It can be seen that compared to the forward training method alone, this bidirectional training strategy can achieve a significant improvement in alignment quality (27.1% vs. 24.3%).
[0095] Table 5 Experimental results of alignment quality
[0096]
[0097] 3) The aforementioned model training method remains effective in extremely low-resource scenarios. Experiments were conducted on the English-Gujarati and Gujarati-English language pairs, low-resource scenarios where BackTranslation failed in the WMT2019 competition. The results are shown in Table 6. While direct application of BackTranslation does result in a slight decrease in translation quality (-0.4 BLEU in the English-Gujarati direction), the aforementioned bidirectional training strategy improves the base model by 1.0 BLEU. Furthermore, applying BackTranslation on top of this bidirectional training strategy yields a +2.8 BLEU improvement, demonstrating that this bidirectional training strategy provides a better base model, enabling the previously ineffective BackTranslation strategy to achieve even better results.
[0098] Table 6 Experimental results in extremely low resource scenarios
[0099]
[0100] Figure 7 This is a flow chart of a machine translation method provided in an embodiment of the present invention. This embodiment is applicable to machine translation scenarios. The method can be performed by a machine translation device provided in an embodiment of the present invention. The device can be implemented using software and / or hardware and can be integrated into an electronic device, such as a user terminal or server.
[0101] See also Figure 7 The method of the embodiment of the present invention specifically includes the following steps:
[0102] S310: Obtain source data to be translated and a machine translation model trained according to the model training method provided in any embodiment of the present invention, wherein the language of the source data to be translated is the same as the language of the original source data in the model training method.
[0103] Among them, the source data to be translated can be data to be translated in the same language as the original source data described above. Since the machine translation model trained above is a model that can machine translate the original source data, the source data to be translated in the same language as the original source data can be machine translated by the machine translation model.
[0104] S320: Input the source data to be translated into the machine translation model, and obtain the translated target data according to the output result of the machine translation model.
[0105] Specifically, for the language corresponding to the original target-end data, the translated target-end data can be understood as the translation of the source-end data to be translated in the language.
[0106] The technical solution of the embodiment of the present invention is that since the machine translation model trained above has good machine translation performance, after the source data to be translated is input into the machine translation model, translated target data with high machine translation accuracy can be obtained, thereby achieving the effect of accurate machine translation.
[0107] Figure 8 This is a structural block diagram of a model training device provided in an embodiment of the present invention, which is used to execute the model training method provided in any of the above embodiments. This device and the model training method in the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the model training device, please refer to the embodiment of the above model training method. Figure 8 The device may specifically include: a corpus acquisition module 410, a bidirectional training module 420 and a forward training module 430.
[0108] The corpus acquisition module 410 is used to acquire original parallel corpora including original source-end data and original target-end data;
[0109] a bidirectional training module 420 for using the original source data as exchange target data and the original target data as exchange source data to obtain exchange parallel corpora, and training the original translation model based on the multiple sets of original parallel corpora and the multiple sets of exchange parallel corpora to obtain an intermediate translation model;
[0110] The forward training module 430 is used to train the intermediate translation model based on multiple sets of original parallel corpora to obtain a machine translation model.
[0111] Optionally, the above-mentioned model training device may further include:
[0112] The original subword representation obtaining module is used to perform word segmentation on the original source data and the original target data in the original parallel corpora before training the original translation model based on the multiple sets of original parallel corpora and the multiple sets of exchanged parallel corpora, and to perform subword segmentation on the obtained word segmentation results to obtain the original subword representation;
[0113] An original parallel corpus updating module, configured to update the original parallel corpus based on original subword representations corresponding to original source-end data and original subword representations corresponding to original target-end data;
[0114] The exchange subword representation obtaining module is used to perform word segmentation on the exchange source end data and the exchange target end data in the exchange parallel corpus, and perform subword segmentation on the obtained word segmentation results to obtain exchange subword representations;
[0115] The exchange parallel corpus updating module is used to update the exchange parallel corpus based on the exchange subword representation corresponding to the exchange source end data and the exchange subword representation corresponding to the exchange target end data.
[0116] Optionally, the above-mentioned model training device may further include:
[0117] an original parallel corpus first enhancement module, configured to, after obtaining the original parallel corpus including the original source-end data and the original target-end data, perform data enhancement on the original source-end data and / or the original target-end data in the original parallel corpus to obtain an enhanced parallel corpus, and use both the enhanced parallel corpus and the original parallel corpus as the original parallel corpus;
[0118] and / or,
[0119] The second enhancement module of the original parallel corpus is used to obtain monolingual source data of the same language as the original source data, and a preliminary translation model that has been preliminarily trained, input the monolingual source data or the original source data into the preliminary translation model to obtain pseudo target data, and use the pseudo parallel corpus composed of the monolingual source data and the pseudo target data, or the original source data and the pseudo target data, and the original parallel corpus as the original parallel corpus.
[0120] Optionally, the above-mentioned model training device may further include:
[0121] A training step ratio acquisition module is used to obtain a preset total number of training steps and a training step ratio for bidirectional training;
[0122] A forward training step obtaining module is used to obtain the bidirectional training steps of the bidirectional training and the forward training steps of the forward training according to the total training steps and the training step ratio;
[0123] The bidirectional training module 420 may include:
[0124] A bidirectional training unit, configured to perform bidirectional training steps on the original translation model based on multiple sets of original parallel corpora and multiple sets of exchanged parallel corpora to obtain an intermediate translation model;
[0125] The forward training module 430 may include:
[0126] The forward training unit is used to train the intermediate translation model for a number of forward training steps based on multiple sets of original parallel corpora to obtain a machine translation model.
[0127] Optionally, the bidirectional loss function matched with the original translation model includes a forward loss function and a reverse loss function; the bidirectional training module 420 may include:
[0128] A forward target end data obtaining unit is used to input the original source end data in the original parallel corpus into the original translation model for each set of original parallel corpus to obtain forward target end data;
[0129] A reverse target-end data obtaining unit is used to input the exchange source-end data in the exchange parallel corpus into the original translation model for each set of exchange parallel corpus to obtain reverse target-end data;
[0130] A reverse loss obtaining unit, configured to combine a forward loss function to obtain a forward loss based on the forward target side data and the original target side data, and to combine a reverse loss function to obtain a reverse loss based on the reverse target side data and the exchanged target side data;
[0131] The intermediate translation model obtaining unit is used to obtain a bidirectional loss according to the forward loss and the reverse loss, and to adjust the network parameters in the original translation model according to the bidirectional loss to obtain the intermediate translation model.
[0132] Optionally, the above-mentioned model training device may further include:
[0133] The translation target data acquisition module is used to obtain the verification parallel corpus after obtaining the machine translation model, input the verification source data in the verification parallel corpus into the machine translation model, and obtain the translation target data based on the output of the machine translation model;
[0134] The machine translation accuracy obtaining module is used to match the translation target end data with the verification target end data in the verification parallel corpus to obtain the machine translation accuracy.
[0135] The model training device provided by the embodiment of the present invention obtains original parallel corpora including original source data and original target data through a corpus acquisition module; uses the original source data as exchange target data and the original target data as exchange source data through a bidirectional training module to obtain exchange parallel corpora, and trains the original translation model based on multiple groups of original parallel corpora and multiple groups of exchange parallel corpora to obtain an intermediate translation model; then, trains the intermediate translation model based on multiple groups of original parallel corpora through a forward training module to obtain a machine translation model. The above-mentioned device solves the problem of insufficient training samples in low-resource scenarios by exchanging the original source data and the original target data and adding the exchange results to the training samples. On this basis, by adding bidirectional training to the forward training, all information in the bilingual data can be fully learned. The two cooperate with each other to improve the model training effect and obtain a machine translation model with better machine translation performance.
[0136] The model training device provided in the embodiment of the present invention can execute the model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0137] It is worth noting that in the embodiment of the above-mentioned model training device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0138] Figure 9 This is a structural block diagram of a machine translation device provided in an embodiment of the present invention. The device is used to execute the machine translation method provided in any of the above embodiments. The device and the machine translation method in the above embodiments belong to the same inventive concept. For details not fully described in the embodiments of the machine translation device, please refer to the embodiments of the above machine translation method. Figure 9 , the device may specifically include: a model acquisition module 510 and a machine translation module 520.
[0139] The model acquisition module 510 is configured to acquire source data to be translated and a machine translation model trained according to the model training method provided in any embodiment of the present invention, wherein the source data to be translated is in the same language as the original source data in the model training method;
[0140] The machine translation module 520 is used to input the source data to be translated into the machine translation model, and obtain the translated target data according to the output result of the machine translation model.
[0141] The machine translation device provided by the embodiment of the present invention cooperates with the model acquisition module and the machine translation module. Since the machine translation model trained above has good machine translation performance, after the source data to be translated is input into the machine translation model, translated target data with high machine translation accuracy can be obtained, thereby achieving the effect of accurate machine translation.
[0142] The machine translation apparatus provided in the embodiment of the present invention can execute the machine translation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0143] It is worth noting that in the embodiment of the above-mentioned machine translation device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be realized; in addition, the specific names of the various functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.
[0144] Figure 10 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0145] like Figure 10 As shown, electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 and a random access memory (RAM) 13, communicatively connected to at least one processor 11. The memory stores a computer program executable by the at least one processor, and processor 11 can perform various appropriate actions and processes according to the computer program stored in read-only memory (ROM) 12 or loaded from storage unit 18 into random access memory (RAM) 13. RAM 13 can also store various programs and data required for the operation of electronic device 10. Processor 11, ROM 12, and RAM 13 are interconnected via bus 14. An input / output (I / O) interface 15 is also connected to bus 14.
[0146] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0147] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model training method or the machine translation method.
[0148] In some embodiments, the model training method or the machine translation method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model training method or the machine translation method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the model training method or the machine translation method in any other appropriate manner (e.g., by means of firmware).
[0149] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0150] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0151] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0153] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0154] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0155] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0156] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A model training method, It is characterized in that include: Obtaining original parallel corpora including original source-end data and original target-end data; Performing word segmentation on the original source-end data and the original target-end data in the original parallel corpus respectively, performing subword segmentation on the obtained word segmentation results, and updating the original parallel corpus based on the obtained original subword representations, wherein the appearance frequency of the original subword representations in the word segmentation results is higher than the appearance frequency of subword representations other than the original subword representations in the word segmentation results; Using the original source-end data as exchange target-end data, and using the original target-end data as exchange source-end data, to obtain an exchange parallel corpus; Performing word segmentation on the exchange source-end data and the exchange target-end data in the exchange parallel corpus respectively, performing subword segmentation on the obtained word segmentation results, and updating the exchange parallel corpus based on the obtained exchange subword representations; The original translation model is trained based on the multiple groups of the original parallel corpora and the multiple groups of the exchanged parallel corpora to obtain an intermediate translation model, and the intermediate translation model is trained based on the multiple groups of the original parallel corpora to obtain a machine translation model.
2. The method according to claim 1, It is characterized in that After obtaining the original parallel corpus including the original source-end data and the original target-end data, the method further includes: Performing data enhancement on the original source-end data and / or the original target-end data in the original parallel corpus to obtain an enhanced parallel corpus, and using both the enhanced parallel corpus and the original parallel corpus as the original parallel corpus; and / or, Obtain monolingual source data of the same language as the original source data, and a preliminary translation model that has been preliminarily trained, input the monolingual source data or the original source data into the preliminary translation model to obtain pseudo target data, and use a pseudo parallel corpus consisting of the monolingual source data and the pseudo target data, or the original source data and the pseudo target data, and the original parallel corpus as the original parallel corpus.
3. The method according to claim 1, It is characterized in that Also includes: Get the preset total number of training steps and the ratio of training steps for bidirectional training; According to the total number of training steps and the ratio of the number of training steps, obtaining the number of bidirectional training steps of the bidirectional training and the number of forward training steps of the forward training; The training of the original translation model based on the multiple groups of the original parallel corpora and the multiple groups of the exchanged parallel corpora to obtain the intermediate translation model includes: Based on the multiple groups of the original parallel corpora and the multiple groups of the exchanged parallel corpora, the original translation model is trained for the two-way training steps to obtain an intermediate translation model; The step of training the intermediate translation model based on the multiple groups of original parallel corpora to obtain a machine translation model includes: The intermediate translation model is trained for the number of forward training steps based on the multiple groups of the original parallel corpora to obtain a machine translation model.
4. The method according to claim 1, It is characterized in that The bidirectional loss function matched with the original translation model includes a forward loss function and a reverse loss function; The training of the original translation model based on the multiple groups of the original parallel corpora and the multiple groups of the exchanged parallel corpora to obtain the intermediate translation model includes: For each group of the original parallel corpora, the original source-end data in the original parallel corpora are input into the original translation model to obtain forward target-end data; For each group of the exchange parallel corpora, the exchange source end data in the exchange parallel corpora are input into the original translation model to obtain reverse target end data; In combination with the forward loss function, a forward loss is obtained based on the forward target end data and the original target end data, and in combination with the reverse loss function, a reverse loss is obtained based on the reverse target end data and the exchange target end data; A bidirectional loss is obtained according to the forward loss and the reverse loss, and network parameters in the original translation model are adjusted according to the bidirectional loss to obtain an intermediate translation model.
5. The method according to claim 1, It is characterized in that After obtaining the machine translation model, the method further includes: Acquire verification parallel corpus, input verification source data in the verification parallel corpus into the machine translation model, and obtain translation target data according to the output result of the machine translation model; The translation target end data is matched with the verification target end data in the verification parallel corpus to obtain the machine translation accuracy.
6. A machine translation method, It is characterized in that include: Obtain source data to be translated, and a machine translation model trained according to the model training method of any one of claims 1 to 5, wherein the language of the source data to be translated is the same as the language of the original source data in the model training method; The source data to be translated is input into the machine translation model, and the translated target data is obtained according to the output result of the machine translation model.
7. A model training device, It is characterized in that include: A corpus acquisition module, used to acquire original parallel corpus including original source-end data and original target-end data; A bidirectional training module, used to use the original source end data as exchange target end data, and use the original target end data as exchange source end data to obtain exchange parallel corpora, and train the original translation model based on multiple groups of the original parallel corpora and multiple groups of the exchange parallel corpora to obtain an intermediate translation model; A forward training module, used for training the intermediate translation model based on multiple groups of the original parallel corpora to obtain a machine translation model; The model training device further comprises: An original subword representation obtaining module is used for, before training the original translation model based on the multiple groups of the original parallel corpora and the multiple groups of the exchanged parallel corpora, respectively performing word segmentation on the original source-end data and the original target-end data in the original parallel corpora, and performing subword segmentation on the obtained word segmentation results to obtain original subword representations, wherein the appearance frequency of the original subword representations in the word segmentation results is higher than the appearance frequency of subword representations other than the original subword representations in the word segmentation results; an original parallel corpus updating module, configured to update the original parallel corpus based on the original subword representation corresponding to the original source-end data and the original subword representation corresponding to the original target-end data; An exchange subword representation obtaining module is used to perform word segmentation on the exchange source-end data and the exchange target-end data in the exchange parallel corpus, and perform subword segmentation on the obtained word segmentation results to obtain exchange subword representations; The exchange parallel corpus updating module is used to update the exchange parallel corpus based on the exchange sub-word representation corresponding to the exchange source end data and the exchange sub-word representation corresponding to the exchange target end data.
8. A machine translation device, It is characterized in that include: A model acquisition module, used to acquire source data to be translated and a machine translation model trained according to the model training method of any one of claims 1 to 5, wherein the language of the source data to be translated is the same as the language of the original source data in the model training method; The machine translation module is used to input the source data to be translated into the machine translation model, and obtain the translated target data according to the output result of the machine translation model.
9. An electronic device, It is characterized in that include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor performs the model training method as described in any one of claims 1-5, or the machine translation method as described in claim 6.
10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a processor to implement the model training method as described in any one of claims 1 to 5, or the machine translation method as described in claim 6 when executed.
Citation Information
Patent Citations
Machine translation model training method, device and system
CN110941966A
Machine translation model training method, language translation method and equipment
CN113705251A
Dictionary generation method and device, storage medium and electronic device
CN114611496A