Machine translation model training method, device, medium and electronic device
By using high-resource corpus to update the vocabulary list of low-resource corpus and adding training data sets, the data sparse problem caused by the scarcity of corpus resources in neural machine translation model training is solved, and the translation quality is improved.
Patent Information
- Application Number
- CN202110920879.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-08-11
AI Technical Summary
In the prior art, the neural machine translation model has a lack of corpus resources during training, resulting in sparse data problems, which in turn affects the translation quality.
By obtaining the training data set and candidate data sets of high-resource corpus and low-resource corpus, the vocabulary of low-resource corpus is updated with the vocabulary of high-resource corpus and calculate the update completion degree, and the low-resource corpus with completion degree greater than the preset threshold value is added to the training data set to continue training the machine translation model.
It effectively avoids the problem of sparse data during training of low-resource corpus, and improves the translation effect of low-resource corpus corresponding to the language to be trained.
Smart Images

Figure CN114358275B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to a method for training a machine translation model, an apparatus for training a machine translation model, a computer-readable medium, and an electronic device. Background Art
[0002] A general neural machine translation model is based on an end-to-end Encoder-Decoder framework. Generally, a large amount of parallel corpus is required for model training. Due to the high cost of manual annotation, for languages with scarce corpus resources, it is impossible to obtain a large amount of parallel corpus, resulting in poor translation quality of the neural machine translation system for languages with scarce corpus resources.
[0003] High-quality parallel corpus often exists only between a small number of languages. For some languages lacking resources, it is difficult to find or obtain available parallel corpus from the Internet. The lack of corpus resources for languages will cause the problem of data sparsity during model training, resulting in poor translation effects of the model for languages with scarce corpus resources. How to improve the translation effect of the machine translation model is one of the problems that need to be solved urgently.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of this application is to provide a method for training a machine translation model, an apparatus for training a machine translation model, a computer-readable medium, and an electronic device, which can at least overcome to some extent the technical problem of how to improve the translation effect of the machine translation model in related technologies.
[0006] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0007] According to one aspect of the embodiments of this application, a method for training a machine translation model is provided. The method for training a machine translation model includes:
[0008] Obtain a training data set and a candidate data set; the training data set includes one or more high-resource corpora and the word list corresponding to each high-resource corpus, the candidate data set includes one or more low-resource corpora and the word list corresponding to each low-resource corpus, the data volume of the high-resource corpus is greater than that of the low-resource corpus, and the word list is used to represent the corresponding relationship between word segmentation and word vectors;
[0009] Train the machine translation model using the corpus in the training dataset, and update the vocabulary corresponding to each piece of corpus in the training dataset during the training process;
[0010] Update the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus, and calculate the update completion degree of the vocabulary corresponding to the low-resource corpus, where the update completion degree is used to represent the progress of updating the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus;
[0011] Add the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold to the training dataset.
[0012] According to one aspect of the embodiments of the present application, a machine translation model training device is provided. The machine translation model training device includes:
[0013] A corpus acquisition module configured to acquire a training dataset and a candidate dataset; the training dataset includes one or more high-resource corpora and the vocabulary corresponding to each high-resource corpus, the candidate dataset includes one or more low-resource corpora and the vocabulary corresponding to each low-resource corpus, the data volume of the high-resource corpus is greater than that of the low-resource corpus, and the vocabulary is used to represent the corresponding relationship between word segmentation and word vectors;
[0014] A model training module configured to train the machine translation model using the corpus in the training dataset, and update the vocabulary corresponding to each piece of corpus in the training dataset during the training process;
[0015] A vocabulary update module configured to update the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus, and calculate the update completion degree of the vocabulary corresponding to the low-resource corpus, where the update completion degree is used to represent the progress of updating the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus;
[0016] A training dataset addition module configured to add the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold to the training dataset.
[0017] In some embodiments of the present application, based on the above technical solutions, the model training module includes:
[0018] A sampling unit configured to sample each piece of corpus in the training dataset according to the sampling weight corresponding to each piece of corpus in the training dataset;
[0019] A model training unit configured to train the machine translation model according to the sampled corpus;
[0020] A sampling weight update unit, configured to calculate the training completion degree corresponding to each piece of corpus in the training dataset during the training process, and update the sampling weight of the corpus according to the training completion degree, where the sampling weight is inversely proportional to the training completion degree.
[0021] In some embodiments of the present application, based on the above technical solutions, the sampling weight update unit includes:
[0022] A first loss function acquisition subunit, configured to acquire a preset first loss function value for indicating the convergence of the machine translation model training;
[0023] A second loss function acquisition subunit, configured to input the validation set corresponding to each piece of corpus in the training dataset into the machine translation model to obtain the second loss function value corresponding to each piece of corpus in the training dataset;
[0024] A training completion degree calculation subunit, configured to calculate the difference between the first loss function value and the second loss function value, and calculate the training completion degree corresponding to each piece of corpus in the training dataset according to the difference.
[0025] In some embodiments of the present application, based on the above technical solutions, the vocabulary update module includes:
[0026] A first same word segmentation quantity acquisition unit, configured to acquire the quantity of the same high-frequency word segmentations in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus, where the high-frequency word segmentations are the word segmentations with the occurrence frequency within a preset ranking in the corpus;
[0027] A first similarity determination unit, configured to respectively determine the similarity between each piece of the high-resource corpus and the low-resource corpus according to the ratio of the word segmentation quantity to the preset ranking;
[0028] A first update completion degree calculation unit, configured to use the training completion degree corresponding to the high-resource corpus with the highest similarity to the low-resource corpus as the update completion degree of the vocabulary corresponding to the low-resource corpus.
[0029] In some embodiments of the present application, based on the above technical solutions, the vocabulary update module further includes:
[0030] A second same word segmentation quantity acquisition unit, configured to acquire the same high-frequency word segmentations in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus, where the high-frequency word segmentations are the word segmentations with the occurrence frequency within a preset ranking in the corpus;
[0031] The second similarity determination unit is configured to determine the similarity between each high-resource corpus and the low-resource corpus respectively according to the ratio of the number of word segments to the preset ranking.
[0032] The second update completion degree calculation unit is configured to use the similarity between each high-resource corpus and the low-resource corpus as a weight, and perform a weighted summation operation on the training completion degrees corresponding to each high-resource corpus to obtain the update completion degree of the word list corresponding to the low-resource corpus.
[0033] In some embodiments of the present application, based on the above technical solution, the word list update module further includes:
[0034] The third same word segment number acquisition unit is configured to acquire the same high-frequency word segments in the word list corresponding to the low-resource corpus and the word list corresponding to the high-resource corpus, and the high-frequency word segments are word segments whose occurrence frequencies in the corpus are within the preset ranking.
[0035] The third similarity determination unit is configured to determine the similarity between each high-resource corpus and the low-resource corpus respectively according to the ratio of the number of word segments to the preset ranking.
[0036] The activation weight calculation unit is configured to calculate the activation weight S corresponding to each high-resource corpus through the following calculation formula i :
[0037]
[0038] wherein, v i and v j are both the similarities between the high-resource corpus and the low-resource corpus, and e is the natural logarithm;
[0039] The third update completion degree calculation unit is configured to perform a weighted summation operation on the training completion degrees corresponding to each high-resource corpus according to the activation weight S i to obtain the update completion degree of the word list corresponding to the low-resource corpus.
[0040] In some embodiments of the present application, based on the above technical solution, the training completion degree calculation subunit includes:
[0041] The specific training completion degree calculation subunit is configured to calculate the training completion degree c corresponding to each corpus in the training dataset through the following calculation formula:
[0042]
[0043] wherein, L * is the value of the first loss function, and L is the value of the second loss function.
[0044] In some embodiments of the present application, based on the above technical solutions, the sampling weight update unit further includes:
[0045] A sampling weight update subunit, configured to calculate the training completion degree corresponding to each piece of corpus in the training dataset at every preset number of steps during the training process, where the step is a preset data volume for inputting the corpus into the machine translation model.
[0046] In some embodiments of the present application, based on the above technical solutions, the machine translation model training device further includes:
[0047] A low-resource corpus vocabulary acquisition unit, configured to perform word segmentation and deduplication on each piece of the low-resource corpus to obtain multiple word segments corresponding to the low-resource corpus, and assign initial word vectors to the multiple word segments corresponding to the low-resource corpus to form a vocabulary corresponding to each piece of the low-resource corpus;
[0048] A high-resource corpus vocabulary acquisition unit, configured to perform word segmentation and deduplication on each piece of the high-resource corpus to obtain multiple word segments corresponding to the high-resource corpus, and assign initial word vectors to the multiple word segments corresponding to the high-resource corpus to form a vocabulary corresponding to each piece of the high-resource corpus.
[0049] In some embodiments of the present application, based on the above technical solutions, the vocabulary update module further includes:
[0050] A coincident word segment acquisition unit, configured to obtain the same word segments in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus as the coincident word segments;
[0051] A word vector update unit, configured to update the coincident word segments and the corresponding word vectors in the vocabulary corresponding to the low-resource corpus according to the coincident word segments and the corresponding word vectors in the vocabulary corresponding to the high-resource corpus.
[0052] In some embodiments of the present application, based on the above technical solutions, the model training module further includes:
[0053] A first word vector update unit, configured to update the word vectors corresponding to the word segments in the vocabulary corresponding to each piece of corpus in the training dataset during the training process to obtain updated word vectors;
[0054] The vocabulary update module further includes:
[0055] A second word vector update unit, configured to update the vocabulary corresponding to the low-resource corpus by using the updated word vectors in the vocabulary corresponding to the high-resource corpus.
[0056] In some embodiments of the present application, based on the above technical solutions, the machine translation model is used to translate the central language into one or more languages, or to translate one or more of the languages into the central language;
[0057] The high-resource corpus includes bilingual parallel corpus corresponding to the preset language and the central language, and the low-resource corpus includes bilingual parallel corpus of the preset language and the central language in a preset domain.
[0058] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the machine translation model training method in the above technical solutions.
[0059] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the machine translation model training method in the above technical solutions by executing the executable instructions.
[0060] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the machine translation model training method in the above technical solutions.
[0061] In the technical solution provided by the embodiments of the present application, first, the machine translation model is trained with the corpus in the training dataset, and the vocabulary corresponding to each corpus in the training dataset is updated during the training process. Then, the vocabulary corresponding to the low-resource corpus is updated with the vocabulary corresponding to the high-resource corpus, and the update completion degree of the vocabulary corresponding to the low-resource corpus is calculated. The low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold is added to the training dataset to continue training the machine translation model with the corpus in the training dataset. Thus, the vocabulary of the low-resource corpus is updated first, and then the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold is added to the training dataset to train the machine translation model with the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold, so that when the amount of corpus resource data of the low-resource corpus is small, the vocabulary of the low-resource corpus can be updated with the vocabulary of the high-resource corpus to avoid the problem of data sparsity during the training of the low-resource corpus, and further improve the training effect of the language to be trained corresponding to the low-resource corpus.
[0062] Moreover, it can be understood that since the machine translation model is first trained with high-resource corpora in the training dataset, and then the low-resource corpus is added to the training dataset to train the machine translation model after the update completion degree of the vocabulary of the low-resource corpus is greater than a preset threshold, the total training duration of the high-resource corpus will be greater than that of the low-resource corpus. As a result, the training process of the model can make full use of the rich corpora of the high-resource corpus, and can improve the training effect of the model for the language to be trained corresponding to the high-resource corpus.
[0063] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings
[0064] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0065] Figure 1 Schematically shows an exemplary device architecture block diagram applying the technical solution of this application.
[0066] Figure 2 Schematically shows the step flow of the machine translation model training method provided by the embodiment of this application.
[0067] Figure 3 Schematically shows a relationship schematic diagram of high-resource corpora and low-resource corpora in the embodiment of this application.
[0068] Figure 4 Schematically shows the step flow of training the machine translation model with the corpora in the training dataset in the embodiment of this application.
[0069] Figure 5 Schematically shows the step flow of calculating the training completion degree corresponding to each corpus in the training dataset during the training process in the embodiment of this application.
[0070] Figure 6 Schematically shows the step flow before updating the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus in the embodiment of this application.
[0071] Figure 7 Schematically shows the step flow of updating the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus in the embodiment of this application.
[0072] Figure 8Schematically shows the step - by - step process of calculating the update completion degree of the vocabulary corresponding to the low - resource corpus in the embodiments of the present application.
[0073] Figure 9 Schematically shows the step - by - step process of calculating the update completion degree of the vocabulary corresponding to the low - resource corpus in the embodiments of the present application.
[0074] Figure 10 Schematically shows the step - by - step process of calculating the update completion degree of the vocabulary corresponding to the low - resource corpus in the embodiments of the present application.
[0075] Figure 11 Schematically shows the structural block diagram of the machine translation model training device provided in the embodiments of the present application.
[0076] Figure 12 Schematically shows the structural block diagram of the electronic device for implementing the embodiments of the present application. Detailed implementation manners
[0077] Now, example embodiments will be described more comprehensively with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.
[0078] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well - known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0079] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0080] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0081] Before elaborating on the technical solutions such as the machine translation model training method and the machine translation model training device provided in the embodiments of the present application, a brief introduction to the artificial intelligence technology and cloud technology involved in some embodiments of the present application will be given first.
[0082] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0083] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0084] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0085] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0086] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.
[0087] Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely expandable, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0088] As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as cloud platform, generally called IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0089] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth rate, and diverse information asset that requires a new processing mode to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process a large amount of data tolerated over time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.
[0090] The solution provided in the embodiments of this application involves the above-related technologies of artificial intelligence and cloud technology, etc., and will be specifically described through the following embodiments.
[0091] The machine translation model training method and device provided in this application will be described in detail below in combination with specific implementation manners.
[0092] Figure 1 An exemplary device architecture block diagram applying the technical solution of this application is schematically shown.
[0093] As Figure 1As shown, the device architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include various electronic devices such as a smartphone, a tablet computer, a laptop computer, and a desktop computer. The server 130 may be an independent physical server, a server cluster or a distributed device composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types capable of providing a communication link between the terminal device 110 and the server 130. For example, it may be a wired communication link or a wireless communication link.
[0094] According to implementation requirements, the device architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 130 may be a server group composed of multiple server devices. In addition, the technical solutions provided in the embodiments of the present application may be applied to the terminal device 110, may also be applied to the server 130, or may be jointly implemented by the terminal device 110 and the server 130. The present application does not make special limitations on this.
[0095] For example, when the terminal device 110 accesses the cloud platform of the server 130, the terminal device 110 may execute the machine translation model training method provided in the present application. First, update the vocabulary of the low-resource corpus, and then add the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold to the training data set, so as to train the machine translation model with the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold. Thus, when the amount of corpus resource data of the low-resource corpus is small, it is possible to update the vocabulary of the low-resource corpus with the vocabulary of the high-resource corpus, so as to avoid the problem of data sparsity in the training process of the low-resource corpus, and further improve the training effect of the language to be trained corresponding to the low-resource corpus.
[0096] In the training of the multi-language translation machine learning model, since the corpus scale corresponding to some languages is large and the corpus scale corresponding to some languages is small, considering the balance of training for different languages, the sampling of the language corpus with a large corpus scale can be reduced, and the sampling of the language corpus with a large corpus scale can be increased. However, doing so may lead to insufficient utilization of the language corpus with a large corpus scale, resulting in the problem that the translation effect of the language with a large corpus scale of the multi-language translation machine learning model is not as good as that of the bilingual translation machine learning model.
[0097] It can be understood that since the present application first trains the machine translation model with the high-resource corpus in the training dataset, and then adds the low-resource corpus to the training dataset to train the machine translation model after the update completion degree of the vocabulary of the low-resource corpus is greater than the preset threshold, therefore, the total training duration of the high-resource corpus will be greater than that of the low-resource corpus, so that the training process of the model can make full use of the rich corpus of the high-resource corpus and improve the training effect of the model on the target training language corresponding to the high-resource corpus.
[0098] In some embodiments, the machine translation model training method of the embodiments of the present application can be implemented by a target terminal.
[0099] Figure 2 Schematically shows the step flow of the machine translation model training method provided by the embodiments of the present application. The execution subject of this machine translation model training method can be a terminal device or a server. As Figure 2 shown, this machine translation model training method mainly may include the following steps S210 to step S240:
[0100] S210. Obtain a training dataset and a candidate dataset. The training dataset includes one or more high-resource corpora and the vocabulary corresponding to each high-resource corpus. The candidate dataset includes one or more low-resource corpora and the vocabulary corresponding to each low-resource corpus. The data volume of the high-resource corpus is greater than that of the low-resource corpus, and the vocabulary is used to represent the corresponding relationship between word segmentation and word vectors.
[0101] The corpus in the training dataset can be directly used for training the machine translation model. The corpus in the candidate dataset can be used for training the machine translation model after meeting the preset conditions. The preset condition can be that the update completion degree of the vocabulary corresponding to the corpus is greater than the preset threshold. In some examples, the preset threshold can be 0.75, 0.8, 0.85, 0.9, 0.95, etc.
[0102] The high-resource corpus can be a corpus with rich corpus resources and a large data volume. The low-resource corpus can be a corpus with scarce corpus resources and a small data volume.
[0103] The vocabulary can include multiple word segmentations and the word vectors corresponding to each word segmentation. The word segmentation can be a word, a sub-word, a phrase or a character, etc., and the word vector is used to represent the characteristics of the corresponding word segmentation.
[0104] In certain embodiments, the machine translation model is used to translate the central language into multiple languages, or to translate multiple languages into the central language. The high-resource corpus includes the bilingual parallel corpus of the first preset language and the central language, and the low-resource corpus includes the bilingual parallel corpus of the second preset language and the central language, and the first preset language is different from the second preset language.
[0105] The central language is the source language or the target language in a machine translation model. The central language can be any preset language, and this application does not limit it. Figure 3 Schematically shows a schematic diagram of the relationship between high-resource corpora and low-resource corpora in an embodiment of this application. For example, as Figure 3 shown in the bipartite graph, the central language is English. The high-resource corpora include Turkish-English bilingual corpora, Russian-English bilingual corpora, Portuguese-English bilingual corpora, and Czech-English bilingual corpora. The low-resource corpora include Azerbaijani-English corpora, Belarusian-English corpora, Galician-English corpora, and Slovak-English corpora. The machine translation model is used to translate English into Turkish, Russian, Portuguese, Czech, Azerbaijani, Belarusian, Galician, and Slovak respectively; or, the machine translation model is used to translate Turkish, Russian, Portuguese, Czech, Azerbaijani, Belarusian, Galician, and Slovak into English respectively.
[0106] In some embodiments, the machine translation model is used to translate the central language into one or more languages, or to translate one or more languages into the central language;
[0107] The high-resource corpora include bilingual corpora corresponding to the preset language and the central language, and the low-resource corpora include bilingual corpora of the preset language and the central language in the preset domain.
[0108] Taking the bilingual corpus of the preset language in the preset domain as the low-resource corpus and taking the bilingual corpus of the whole domain of the preset language, that is, the non-fixed domain, as the high-resource corpus can update the vocabulary of the bilingual corpus of the preset language in the preset domain with the vocabulary of the bilingual corpus of the preset language in the whole domain, so as to avoid the problem of data sparsity in the training process of the bilingual corpus of the preset language in the preset domain with less resources, and then improve the training effect on the language to be trained corresponding to the bilingual corpus of the preset language in the preset domain with less resources.
[0109] The central language is the source language or the target language in a machine translation model. The central language can be any pre-set language, and this application does not limit it. For example, the central language can be Chinese, and the pre-set language can be English. The pre-set fields can include the medical field or the chip field. The high-resource corpus can include English-Chinese bilingual parallel corpus, German-Chinese bilingual parallel corpus, Arabic-Chinese bilingual parallel corpus, and French-Chinese bilingual parallel corpus. The low-resource corpus can include English-Chinese parallel corpus in the medical field, English-Chinese parallel corpus in the chip field, Azerbaijani-Chinese parallel corpus, and Belarusian-Chinese parallel corpus. The machine translation model is used to translate English, German, Arabic, French, English in the medical field, English in the chip field, Azerbaijani, and Belarusian into Chinese respectively; or, the machine translation model is used to translate Chinese into English, German, Arabic, French, English in the medical field, English in the chip field, Azerbaijani, and Belarusian respectively.
[0110] S220. Train the machine translation model using the corpus in the training dataset, and update the vocabulary corresponding to each corpus in the training dataset during the training process.
[0111] In the initial state of training, the training dataset can only include high-resource corpus. Input the high-resource corpus in the training dataset into the machine translation model to implement the training of the machine translation model. The machine translation model can be a machine translation model based on a deep conversion encoder or based on a neural network.
[0112] During the training of the machine translation model, since the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold in S240 below is added to the training dataset, the training dataset can include both high-resource corpus and low-resource corpus with the update completion degree greater than the preset threshold. Input the high-resource corpus and the low-resource corpus with the update completion degree greater than the preset threshold into the machine translation model to implement the training of the machine translation model.
[0113] During the process of training the machine translation model using the corpus, the vocabulary corresponding to each corpus in the training dataset can be updated, so as to optimize the word vectors of the vocabulary corresponding to each corpus, and the training effect of the model can be improved.
[0114] Figure 4 Schematically shows the step flow of training the machine translation model using the corpus in the embodiments of the present application. As Figure 4 shown, on the basis of the above embodiments, in some embodiments, training the machine translation model using the corpus in step S220 can further include the following steps S410 to step S430:
[0115] S410. Sample each piece of corpus in the training dataset according to the sampling weights corresponding to each piece of corpus in the training dataset.
[0116] S420. Train the machine translation model according to the sampled corpus.
[0117] S430. During the training process, calculate the training completion degree corresponding to each piece of corpus in the training dataset, and update the sampling weights of the corpus according to the training completion degree. The sampling weights are inversely proportional to the training completion degree.
[0118] Since the training dataset may include multiple pieces of corpus, when training the machine translation model using the corpus in the training dataset, each piece of corpus in the training dataset can be sampled according to the sampling weights corresponding to each piece of corpus in the training dataset first, and then the machine translation model can be trained according to the sampled corpus. Specifically, in the initial state of training, a preset initial value can be used as the sampling weights corresponding to each piece of corpus in the first corpus dataset. In some examples, in the initial state of training, the sampling weights corresponding to each piece of corpus can be the same, and the sampling weights corresponding to each piece of corpus can be the ratio of a preset value such as 1, 2, 10, 20, 100, etc. to the total number of pieces of corpus in the training dataset.
[0119] During the training process, calculate the training completion degree corresponding to each piece of corpus in the training dataset, and update the sampling weights of the corpus according to the training completion degree, so that the sampling weights are inversely proportional to the training completion degree. Thereby, it is possible to reduce the sampling of the corpus with a high training completion degree and increase the sampling of the corpus with a low training completion degree, so as to achieve the training balance among each piece of corpus, and further balance the training completion degree of the machine translation model for the languages corresponding to each piece of corpus, and improve the overall translation effect of the model.
[0120] Based on the above embodiments, in some embodiments, calculating the training completion degree corresponding to each piece of corpus in the training dataset in step S430 during the training process may further include the following steps:
[0121] During the training process, calculate the training completion degree corresponding to each piece of corpus in the training dataset every preset number of steps. Wherein, the step size is the preset data volume for inputting the corpus into the machine translation model.
[0122] At every preset number of steps, calculate the training completion degree corresponding to each corpus in the training dataset, and update the sampling weight of the corpus according to the training completion degree, so that the sampling weight of the corpus is inversely proportional to the training completion degree. Thus, the model can update the sampling weight of the corpus once every preset number of steps, thereby realizing dynamic sampling, making the sampling weight of the corpus inversely proportional to the training completion degree, and thus achieving training balance among various corpora and improving the overall translation effect of the model.
[0123] In a specific example, the step size can be the data volume of 100, 200, or 500 pieces of corpus data. Alternatively, the step size can be the data volume of corpus data with a size of 100KB, 200KB, 300KB, or 500KB.
[0124] Based on the above embodiments, in some embodiments, the step of updating the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus and calculating the update completion degree of the vocabulary corresponding to the low-resource corpus in step S230 may include:
[0125] During the training process, at every preset number of steps, update the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus, and calculate the update completion degree of the vocabulary corresponding to the low-resource corpus.
[0126] Thus, the update completion degree can be updated in a timely manner according to the preset period, and then the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold can be added to the training dataset in a timely manner.
[0127] Furthermore, when calculating the training completion degree corresponding to each corpus in the training dataset at every preset number of steps, the vocabulary corresponding to the low-resource corpus can be updated with the vocabulary corresponding to the high-resource corpus, and the update completion degree of the vocabulary corresponding to the low-resource corpus can be calculated. Thus, both the training completion degree and the update completion degree can be updated according to the preset period, and then the low-resource corpus with the update completion degree of the vocabulary greater than the preset threshold can be added to the training dataset in a timely manner.
[0128] Figure 5 Schematically shows the step flow of calculating the training completion degree corresponding to each corpus in the training dataset in the embodiments of the present application during the training process. As Figure 5 shown, based on the above embodiments, in some embodiments, the step of calculating the training completion degree corresponding to each corpus in the training dataset in step S430 may further include the following steps S510 to S530:
[0129] S510. Obtain a preset first loss function value for indicating the convergence of the machine translation model training.
[0130] S520. Input the validation set corresponding to each corpus in the training dataset into the machine translation model to obtain the second loss function value corresponding to each corpus in the training dataset.
[0131] S530. Calculate the difference between the first loss function value and the second loss function value, and calculate the training completion degree corresponding to each corpus in the training dataset according to the difference.
[0132] The first loss function value used to indicate the convergence of the machine translation model training can be a preset reference value, which is related to the specific type of the model. Inputting the validation set corresponding to each corpus in the training dataset into the machine translation model can obtain the second loss function value corresponding to each corpus in the training dataset in the current state of the machine translation model. Calculate the difference between the first loss function value and the second loss function value, and the training completion degree corresponding to each corpus in the training dataset can be calculated according to the difference.
[0133] Based on the above embodiments, in some embodiments, calculating the training completion degree corresponding to each corpus in the training dataset according to the difference in step S530 may further include the following steps:
[0134] Calculate the training completion degree c corresponding to each corpus in the training dataset through the following calculation formula:
[0135]
[0136] where, L * is the first loss function value, and L is the second loss function value.
[0137] In some embodiments, it is also possible to first obtain a preset first likelihood score used to indicate the convergence of the machine translation model training, that is, the reference value of the likelihood score. Then input the validation set corresponding to each corpus in the training dataset into the machine translation model, and the second likelihood score corresponding to each corpus in the training dataset in the current state of the machine translation model can be obtained. Calculate the ratio of the second likelihood score to the first likelihood score, and calculate the training completion degree c corresponding to each corpus in the training dataset according to the ratio. The specific calculation method can be as shown in the following calculation formula:
[0138]
[0139] where, M * is the first likelihood score, and M is the second likelihood score.
[0140] It should be noted that the first likelihood score M * and the first loss function value L* may have the following calculation relationship:
[0141]
[0142] The second likelihood score M and the second function loss value L may have the following calculation relationship:
[0143]
[0144] where p i may be the true probability distribution of the validation set, or the true probability distribution of the validation set after label smoothing is applied to the validation set. q i may be the probability distribution output by the machine translation model when the validation set is input into the machine translation model.
[0145] In some embodiments, it is also possible to obtain the BLEU (Bilingual Evaluation Understudy, an evaluation metric for machine translation) value of the machine translation model in its current state when the validation sets corresponding to each piece of corpus in the training dataset are input into the machine translation model, and then calculate the ratio between the BLEU value and the target value of BLEU, and use the ratio as the training completion degree corresponding to each piece of corpus in the training dataset.
[0146] S230. Update the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus, and calculate the update completion degree of the vocabulary corresponding to the low-resource corpus. The update completion degree is used to represent the progress of updating the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus.
[0147] Updating the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus enables the word vectors corresponding to the word segmentations in the vocabulary corresponding to the high-resource corpus to be shared into the vocabulary corresponding to the low-resource corpus. Thus, the vocabulary corresponding to the low-resource corpus can be initialized through the vocabulary corresponding to the high-resource corpus, realizing curriculum learning of the vocabulary corresponding to the low-resource corpus with respect to the vocabulary corresponding to the high-resource corpus. In subsequent steps, the low-resource corpus with an update completion degree of the vocabulary greater than a preset threshold is added to the training dataset, thereby avoiding the problem of data sparsity during the training of the low-resource corpus and being beneficial to improving the training effect of the language to be trained corresponding to the low-resource corpus.
[0148] In a specific embodiment, please continue to refer to Figure 3, the state (a) is the initial state of model training. At this time, the training dataset includes four high-resource corpora: Turkish-English bilingual corpus, Russian-English bilingual corpus, Portuguese-English bilingual corpus, and Czech-English bilingual corpus. The candidate dataset includes four low-resource corpora: Azerbaijani-English corpus, Belarusian-English corpus, Galician-English corpus, and Slovak-English corpus. On the right side of the state (a), the update completion degree of the vocabulary list corresponding to the Azerbaijani-English corpus is 0.0, the update completion degree of the vocabulary list corresponding to the Belarusian-English corpus is 0.0, the update completion degree of the vocabulary list corresponding to the Galician-English corpus is 0.0, and the update completion degree of the vocabulary list corresponding to the Slovak-English corpus is 0.0.
[0149] After the training starts with the corpora in the training dataset to train the machine translation model, during the training process, update the vocabulary lists corresponding to each corpus in the training dataset, and use the vocabulary lists corresponding to the high-resource corpora to update the vocabulary lists corresponding to the low-resource corpora, and calculate the update completion degree of the vocabulary lists corresponding to the low-resource corpora. In the state (b) in the middle of the model training, at this time, the update completion degree of the vocabulary list corresponding to the Azerbaijani-English corpus is 0.69, the update completion degree of the vocabulary list corresponding to the Belarusian-English corpus is 0.65, the update completion degree of the vocabulary list corresponding to the Galician-English corpus is 0.71, and the update completion degree of the vocabulary list corresponding to the Slovak-English corpus is 0.81. In this embodiment, the preset threshold can be 0.8. Since the update completion degree of the vocabulary list corresponding to the Slovak-English corpus is greater than the preset threshold, the Slovak-English corpus is added to the training dataset.
[0150] In the state (c) in the later stage of the model training, at this time, the update completion degree of the vocabulary list corresponding to the Azerbaijani-English corpus is 0.88, the update completion degree of the vocabulary list corresponding to the Belarusian-English corpus is 0.76, the update completion degree of the vocabulary list corresponding to the Galician-English corpus is 0.85, and the update completion degree of the vocabulary list corresponding to the Slovak-English corpus is 0.96. Except for the Slovak-English corpus that has been added to the training dataset, since the update completion degrees of the vocabulary lists corresponding to the Azerbaijani-English corpus and the Galician-English corpus are both greater than the preset threshold, the Azerbaijani-English corpus and the Galician-English corpus are both added to the training dataset.
[0151] In the (d) state at the end of model training, at this time, the update completion degree of the vocabulary corresponding to the Azerbaijani-English parallel corpus is 0.98, the update completion degree of the vocabulary corresponding to the Belarusian-English parallel corpus is 0.90, the update completion degree of the vocabulary corresponding to the Galician-English parallel corpus is 0.95, and the update completion degree of the vocabulary corresponding to the Slovak-English parallel corpus is 1.00. In addition to the Slovak-English parallel corpus, Azerbaijani-English parallel corpus, and Galician-English corpus that have been added to the training dataset, since the update completion degree of the vocabulary corresponding to the Belarusian-English parallel corpus is also greater than the preset threshold, the Belarusian-English parallel corpus was added to the training dataset.
[0152] It should be noted that in some embodiments, the update completion degree of the vocabulary corresponding to the low-resource corpus can also be greater than 1, depending on the specific calculation methods of the update completion degree and the training completion degree.
[0153] Figure 3 The shown vocabulary update direction is from the high-resource corpus to the low-resource corpus. That is to say, the vocabulary corresponding to the high-resource corpus is used to update the vocabulary corresponding to the low-resource corpus. And, in Figure 3 In the shown embodiment, any high-resource corpus updates the vocabulary of all low-resource corpora to initialize the low-resource corpora with the high-resource corpus, enabling the sharing of word vectors of the vocabulary of the high-resource corpus into the low-resource corpora, avoiding the problem of data sparsity during the training process of the low-resource corpora, and being beneficial to improving the training effect of the to-be-trained language corresponding to the low-resource corpora.
[0154] In some implementation manners, based on the above embodiments, updating the vocabulary corresponding to each corpus in the training dataset during the training process in step S220 may further include the following steps:
[0155] Update the word vectors corresponding to the word segments in the vocabulary corresponding to each corpus in the training dataset during the training process to obtain the updated word vectors.
[0156] And updating the vocabulary corresponding to the low-resource corpus with the vocabulary corresponding to the high-resource corpus in step S230 may further include the following steps:
[0157] Update the vocabulary corresponding to the low-resource corpus with the updated word vectors in the vocabulary corresponding to the high-resource corpus.
[0158] That is to say, the word vectors of the segmented words updated in model training using the vocabulary corresponding to the high-resource corpus can be used to update the vocabulary of the low-resource corpus, while the word vectors of the segmented words that have not been updated in model training using the vocabulary corresponding to the high-resource corpus will not be used to update the vocabulary of the low-resource corpus. It can be understood that the word vectors of the segmented words that have not been updated in model training using the vocabulary corresponding to the high-resource corpus are initial word vectors that have not been optimized during the model training process, and the representation of these initial word vectors is not as accurate as that of the word vectors of the segmented words updated in model training. Using the updated word vectors in the vocabulary corresponding to the high-resource corpus to update the vocabulary corresponding to the low-resource corpus can improve the update efficiency of updating the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus and, at the same time, improve the representational accuracy of the vocabulary of the low-resource corpus, thus ensuring the translation effect of the machine translation model.
[0159] Figure 6 Schematically shows the step flow before updating the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus in the embodiments of the present application. As Figure 6 shown, based on the above embodiments, in some embodiments, before updating the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus in step S230, the following steps S610 to S620 may be further included:
[0160] S610. Segment and deduplicate each low-resource corpus to obtain multiple segmented words corresponding to the low-resource corpus, and assign initial word vectors to the multiple segmented words corresponding to the low-resource corpus to form a vocabulary corresponding to each low-resource corpus respectively.
[0161] S620. Segment and deduplicate each high-resource corpus to obtain multiple segmented words corresponding to the high-resource corpus, and assign initial word vectors to the multiple segmented words corresponding to the high-resource corpus to form a vocabulary corresponding to each high-resource corpus respectively.
[0162] Specifically, after obtaining each corpus, the text of the corpus can be segmented. A tokenizer can be used to segment the text of the corpus. After obtaining the segmented words corresponding to each corpus, duplicate segmented words in the same corpus can be deduplicated, and only one of the multiple identical segmented words in the same corpus is retained, and the total number of occurrences of the segmented word in the corpus is calculated. In a specific implementation, SentencePiece can be used to segment and deduplicate each low-resource corpus and each high-resource corpus to obtain multiple segmented words corresponding to each corpus. Then, initial word vectors are assigned to the segmented words corresponding to each corpus respectively to form a vocabulary corresponding to each corpus respectively.
[0163] Figure 7Schematically shown is the step flow of updating the word list corresponding to the low-resource corpus with the word list corresponding to the high-resource corpus in the embodiments of the present application. As Figure 7 shown, based on the above embodiments, in some embodiments, the step of updating the word list corresponding to the low-resource corpus with the word list corresponding to the high-resource corpus in step S230 may further include the following steps S710 to S720:
[0164] S710. Obtain the same word segments in the word list corresponding to the low-resource corpus and the word list corresponding to the high-resource corpus as the overlapping word segments.
[0165] S720. Update the overlapping word segments and their corresponding word vectors in the word list corresponding to the low-resource corpus according to the overlapping word segments and their corresponding word vectors in the word list corresponding to the high-resource corpus.
[0166] Specifically, it may be that the overlapping word segments and their corresponding word vectors in the word list corresponding to the high-resource corpus are used as the overlapping word segments and their corresponding word vectors in the word list corresponding to the low-resource corpus, so that the word vectors of the overlapping word segments in the word list corresponding to the high-resource corpus are the same as the word vectors of the overlapping word segments in the word list corresponding to the low-resource corpus.
[0167] Alternatively, it may be that the overlapping word segments and their corresponding word vectors in the word list corresponding to the low-resource corpus are updated according to the overlapping word segments and their corresponding word vectors in the word list corresponding to the high-resource corpus, so that the word vectors of the overlapping word segments in the word list corresponding to the high-resource corpus are similar to the word vectors of the overlapping word segments in the word list corresponding to the low-resource corpus.
[0168] Updating the overlapping word segments and their corresponding word vectors in the word list corresponding to the low-resource corpus according to the overlapping word segments and their corresponding word vectors in the word list corresponding to the high-resource corpus can make the word vectors of the overlapping word segments in the word list corresponding to the high-resource corpus similar to or the same as the word vectors of the overlapping word segments in the word list corresponding to the low-resource corpus. It can be understood that in different languages, there may be the same word roots, and the meanings of the same word roots may be close or the same. Thus, it can make the training of the machine translation model more in line with the actual logic of the language and improve the translation effect of the machine translation model.
[0169] Moreover, updating the overlapping word segments and their corresponding word vectors in the word list corresponding to the low-resource corpus according to the overlapping word segments and their corresponding word vectors in the word list corresponding to the high-resource corpus can share the word vector features obtained from the high-resource corpus into the low-resource corpus, avoid the problem of data sparsity in the training process of the low-resource corpus, and is beneficial to improving the training effect of the language to be trained corresponding to the low-resource corpus.
[0170] Figure 8Schematically shows the step - by - step process of calculating the update completion degree of the word list corresponding to the low - resource corpus in the embodiments of the present application. As Figure 8 shown, based on the above - mentioned embodiments, in some embodiments, calculating the update completion degree of the word list corresponding to the low - resource corpus in step S230 may further include the following steps S810 to S830:
[0171] S810. Obtain the number of segmented words of the same high - frequency segmented words in the word list corresponding to the low - resource corpus and the word list corresponding to the high - resource corpus. The high - frequency segmented words are the segmented words whose occurrence frequencies in the corpus are within the preset ranking.
[0172] S820. Determine the similarity between each high - resource corpus and the low - resource corpus respectively according to the ratio of the number of segmented words to the preset ranking.
[0173] S830. Take the training completion degree corresponding to the high - resource corpus with the highest similarity to the low - resource corpus as the update completion degree of the word list corresponding to the low - resource corpus.
[0174] Thus, the update completion degree of the word list corresponding to the low - resource corpus can be reflected by the training completion degree corresponding to the high - resource corpus with the highest similarity to the low - resource corpus. It can be understood that the word list corresponding to the high - resource corpus with the highest similarity to the low - resource corpus has a relatively large number of the same high - frequency segmented words as the word list corresponding to the low - resource corpus. Therefore, when the training completion degree of the word list corresponding to the high - resource corpus is relatively high, it indicates that most of the word vectors in the word list corresponding to the high - resource corpus have been updated, and by updating the word list corresponding to the low - resource corpus with the word list corresponding to the high - resource corpus, a relatively large number of word vectors in the word list corresponding to the low - resource corpus have been updated.
[0175] Therefore, taking the training completion degree corresponding to the high - resource corpus with the highest similarity to the low - resource corpus can better reflect the update completion degree of the word list corresponding to the low - resource corpus, so as to accurately evaluate the update completion degree of the word list corresponding to the low - resource corpus, and further enable the execution timing of adding the low - resource corpus with the update completion degree of the word list greater than the preset threshold to the training data set to be more accurate. Thus, it can further avoid the problem of data sparsity in the training process of the low - resource corpus, which is beneficial to improving the training effect of the to - be - trained language corresponding to the low - resource corpus.
[0176] Specifically, the similarity between the low - resource corpus i and the high - resource corpus j can be calculated by the following formula:
[0177]
[0178] where k is the set number, vocab k (i) is the segmented words ranked in the top k positions in the word list corresponding to the low - resource corpus i, vocabk (j) is a word segment that ranks within the top k in terms of the frequency of occurrence in the vocabulary corresponding to the high-resource corpus |vocab k (i) ∩ vocab k (j)| is the number of word segments that are the same high-frequency word segments in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus. High-frequency word segments are word segments that rank within the top k in terms of the frequency of occurrence in the corpus.
[0179] Figure 9 Schematically shows the step flow of calculating the update completion degree of the vocabulary corresponding to the low-resource corpus in the embodiments of the present application. As Figure 9 shown, on the basis of the above embodiments, in some embodiments, calculating the update completion degree of the vocabulary corresponding to the low-resource corpus in step S230 may further include the following steps S910 to step S930:
[0180] S910. Obtain the same high-frequency word segments in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus. High-frequency word segments are word segments that rank within a preset ranking in terms of the frequency of occurrence in the corpus.
[0181] S920. Determine the similarity between each high-resource corpus and the low-resource corpus respectively according to the ratio of the number of word segments to the preset ranking.
[0182] S930. Use the similarity between each high-resource corpus and the low-resource corpus as a weight, and perform a weighted summation operation on the training completion degree corresponding to each high-resource corpus to obtain the update completion degree of the vocabulary corresponding to the low-resource corpus.
[0183] Thus, it is possible to reflect the update completion degree of the vocabulary corresponding to the low-resource corpus through the result of the weighted summation operation of the training completion degree corresponding to each high-resource corpus using the similarity between each high-resource corpus and the low-resource corpus as a weight. It can be understood that using the similarity between each high-resource corpus and the low-resource corpus as a weight can make the weight corresponding to the high-resource corpus with a higher similarity to the low-resource corpus larger, and the weight corresponding to the high-resource corpus with a lower similarity to the low-resource corpus smaller. Therefore, using the similarity between each high-resource corpus and the low-resource corpus as a weight and performing a weighted summation operation on the training completion degree corresponding to each high-resource corpus can better reflect the update completion degree of the vocabulary corresponding to the low-resource corpus, thereby being able to accurately evaluate the update completion degree of the vocabulary corresponding to the low-resource corpus, and further enabling the execution timing of adding the low-resource corpus with an update completion degree greater than the preset threshold to the training data set to be more accurate. Thus, it is possible to further avoid the problem of data sparsity during the training process of the low-resource corpus, which is beneficial to improving the training effect of the language to be trained corresponding to the low-resource corpus.
[0184] Specifically, taking the similarity between each high-resource corpus and the low-resource corpus as a weight, a weighted sum operation is performed on the training completion degrees corresponding to each high-resource corpus to obtain the initialization completion degree corresponding to the low-resource corpus, which can be calculated by the following calculation formula:
[0185]
[0186] Among them, c avg is the initialization completion degree corresponding to the low-resource corpus i, and c j is the training completion degree corresponding to the high-resource corpus j. The similarity between the low-resource corpus i and the high-resource corpus j can be calculated by the calculation formula (5) in the above text, or can be calculated by other methods, which will not be elaborated here.
[0187] Figure 10 Schematically shows the step flow of calculating the update completion degree of the vocabulary corresponding to the low-resource corpus in the embodiment of the present application. As Figure 10 shown, on the basis of the above embodiments, in some embodiments, calculating the update completion degree of the vocabulary corresponding to the low-resource corpus in step S230 may further include the following steps S1010 to step S1040:
[0188] S1010. Obtain the same high-frequency word segments in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus. The high-frequency word segments are the word segments whose occurrence frequencies in the corpus are within the preset ranking.
[0189] S1020. Determine the similarity between each high-resource corpus and the low-resource corpus respectively according to the ratio of the number of word segments to the preset ranking.
[0190] S1030. Calculate the activation weight S i corresponding to each high-resource corpus through the following calculation formula:
[0191]
[0192] Among them, v i and v j are both the similarities between the high-resource corpus and the low-resource corpus, and e is the natural logarithm.
[0193] S1040. According to the activation weight S i , perform a weighted sum operation on the training completion degrees corresponding to each high-resource corpus to obtain the update completion degree of the vocabulary corresponding to the low-resource corpus.
[0194] Thus, the result obtained by performing a weighted sum operation on the training completion degrees corresponding to each high-resource corpus according to the activation weight S i can reflect the update completion degree of the vocabulary corresponding to the low-resource corpus. It can be understood that the activation weight S iis proportional to the similarity between the high-resource corpus and the low-resource corpus, and the corresponding activation weight S of each high-resource corpus i As the weight, it can make the corresponding activation weight S of the high-resource corpus with a higher similarity to the low-resource corpus i larger, and the corresponding activation weight S of the high-resource corpus with a higher similarity to the low-resource corpus i smaller. Therefore, according to the activation weight S i , perform a weighted summation operation on the training completion degrees corresponding to each high-resource corpus to obtain the update completion degree of the word list corresponding to the low-resource corpus, which can better reflect the update completion degree of the word list corresponding to the low-resource corpus, so as to accurately evaluate the update completion degree of the word list corresponding to the low-resource corpus, and further enable the execution timing of adding the low-resource corpus with the update completion degree of the word list greater than the preset threshold to the training data set to be more accurate. Thus, it can further avoid the problem of data sparsity in the training process of the low-resource corpus, which is beneficial to improving the training effect of the to-be-trained language corresponding to the low-resource corpus.
[0195] Specifically, the similarity between the high-resource corpus and the low-resource corpus can be calculated by the calculation formula (5) in the above text, or can be calculated by other methods, which will not be elaborated here.
[0196] S240. Add the low-resource corpus with the update completion degree of the word list greater than the preset threshold to the training data set.
[0197] Add the low-resource corpus with the update completion degree of the word list greater than the preset threshold to the training data set to continue training the machine translation model according to the corpus in the training data set, and update the word list corresponding to the high-resource corpus and the word list corresponding to the low-resource corpus in the training data set during the training process. Thus, the low-resource corpus in the candidate data set can be added to the training data set when the update completion degree of the word list is greater than the preset threshold, start training the model, and update the word list corresponding to the low-resource corpus during the training process.
[0198] Specifically, please continue to refer to Figure 3 , when the update completion degree of the word list corresponding to the low-resource corpus is greater than the preset threshold, add the low-resource corpus to the training data set to train the model, and update the word list corresponding to the low-resource corpus during the training process. The above has already described Figure 3 the states (b), (c), and (d) in which the low-resource corpus with the update completion degree of the word list greater than the preset threshold is added to the training data set, which will not be elaborated here.
[0199] In some embodiments, after the update completion degree of the vocabulary corresponding to the low-resource corpus is greater than a preset threshold, the update of the vocabulary corresponding to the low-resource corpus using the vocabulary corresponding to the high-resource corpus can be stopped, and the update completion degree of the vocabulary corresponding to the low-resource corpus can be calculated. However, during the model training process, the low-resource corpus with an update completion degree of the vocabulary greater than the preset threshold and the high-resource corpus are jointly used to train the model. During this process, the word vectors of the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus with an update completion degree of the vocabulary greater than the preset threshold can be shared bidirectionally, so that the word vectors of the same overlapping word segment in the vocabulary corresponding to the high-resource corpus and the vocabulary corresponding to the low-resource corpus are the same, which can make the word vectors more accurately represent the features of the word segments, thereby improving the translation effect of the machine translation model.
[0200] In some embodiments, after adding the low-resource corpus with an update completion degree of the vocabulary greater than the preset threshold to the training data set in step S240, the present application may further include:
[0201] When the state where only the first preset number of low-resource corpora remain to be added to the training data set lasts for a preset time or a second preset number of steps, the first preset number of low-resource corpora are added to the training data set.
[0202] Thus, it is possible to avoid the situation that at the end of the model training, some low-resource corpora are still not added to the training data set due to insufficient update completion degree of the vocabulary or other limiting reasons, thereby avoiding the situation that some low-resource corpora are still not input into the model for training at the end of the model training. Thus, the training duration of each corpus can be balanced, and the training effect on the machine translation model can be improved. Specifically, the first preset number can be 1, 2, 3, etc.
[0203] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution, etc.
[0204] The following introduces the apparatus embodiments of the present application, which can be used to execute the machine translation model training method in the above embodiments of the present application. Figure 11 Schematically shows the structural block diagram of the machine translation model training apparatus provided by the embodiments of the present application. As Figure 11 shown, the machine translation model training apparatus 1100 may include:
[0205] The corpus acquisition module 1110 is configured to acquire a training data set and a candidate data set; the training data set includes one or more high-resource corpora and the word lists corresponding to each high-resource corpus, and the candidate data set includes one or more low-resource corpora and the word lists corresponding to each low-resource corpus. The data volume of the high-resource corpus is greater than that of the low-resource corpus, and the word list is used to represent the corresponding relationship between word segmentation and word vectors;
[0206] The model training module 1120 is configured to train the machine translation model using the corpora in the training data set and update the word lists corresponding to each corpus in the training data set during the training process;
[0207] The word list update module 1130 is configured to update the word list corresponding to the low-resource corpus using the word list corresponding to the high-resource corpus and calculate the update completion degree of the word list corresponding to the low-resource corpus. The update completion degree is used to represent the progress of updating the word list corresponding to the low-resource corpus using the word list corresponding to the high-resource corpus;
[0208] The training data set addition module 1140 is configured to add the low-resource corpora with the update completion degree of the word list greater than a preset threshold to the training data set.
[0209] In some embodiments of the present application, based on the above embodiments, the model training module includes:
[0210] The sampling unit is configured to sample each corpus in the training data set according to the sampling weights corresponding to each corpus in the training data set;
[0211] The model training unit is configured to train the machine translation model according to the sampled corpus;
[0212] The sampling weight update unit is configured to calculate the training completion degree corresponding to each corpus in the training data set during the training process and update the sampling weights of the corpora according to the training completion degree. The sampling weight is inversely proportional to the training completion degree.
[0213] In some embodiments of the present application, based on the above embodiments, the sampling weight update unit includes:
[0214] The first loss function acquisition subunit is configured to acquire a preset first loss function value for representing the convergence of the machine translation model training;
[0215] The second loss function acquisition subunit is configured to input the validation sets corresponding to each corpus in the training data set into the machine translation model to obtain the second loss function values corresponding to each corpus in the training data set;
[0216] A training completion calculation subunit, configured to calculate the difference between the first loss function value and the second loss function value, and calculate the training completion corresponding to each piece of corpus in the training dataset according to the difference.
[0217] In some embodiments of the present application, based on the above embodiments, the vocabulary update module includes:
[0218] A first same word segmentation quantity acquisition unit, configured to acquire the word segmentation quantity of the same high-frequency word segments in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus, where the high-frequency word segments are word segments whose occurrence frequency in the corpus is within a preset ranking;
[0219] A first similarity determination unit, configured to determine the similarity between each piece of high-resource corpus and the low-resource corpus respectively according to the ratio of the word segmentation quantity to the preset ranking;
[0220] A first update completion calculation unit, configured to use the training completion corresponding to the high-resource corpus with the highest similarity to the low-resource corpus as the update completion of the vocabulary corresponding to the low-resource corpus.
[0221] In some embodiments of the present application, based on the above embodiments, the vocabulary update module further includes:
[0222] A second same word segmentation quantity acquisition unit, configured to acquire the same high-frequency word segments in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus, where the high-frequency word segments are word segments whose occurrence frequency in the corpus is within a preset ranking;
[0223] A second similarity determination unit, configured to determine the similarity between each piece of high-resource corpus and the low-resource corpus respectively according to the ratio of the word segmentation quantity to the preset ranking;
[0224] A second update completion calculation unit, configured to use the similarity between each piece of high-resource corpus and the low-resource corpus as a weight, perform a weighted summation operation on the training completion corresponding to each piece of high-resource corpus, and obtain the update completion of the vocabulary corresponding to the low-resource corpus.
[0225] In some embodiments of the present application, based on the above embodiments, the vocabulary update module further includes:
[0226] A third same word segmentation quantity acquisition unit, configured to acquire the same high-frequency word segments in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus, where the high-frequency word segments are word segments whose occurrence frequency in the corpus is within a preset ranking;
[0227] A third similarity determination unit, configured to determine the similarity between each piece of high-resource corpus and the low-resource corpus respectively according to the ratio of the word segmentation quantity to the preset ranking;
[0228] An activation weight calculation unit, configured to calculate the activation weight S corresponding to each piece of high-resource corpus through the following calculation formula i :
[0229]
[0230] wherein, v i and v j are both the similarities between the high-resource corpus and the low-resource corpus, and e is the natural logarithm;
[0231] A third update completion degree calculation unit, configured to perform a weighted summation operation on the training completion degrees corresponding to each piece of high-resource corpus according to the activation weight S i to obtain the update completion degree of the vocabulary corresponding to the low-resource corpus.
[0232] In some embodiments of the present application, based on the above embodiments, the training completion degree calculation subunit includes:
[0233] A specific training completion degree calculation subunit, configured to calculate the training completion degree c corresponding to each piece of corpus in the training dataset through the following calculation formula:
[0234]
[0235] wherein, L * is the first loss function value, and L is the second loss function value.
[0236] In some embodiments of the present application, based on the above embodiments, the sampling weight update unit further includes:
[0237] A sampling weight update subunit, configured to calculate the training completion degrees corresponding to each piece of corpus in the training dataset at every preset number of steps during the training process, where the step size is the preset data volume for inputting the corpus into the machine translation model.
[0238] In some embodiments of the present application, based on the above embodiments, the machine translation model training device further includes:
[0239] A low-resource corpus vocabulary acquisition unit, configured to perform word segmentation and deduplication on each piece of low-resource corpus, obtain multiple word segments corresponding to the low-resource corpus, and assign initial word vectors to the multiple word segments corresponding to the low-resource corpus to form a vocabulary corresponding to each piece of low-resource corpus;
[0240] A high-resource corpus vocabulary acquisition unit, configured to perform word segmentation and deduplication on each piece of high-resource corpus, obtain multiple word segments corresponding to the high-resource corpus, and assign initial word vectors to the multiple word segments corresponding to the high-resource corpus to form a vocabulary corresponding to each piece of high-resource corpus.
[0241] In some embodiments of the present application, based on the above embodiments, the vocabulary update module further includes:
[0242] A coincident word segmentation acquisition unit, configured to obtain the same word segmentations in the vocabulary corresponding to the low-resource corpus and the vocabulary corresponding to the high-resource corpus as the coincident word segmentations;
[0243] A word vector update unit, configured to update the coincident word segmentations and the corresponding word vectors in the vocabulary corresponding to the low-resource corpus according to the coincident word segmentations and the corresponding word vectors in the vocabulary corresponding to the high-resource corpus.
[0244] In some embodiments of the present application, based on the above embodiments, the model training module further includes:
[0245] A first word vector update unit, configured to update the word vectors corresponding to the word segmentations in the vocabulary corresponding to each corpus in the training dataset during the training process to obtain updated word vectors;
[0246] The vocabulary update module further includes:
[0247] A second word vector update unit, configured to update the vocabulary corresponding to the low-resource corpus with the updated word vectors in the vocabulary corresponding to the high-resource corpus.
[0248] In some embodiments of the present application, based on the above embodiments, the machine translation model is used to translate the central language into one or more languages, or to translate one or more languages into the central language;
[0249] The high-resource corpus includes a bilingual parallel corpus corresponding to the preset language and the central language, and the low-resource corpus includes a bilingual parallel corpus corresponding to the preset language and the central language in the preset domain.
[0250] The specific details of the machine translation model training device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments, and will not be repeated here.
[0251] Figure 12 Schematically shows a structural block diagram of an electronic device for implementing the embodiments of the present application.
[0252] It should be noted that Figure 12 The illustrated electronic device 1200 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0253] Such as Figure 12As shown, the electronic device 1200 includes a central processing unit 1201 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1202 (ROM) or the program loaded from the storage section 1208 into the random access memory 1203 (RAM). In the random access memory 1203, various programs and data required for the operation of the device are also stored. The central processing unit 1201, the read-only memory 1202, and the random access memory 1203 are connected to each other via a bus 1204. An input / output interface 1205 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1204.
[0254] The following components are connected to the input / output interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a local area network card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage section 1208 as needed.
[0255] Specifically, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit 1201, various functions defined in the device of the present application are executed.
[0256] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor device, apparatus, or component, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution device, apparatus, or component. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0257] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based device for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0258] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0259] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0260] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0261] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for training a machine translation model, characterized in that, the method includes: Obtain a training data set and a candidate data set; the training data set includes one or more high-resource corpora and the word lists corresponding to each of the high-resource corpora, the candidate data set includes one or more low-resource corpora and the word lists corresponding to each of the low-resource corpora, the data volume of the high-resource corpus is greater than that of the low-resource corpus, and the word list is used to represent the correspondence between word segmentation and word vectors; Use the corpora in the training data set to train the machine translation model, and update the word lists corresponding to each corpus in the training data set during the training process; Update the word list corresponding to the low-resource corpus with the word list corresponding to the high-resource corpus, and calculate the update completion degree of the word list corresponding to the low-resource corpus, where the update completion degree is used to represent the progress of updating the word list corresponding to the low-resource corpus with the word list corresponding to the high-resource corpus; Add the low-resource corpora with the update completion degree of the word list greater than the preset threshold to the training data set.
2. The method according to claim 1, characterized in that, the step of using the corpora in the training data set to train the machine translation model includes: Sampling each corpus in the training data set according to the sampling weights corresponding to each corpus in the training data set; Train the machine translation model according to the sampled corpora; During the training process, calculate the training completion degrees corresponding to each corpus in the training data set, and update the sampling weights of the corpora according to the training completion degrees, where the sampling weights are inversely proportional to the training completion degrees.
3. The method according to claim 2, characterized in that, the step of calculating the training completion degrees corresponding to each corpus in the training data set during the training process includes: Obtain a preset first loss function value for indicating the convergence of the machine translation model training; Input the validation sets corresponding to each corpus in the training data set into the machine translation model to obtain the second loss function values corresponding to each corpus in the training data set; Calculate the difference between the first loss function value and the second loss function value, and calculate the training completion degrees corresponding to each corpus in the training data set according to the difference.
4. The method according to claim 3, characterized in that, the step of calculating the update completion degree of the word list corresponding to the low-resource corpus includes: Obtain the number of segmented words of the same high-frequency segmented words in the word list corresponding to the low-resource corpus and the word list corresponding to the high-resource corpus, where the high-frequency segmented words are the segmented words with the occurrence frequency within a preset ranking in the corpus; Determine the similarity between each high-resource corpus and the low-resource corpus according to the ratio of the number of segmented words to the preset ranking; Use the training completion degree corresponding to the high-resource corpus with the highest similarity to the low-resource corpus as the update completion degree of the word list corresponding to the low-resource corpus.
5. The method according to claim 3, characterized in that, the step of calculating the update completion degree of the word list corresponding to the low-resource corpus includes: Obtain the same high-frequency word segments in the word list corresponding to the low-resource corpus and the word list corresponding to the high-resource corpus, where the high-frequency word segments are word segments whose occurrence frequencies in the corpus are within a preset ranking; Determine the similarity between each high-resource corpus and the low-resource corpus respectively according to the ratio of the number of word segments to the preset ranking; Use the similarity between each high-resource corpus and the low-resource corpus as a weight, and perform a weighted summation operation on the training completion degrees corresponding to each high-resource corpus to obtain the update completion degree of the word list corresponding to the low-resource corpus.
6. The method according to claim 3, wherein, the calculating the update completion degree of the word list corresponding to the low-resource corpus includes: Obtain the same high-frequency word segments in the word list corresponding to the low-resource corpus and the word list corresponding to the high-resource corpus, where the high-frequency word segments are word segments whose occurrence frequencies in the corpus are within a preset ranking; Determine the similarity between each high-resource corpus and the low-resource corpus respectively according to the ratio of the number of word segments to the preset ranking; Calculate the activation weight S corresponding to each portion of the high-resource corpus through the following calculation formula i : where v i and v j are both the similarities between the high-resource corpus and the low-resource corpus, and e is the natural logarithm; According to the activation weight S i , perform a weighted summation operation on the training completion degrees corresponding to each piece of the high-resource corpus to obtain the update completion degree of the vocabulary corresponding to the low-resource corpus.
7. The method according to claim 3, wherein, the calculating the training completion degree corresponding to each corpus in the training dataset according to the difference includes: Calculate the training completion degree c corresponding to each corpus in the training dataset through the following calculation formula: wherein, L * is the first loss function value, and L is the second loss function value.
8. The method according to claim 2, wherein, during the training process, calculating the training completion degree corresponding to each corpus in the training dataset includes: During the training process, every preset number of steps, calculate the training completion degree corresponding to each corpus in the training dataset, where the step is the preset data volume for inputting the corpus into the machine translation model.
9. The method according to claim 1, wherein, before updating the word list corresponding to the low-resource corpus with the word list corresponding to the high-resource corpus, the method further includes: Perform word segmentation and deduplication on each low-resource corpus to obtain multiple word segments corresponding to the low-resource corpus, and assign initial word vectors to the multiple word segments corresponding to the low-resource corpus to form a word list corresponding to each low-resource corpus respectively; Perform word segmentation and deduplication on each high-resource corpus to obtain multiple word segments corresponding to the high-resource corpus, and assign initial word vectors to the multiple word segments corresponding to the high-resource corpus to form a word list corresponding to each high-resource corpus respectively.
10. The method according to claim 1, wherein, the updating the word list corresponding to the low-resource corpus with the word list corresponding to the high-resource corpus includes: Obtain the same word segments in the word list corresponding to the low-resource corpus and the word list corresponding to the high-resource corpus as the overlapping word segments; Update the overlapping word segments and the corresponding word vectors in the word list corresponding to the low-resource corpus according to the overlapping word segments and the corresponding word vectors in the word list corresponding to the high-resource corpus.
11. The method according to claim 1, wherein, the updating the word list corresponding to each corpus in the training dataset during the training process includes: During the training process, update the word vectors corresponding to the word segmentations in the vocabulary corresponding to each corpus in the training dataset to obtain updated word vectors; The updating the vocabulary corresponding to the low-resource corpus by using the vocabulary corresponding to the high-resource corpus includes: Updating the vocabulary corresponding to the low-resource corpus by using the updated word vectors in the vocabulary corresponding to the high-resource corpus.
12. The method according to any one of claims 1-11, wherein, The machine translation model is used to translate the central language into one or more languages, or to translate one or more of the languages into the central language; The high-resource corpus includes bilingual parallel corpus corresponding to the preset language and the central language, and the low-resource corpus includes bilingual parallel corpus corresponding to the preset language and the central language in a preset domain.
13. A machine translation model training device, wherein, The device includes: A corpus acquisition module configured to acquire a training dataset and a candidate dataset; the training dataset includes one or more high-resource corpora and the vocabulary corresponding to each high-resource corpus, the candidate dataset includes one or more low-resource corpora and the vocabulary corresponding to each low-resource corpus, the data volume of the high-resource corpus is greater than that of the low-resource corpus, and the vocabulary is used to represent the correspondence between word segmentations and word vectors; A model training module configured to train the machine translation model by using the corpora in the training dataset and update the vocabulary corresponding to each corpus in the training dataset during the training process; A vocabulary update module configured to update the vocabulary corresponding to the low-resource corpus by using the vocabulary corresponding to the high-resource corpus and calculate the update completion degree of the vocabulary corresponding to the low-resource corpus, where the update completion degree is used to represent the progress of updating the vocabulary corresponding to the low-resource corpus by using the vocabulary corresponding to the high-resource corpus; A training dataset addition module configured to add the low-resource corpus with the update completion degree of the vocabulary greater than a preset threshold to the training dataset.
14. A computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the machine translation model training method according to any one of claims 1 to 12 is implemented.
15. An electronic device, wherein, includes: A processor; and A memory for storing executable instructions of the processor; wherein, the processor is configured to execute the machine translation model training method according to any one of claims 1 to 12 by executing the executable instructions.
Citation Information
Patent Citations
Training method and device for neural network machine translation model
CN109117483A
German lexical analysis method and system for neural network machine translation
CN110765766A