A model training method, device, computing device and readable storage medium

By training the translation model and word vector transformation components, the third translation model is constructed, which solves the problem of overfitting neural machine translation in the translation task between languages ​​with fewer bilingual corpus, improves translation quality and reduces training costs.

CN113688637BActive Publication Date: 2025-05-06ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010424684.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-19
Publication Date
2025-05-06
Estimated Expiration
2040-05-19

AI Technical Summary

Technical Problem

Neural machine translation is prone to overfitting in translation tasks between languages ​​with fewer bilingual corpus, resulting in poor translation results and requires retraining each time you face a brand new language, which is costly.

Method used

By training the first and second translation models, and training the source language and target language transformation components using word vectors of the source language and the target language, a third translation model is built for translating text from the second source language into the second target language.

Benefits of technology

It improves the translation quality of languages ​​with fewer bilingual corpus, solves the problem of data sparsity and overfitting, and does not require repeated training, which is less expensive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113688637B_ABST
    Figure CN113688637B_ABST
Patent Text Reader

Abstract

The present invention discloses a model training method, including: training a first translation model, the first translation model is used to translate a text of a first source language into a first target language; training a second translation model, the second translation model is used to translate a text of a second source language into a second target language; training a source language conversion component using a word vector of the first source language and a word vector of the second source language; training a target language conversion component using a word vector of the first target language and a word vector of the second target language; and constructing a third translation model based on the trained first translation model, the second translation model, the source language conversion component and the target language conversion component, the third translation model is used to translate the text of the second source language into a second target language. The present invention also discloses a corresponding model training device, a translation device, a computing device and a readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model training method, device, computing equipment and readable storage medium. Background Art

[0002] Machine translation is the process of using computers to convert one natural language (source language) into another natural language (target language). In recent years, machine translation based on neural networks, referred to as neural machine translation, has developed rapidly and has been widely used in the translation field.

[0003] Since the number of translation model parameters of neural machine translation is huge, a large amount of bilingual parallel corpus, such as millions of sentence pairs, is required to train a relatively ideal translation model. Therefore, neural machine translation can achieve good translation results in translation tasks between languages ​​with a large amount of bilingual corpus, such as English, German, and French. However, for translation tasks between languages ​​with less or even scarce bilingual corpus, the translation model of neural machine translation is prone to serious overfitting, resulting in poor translation results.

[0004] At present, it is expected to use the rich bilingual corpus of other languages ​​in a unified model to improve the translation quality of languages ​​with less bilingual corpus. Therefore, it is usually directly built to translate between languages ​​with rich bilingual corpus and between languages ​​with less bilingual corpus. However, the disadvantage of this method is that it needs to be retrained every time a new language is encountered, which is costly.

[0005] Therefore, there is an urgent need for a more advanced model training scheme to improve the translation quality of languages ​​with less bilingual corpus. Summary of the invention

[0006] To this end, embodiments of the present invention provide a model training method, apparatus, computing device, and readable storage medium in an effort to solve or at least alleviate at least one of the above problems.

[0007] According to one aspect of an embodiment of the present invention, a model training method is provided, including: training a first translation model, the first translation model is used to translate text in a first source language into a first target language; training a second translation model, the second translation model is used to translate text in a second source language into a second target language; training a source language conversion component using word vectors of the first source language and word vectors of the second source language; training a target language conversion component using word vectors of the first target language and word vectors of the second target language; and constructing a third translation model based on the trained first translation model, the second translation model, the source language conversion component and the target language conversion component, the third translation model being used to translate text in the second source language into the second target language.

[0008] Optionally, in the method according to an embodiment of the present invention, the source language conversion component is trained using the word vectors of the first source language and the word vectors of the second source language, including: calculating the nonlinear conversion of the word vectors of the second source language to the word vectors of the first source language to obtain nonlinear conversion relationship one.

[0009] Optionally, in the method according to the embodiment of the present invention, calculating the nonlinear conversion of the word vector of the second source language to the word vector of the first source language includes: using at least one layer of feedforward neural network to learn the nonlinear conversion relationship, and the feedforward neural network adopts a nonlinear activation function.

[0010] Optionally, in the method according to an embodiment of the present invention, the target language conversion component is trained using the word vectors of the first target language and the word vectors of the second target language, including: calculating the nonlinear conversion of the word vectors of the second target language to the word vectors of the first target language to obtain nonlinear conversion relationship 2.

[0011] Optionally, in the method according to the embodiment of the present invention, calculating the nonlinear conversion of the word vector of the second target language to the word vector of the first target language includes: using at least one layer of feedforward neural network to learn the nonlinear conversion relationship 2, and the feedforward neural network adopts a nonlinear activation function.

[0012] Optionally, in the method according to an embodiment of the present invention, a third translation model is constructed based on the trained first translation model, second translation model, source language conversion component and target language conversion component, including: applying non-linear conversion relationship 1 to the word vector of the second source language of the second translation model to obtain the word vector of the first source language of the third translation model; applying non-linear conversion relationship 2 to the word vector of the second target language of the second translation model to obtain the word vector of the first target language of the third translation model; using the encoder of the first translation model as the encoder of the third translation model; and using the decoder of the first translation model as the decoder of the third translation model.

[0013] Optionally, in the method according to an embodiment of the present invention, the decoder of the first translation model and / or the encoder of the first translation model are constructed based on a long short-term memory neural network model or a Transformer model.

[0014] Optionally, in the method according to the embodiment of the present invention, training the first translation model includes: training the first translation model using a first bilingual corpus, the first bilingual corpus including texts written in a first source language and a first target language that have a translation relationship with each other.

[0015] Optionally, in the method according to the embodiment of the present invention, training the second translation model includes: training the second translation model in a multi-task learning manner.

[0016] Optionally, in the method according to the embodiment of the present invention, the second translation model includes a second source language word embedding component and a second target language word embedding component, the second source language word embedding component is used to generate word vectors for words in the second source language, and the second target language word embedding component is used to generate word vectors for words in the second target language, and the second translation model is trained in a multi-task learning manner, including: pre-training the second source language word embedding component using monolingual corpus in the second source language; pre-training the second target language word embedding component using monolingual corpus in the second target language; training the second translation model using second bilingual corpus, the second bilingual corpus including texts written in the second source language and the second target language that have a translation relationship with each other.

[0017] Optionally, in the method according to the embodiment of the present invention, the first translation model includes a first source language word embedding component, the first source language word embedding component is used to generate word vectors for words in the first source language, and the source language conversion component is trained using the word vectors of the first source language and the word vectors of the second source language, including: generating the word vectors of the first source language via the trained first source language word embedding component; generating the word vectors of the second source language via the trained second source language word embedding component.

[0018] Optionally, in the method according to the embodiment of the present invention, the first translation model includes a first target language word embedding component, the first target language word embedding component is used to generate word vectors for words in the first target language, and the target language conversion component is trained using the word vectors of the first target language and the word vectors of the second target language, including: generating the word vectors of the first target language via the trained first target language word embedding component; generating the word vectors of the second target language via the trained second target language word embedding component.

[0019] Optionally, in the method according to the embodiment of the present invention, after constructing the third translation model, the method further includes: training the third translation model using the second bilingual corpus.

[0020] Optionally, in the method according to the embodiment of the present invention, the amount of the second bilingual corpus is smaller than that of the first bilingual corpus.

[0021] According to another aspect of an embodiment of the present invention, a translation method is provided, which is suitable for translating a text in a second source language into a second target language using a third translation model, wherein the third translation model is trained using a model training method according to an embodiment of the present invention, and comprises: generating word vectors for words in the second source language; applying a non-linear conversion relationship 1 to the word vectors of the second source language to obtain word vectors of the first source language; receiving the word vectors of the first source language via an encoder of the third translation model and encoding them; generating word vectors for words in the second target language previously output by a decoder of the third translation model; applying a non-linear conversion relationship 2 to the word vectors of the second target language to obtain word vectors of the first target language; and decoding and outputting the current word in the second target language via the decoder of the third translation model based on the state of the encoder of the third translation model and the word vectors of the first target language.

[0022] According to another aspect of an embodiment of the present invention, a model training device is provided, including: a first model training module, suitable for training a first translation model, the first translation model is used to translate a text in a first source language into a first target language; a second model training module, suitable for training a second translation model, the second translation model is used to translate a text in a second source language into a second target language; a word vector conversion training module, suitable for training a source language conversion component using word vectors of the first source language and word vectors of the second source language; and suitable for training a target language conversion component using word vectors of the first target language and word vectors of the second target language; and a third model training module, suitable for constructing a third translation model based on the trained first translation model, the second translation model, the source language conversion component and the target language conversion component, the third translation model is used to translate the text in the second source language into the second target language.

[0023] According to another aspect of an embodiment of the present invention, a translation device is provided, which is suitable for translating a text in a second source language into a second target language by using a third translation model, wherein the third translation model is trained by using a model training method according to an embodiment of the present invention, and comprises a third source language word embedding component, a third encoder, a third decoder, a third target language word embedding component, a source language conversion component and a target language conversion component, wherein the third source language word embedding component is suitable for generating word vectors for words in the second source language; the source language conversion component is suitable for receiving word vectors in the second source language and performing non-linear conversion on the word vectors in the second source language; the third target language word embedding component is suitable for generating word vectors for words in the second target language; the target language conversion component is suitable for performing non-linear conversion on the word vectors in the second target language; the third encoder is suitable for receiving and encoding the word vectors converted by the source language conversion component; and the third decoder is suitable for decoding and outputting the current word in the second target language based on the state of the third encoder and the word vectors converted by the target language conversion component of the word previously output by the third decoder.

[0024] According to another aspect of an embodiment of the present invention, a computing device is provided, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the model training method and the translation method according to an embodiment of the present invention.

[0025] According to another aspect of an embodiment of the present invention, a readable storage medium storing a program is provided, wherein the program includes instructions, which, when executed by a computing device, causes the computing device to execute any one of the model training method and the translation method according to an embodiment of the present invention.

[0026] The model training method according to the embodiment of the present invention constructs a translation model of the desired language (e.g., the third translation model 300) by utilizing a translation model of other languages ​​with rich bilingual corpora (e.g., the first translation model 200a) and learning a word vector mapping relationship from other languages ​​to the desired language (e.g., the source language conversion component 350 and the target language conversion component 360), thereby achieving a significant improvement in the translation performance and translation quality of the translation model of the desired language, and solving the problems of data sparsity and overfitting. In addition, the entire process does not require repeated training, and the cost is relatively low.

[0027] Furthermore, by adopting a multi-task learning approach to train the word vectors of the desired language (e.g., the second source language word embedding component 210b and the second target language word embedding component 240b in the second translation model 200b), the model can learn the translation information while learning the word vectors of the desired language, which can improve the learning effect and make the word vectors finally learned more suitable for the translation task of the desired language. At the same time, it also reduces network overfitting and improves the generalization effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To achieve the above and related purposes, certain illustrative aspects are described herein in conjunction with the following description and accompanying drawings, which indicate various ways in which the principles disclosed herein can be practiced, and all aspects and their equivalents are intended to fall within the scope of the claimed subject matter. The above and other purposes, features and advantages of the present disclosure will become more apparent by reading the following detailed description in conjunction with the accompanying drawings. Throughout the present disclosure, the same reference numerals generally refer to the same parts or elements.

[0029] Figure 1 A schematic diagram of a translation system 100 according to an embodiment of the present invention is shown;

[0030] Figure 2 A schematic diagram of a translation model 200 according to an embodiment of the present invention is shown;

[0031] Figure 3 A schematic diagram showing a first translation model 200a, a second translation model 200b and a third translation model 300 according to an embodiment of the present invention is shown;

[0032] Figure 4 A schematic diagram of a computing device 400 according to one embodiment of the present invention is shown;

[0033] Figure 5 A flowchart of a model training method 500 according to an embodiment of the present invention is shown; and

[0034] Figure 6 A structural block diagram of a model training device 600 according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0035] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0036] Figure 1 FIG. 1 is a schematic diagram of a translation system 100 according to an embodiment of the present invention. Figure 1 As shown, the translation system 100 includes a translation front end 110, a translation device 120 and a model training device 600. The translation front end 110 is any requesting party that needs to translate the text of the source language into the target language. For example, in one way, the translation front end 110 can be a part of an instant messaging system. The instant messaging system can receive a message text input by a user. If the user indicates that the message text needs to be translated, the recognition front end 110 can send the message text to the translation device 120 for translation. In another way, the translation front end 110 can be a part of a text display system. The text display system can display network text. If the user indicates that the network text needs to be translated, the recognition front end 110 can send the network text to the translation device 120 for translation.

[0037] The present invention is not limited to the specific form of the translation front end 110. The translation device 120 may receive the request of the translation front end 110 in various ways. For example, the translation device 120 may provide an application program interface (API) with a predetermined format definition to facilitate the translation front end 110 to organize the translation request according to the definition and send it to the translation device 120.

[0038] The translation device 120 uses a translation model based on a neural network to perform translation, that is, a text in a source language can be input into the translation model so that the translation model outputs a translation text of the text in the source language to the target language. The text in the source language can include multiple words in the source language, and the translation text in the target language can include multiple words in the target language.

[0039] Figure 2 FIG. 2 shows a schematic diagram of a translation model 200 according to an embodiment of the present invention. Figure 2 As shown, the translation model 200 may include a source language word embedding component 210, an encoder 220, a decoder 230, and a target language word embedding component 240. In some cases, the translation model 200 may also include other different or more components.

[0040] The source language word embedding component 210 can generate a word vector for a word in the first source language, and then output it to the encoder 220 for encoding. The target language word embedding component 240 can generate a word vector for a word in the target language output by the decoder 230, and then output it to the decoder 230. The decoder 230 can decode and output the current word in the target language based on the state of the encoder 220 and the word vector of the word in the target language previously output by the decoder 230.

[0041] It should be noted that the embodiments of the present invention do not limit the specific word embedding algorithm. For example, the Skip-Gram model and CBOW (continuous bags of word) model in the word2vec tool, the Glove model, and various neural network language models, etc. can be used.

[0042] It should also be noted that the encoder 220 may have multiple (multi-layer) layers, and the corresponding decoder 230 may also have multiple (multi-layer) layers. In addition, various neural network models may be used to construct the decoder 220 or the encoder 230, such as various recurrent neural network models (RNN), long short-term memory network models (LSTM), Transformer models, Bert models, and convolutional neural network models, etc.

[0043] In some embodiments, the translation model 200 may also use an attention mechanism, for example, may include an attention component. The attention mechanism is well known to those skilled in the art and will not be described in detail herein.

[0044] The translation model used by the translation device 120 can be trained by the model training device 600. Generally speaking, the model training device 600 can use bilingual parallel texts of the source language and the target language to train the translation model. The bilingual parallel texts include texts written in the source language and the target language that have a translation relationship with each other. In other words, it can include the text in the source language and the translation text of the text in the source language to the target language.

[0045] The following is an example of a bilingual corpus in English and Chinese:

[0046] “I need water.

[0047] I need water.”

[0048] Model training device 600 can use the text of source language as the input of model, and use the translation text from source language to target language as the output of model to train the model, and adjust the parameters of model. It should be noted that in some cases, the bilingual corpus written in source language and target language is quite abundant, such as bilingual corpus of major languages ​​such as English-French, English-German, Chinese-English. Therefore, the model translation effect obtained by training the translation model with a large amount of high-quality bilingual corpus is good. In other cases, the bilingual corpus written in source language and target language is very small (or of low quality), such as bilingual corpus between small languages ​​such as Mongolian and Tibetan to any other. If this small sample set is used to train the translation model, serious overfitting and sparsity problems will occur, and the translation performance is poor.

[0049] Therefore, in order to solve the above problems, the model training device 300 according to an embodiment of the present invention can use the translation model 200a of other languages ​​(usually languages ​​with abundant bilingual corpora) and the translation model 200b of the desired language to construct the final translation model 300 of the desired language.

[0050] For ease of description, the translation model 200a of other languages ​​is referred to as the first translation model 200a hereinafter, which is used for translating the text of the first source language into the first target language. The translation model 200b of the desired language is referred to as the second translation model 200b, which is used for translating the text of the second source language into the second target language. The translation model 300 finally obtained is referred to as the 3rd translation model 300, which is also used for translating the text of the second source language into the second target language. In some cases, the second source language is the same language as the first source language, and the second target language is different languages ​​from the first target language. In other cases, the second source language is different languages ​​from the first source language, and the second target language is the same language as the first target language. In some other cases, the second source language is different languages ​​from the first source language, and the second target language is also different languages ​​from the first target language.

[0051] Figure 3 A schematic diagram of a first translation model 200a, a second translation model 200b and a third translation model 300 according to an embodiment of the present invention is shown.

[0052] like Figure 3 As shown, similar to the translation model 200, the first translation model 200a may include a first source language word embedding component 210a, a first encoder 220a, a first decoder 230a, and a first target language word embedding component 240a. The first source language word embedding component 210a may generate a word vector for a word in the first source language, and then output it to the first encoder 220a for encoding. The first target language word embedding component 240a may generate a word vector for a word in the first target language output by the first decoder 230a, and then output it to the first decoder 230a. The first decoder 230a may decode and output the current word in the first target language based on the state of the first encoder 220a and the word vector of the word in the first target language previously output by the first decoder 230a.

[0053] Similarly to the translation model 200, the second translation model 200b may include a second source language word embedding component 210b, a second encoder 220b, a second decoder 230b, and a second target language word embedding component 240b. The second source language word embedding component 210b can generate a word vector for a word in the second source language, and then output it to the second encoder 220b for encoding. The second target language word embedding component 240b can generate a word vector for a word in the second target language output by the second decoder 230b, and then output it to the second decoder 230b. The second decoder 230b can decode and output the current word in the second target language based on the state of the second encoder 220a and the word vector of the word in the second target language previously output by the second decoder 230b.

[0054] The third translation model 300 may include a third source language word embedding component 310, a third encoder 320, a third decoder 330, a third target language word embedding component 340, a source language conversion component 350, and a target language conversion component 360. The third source language word embedding component 310 may generate a word vector for a word in a second source language, and then output it to the source language conversion component 350. The source language conversion component 350 may perform a linear and / or nonlinear conversion on the word vector of the second source language to map the word vector of the second source language to the word vector of the first source language. The third encoder 320 receives the word vector converted by the source language conversion component 350 and encodes it. The third target language word embedding component 340 may generate a word vector for the word in the second target language output by the third decoder 330, and then output it to the target language conversion component 360. The target language conversion component 360 may perform a linear and / or nonlinear conversion on the word vector of the second target language to map the word vector of the second target language to the word vector of the first target language.

[0055] The third decoder 330 may decode and output the current word in the second target language based on the state of the third encoder 320 and the word vector of the word previously output by the third decoder 330 converted by the target language conversion component 360 .

[0056] The specific structures of at least some components (such as word embedding components, decoders, and encoders, etc.) in the first translation model 200a, the second translation model 200b, and the third translation model 300 can be referred to in conjunction with Figure 2 The translation model 200 described is not described in detail here.

[0057] The following will describe in detail the process of the model training device 600 constructing the final third translation model 300 based on the first translation model 200a and the second translation model 200b.

[0058] According to an embodiment of the present invention, the model training apparatus 600 may be implemented by a computing device 400 as described below.

[0059] Figure 4 FIG. 4 is a schematic diagram of a computing device 400 according to an embodiment of the present invention. Figure 4 As shown, in a basic configuration 402, computing device 400 typically includes a system memory 406 and one or more processors 404. A memory bus 408 may be used for communication between processor 404 and system memory 406.

[0060] Depending on the desired configuration, the processor 404 can be any type of process, including but not limited to: a microprocessor (μP), a microcontroller (μC), a digital information processor (DSP), or any combination thereof. The processor 404 can include one or more levels of cache such as a primary cache 410 and a secondary cache 412, a processor core 414, and registers 416. An example processor core 414 can include an arithmetic logic unit (ALU), a floating point unit (FPU), a digital signal processing core (DSP core), or any combination thereof. An example memory controller 418 can be used with the processor 404, or in some implementations, the memory controller 418 can be an internal part of the processor 404.

[0061] Depending on the desired configuration, system memory 406 may be any type of memory, including but not limited to: volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.), or any combination thereof. System memory 406 may include an operating system 420, one or more applications 422, and program data 424. In some implementations, application 422 may be arranged to execute instructions on the operating system by one or more processors 404 using program data 424.

[0062] The computing device 400 may also include an interface bus 440 that facilitates communication from various interface devices (e.g., output devices 442, peripheral interfaces 444, and communication devices 446) to the basic configuration 402 via the bus / interface controller 430. Example output devices 442 include a graphics processing unit 448 and an audio processing unit 450. They can be configured to facilitate communication with various external devices such as a display or speakers via one or more A / V ports 452. Example peripheral interfaces 444 may include a serial interface controller 454 and a parallel interface controller 456, which may be configured to facilitate communication with external devices such as input devices (e.g., keyboards, mice, pens, voice input devices, touch input devices) or other peripherals (e.g., printers, scanners, etc.) via one or more I / O ports 458. Example communication devices 446 may include a network controller 460, which may be arranged to facilitate communication with one or more other computing devices 462 via a network communication link via one or more communication ports 464.

[0063] A network communication link can be an example of a communication medium. Communication media can generally be embodied as computer-readable instructions, data structures, program modules in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. A "modulated data signal" can be a signal in which one or more of its data sets or its changes can be carried out in a manner that encodes information in the signal. As non-limiting examples, communication media can include wired media such as a wired network or a dedicated line network, and various wireless media such as sound, radio frequency (RF), microwave, infrared (IR) or other wireless media. The term computer-readable medium used herein can include both storage media and communication media.

[0064] The computing device 400 can be implemented as a server, such as a database server, an application server, a WEB server, etc., or as a personal computer including a desktop computer and a notebook computer. Of course, the computing device 400 can also be implemented as a part of a small-sized portable (or mobile) electronic device.

[0065] In an embodiment according to the present invention, the computing device 400 is implemented as a model training device 600 according to an embodiment of the present invention, and is configured to execute the model training method 500 according to an embodiment of the present invention. Among them, the application 422 of the computing device 400 includes multiple program instructions for executing the model training method 500 according to an embodiment of the present invention, and the program data 424 can also store configuration information of the model training device 600, etc.

[0066] Figure 5 FIG. 5 shows a flow chart of a model training method 500 according to an embodiment of the present invention. Figure 5 As shown, the model training method 500 starts with step S510.

[0067] In step S510, a first translation model 200a for translating a text in a first source language into a first target language is trained. As described above, the first translation model 200a may include a first source language word embedding component 210a, a first encoder 220a, a first decoder 230a, and a first target language word embedding component 240a. The first source language word embedding component 210a is used to generate a word vector for a word in the first source language. The first target language word embedding component 240a is used to generate a word vector for a word in the first target language.

[0068] It can be understood that the bilingual corpus including texts written in the first source language and the first target language and having a translation relationship between them is referred to as the first bilingual corpus. The first bilingual corpus is usually abundant, so the first translation model 200a can be constructed first, and the first translation model 200a can be trained using a large amount of the first bilingual corpus. During the training process, the text in the first source language in the first bilingual corpus is used as the input of the first translation model 200a, and the translation text of the text in the first source language input in the first bilingual corpus to the first target language is used as the output of the first translation model 200a.

[0069] Any suitable loss function and training algorithm in the art may be used to train the first translation model 200a. For example, a function such as cross entropy or negative log likelihood may be used as a loss function, and an optimization algorithm such as Adam (Adaptive Moment Estimation), AdaGrad (Adaptive Gradient Algorithm) and SGD (Stochastic Gradient Descent) may be used for training. The embodiment of the present invention does not limit the specific loss function and training algorithm.

[0070] In step S520, a second translation model 200b for translating text in a second source language into a second target language may also be trained. As described above, the second translation model 200b may include a second source language word embedding component 210b, a second encoder 220b, a second decoder 230b, and a second target language word embedding component 240b. The second source language word embedding component 210b is used to generate a word vector for a word in the second source language. The second target language word embedding component 240b is used to generate a word vector for a word in the second target language.

[0071] Understandably, the bilingual corpus including texts written in the second source language and the second target language and having a translation relationship with each other is called the second bilingual corpus. The amount of the second bilingual corpus is usually relatively small, less than or even much less than the amount of the first bilingual corpus.

[0072] In an embodiment of the present invention, the second translation model 200b can be trained in a multi-task learning manner. Specifically, there are the following interrelated tasks: a word vector learning task of a second source language, a word vector learning task of a second target language, and a translation task from a second source language to a second target language. Therefore, the second source language word embedding component 210b can be pre-trained using a monolingual corpus of the second source language, and the second target language word embedding component 240b can be pre-trained using a monolingual corpus of the second target language. Among them, the monolingual corpus of the second source language is a corpus that only includes text written in the second source language, and the monolingual corpus of the second target language is a corpus that only includes text written in the second target language.

[0073] Then, the second bilingual corpus is used to train the entire second translation model 200b, thereby further adjusting the parameters of the entire model 200b, especially the word embedding component thereof. During the training process, the text of the second source language in the second bilingual corpus is used as the input of the second translation model 200b, and the translation text of the second source language input in the second bilingual corpus to the second target language is used as the output of the second translation model 200b.

[0074] Any suitable loss function and training algorithm in the art may be used to train the second translation model 200b. For example, a function such as cross entropy or negative log likelihood may be used as a loss function, and an optimization algorithm such as Adam (Adaptive Moment Estimation), AdaGrad (Adaptive Gradient Algorithm) and SGD (Stochastic Gradient Descent) may be used for training. The embodiment of the present invention does not limit the specific loss function and training algorithm.

[0075] It can be understood that by training the second translation model 200b in a multi-task learning manner, the model can learn translation information while learning the word vectors of the second source language and the second target language, which can improve the learning effect and make the final learned word vector more suitable for the translation task. At the same time, it also reduces network overfitting and improves the generalization effect.

[0076] After the first translation model 200a and the second translation model 200b are trained, the source language conversion component 350 can be trained using the word vectors of the first source language and the word vectors of the second source language in step S530. The source language conversion component 350 is used to perform linear and / or nonlinear conversion on the word vectors of the second source language so as to map the word vectors of the second source language to the word vectors of the first source language. That is, the nonlinear conversion of the word vectors of the second source language to the word vectors of the first source language can be calculated to obtain a nonlinear conversion relationship one.

[0077] Specifically, the source language conversion component 350 can be constructed based on various neural network models (i.e., various neural network models can be used to learn nonlinear conversion relationships), such as single-layer feedforward neural network models, multi-layer feedforward neural network models, and other complex neural network models. In some embodiments, the neural network model used by the source language conversion component 350 can use functions such as ReLU, ELU, and PReLU as activation functions. The embodiment of the present invention does not limit the specific network model and activation function of the source language conversion component 350.

[0078] Next, the word vector of the first source language can be generated through the trained first source language word embedding component 210a, and the word vector of the second source language can be generated through the trained second source language word embedding component 210b. The word vectors of the first source language and the word vectors of the second source language generated are used to train the source language conversion component 350. During the training process, the word vector of the second source language is used as the input of the source language conversion component 350, and the word vector of the first source language is used as the output of the source language conversion component 350.

[0079] It should be noted that any suitable loss function and training algorithm in the art can be used to train the source language conversion component 350. For example, a function such as cross entropy or mean square error can be used as a loss function, and an optimization algorithm such as Adam (Adaptive Moment Estimation), AdaGrad (Adaptive Gradient Algorithm) and SGD (Stochastic Gradient Descent) can be used for training. The embodiment of the present invention does not limit the specific loss function and training algorithm.

[0080] After the first translation model 200a and the second translation model 200b are trained, the target language conversion component 360 can also be trained using the word vectors of the first target language and the word vectors of the second target language in step S540. The target language conversion component 360 is used to perform linear and / or nonlinear conversion on the word vectors of the second target language so as to map the word vectors of the second target language to the word vectors of the first target language. That is, the nonlinear conversion of the word vectors of the second target language to the word vectors of the first target language is calculated to obtain the second nonlinear conversion relationship.

[0081] Specifically, the target language conversion component 360 can be constructed based on various neural network models (i.e., various neural network models can be used to learn nonlinear conversion relationship 2), such as a single-layer feedforward neural network model, a multi-layer feedforward neural network model, and other complex neural network models. In some embodiments, the neural network model used by the target language conversion component 360 can use a nonlinear function or a linear function such as a Sigmoid function, a tanh function, a Relu function, etc. as an activation function. The embodiment of the present invention does not limit the specific network model and activation function of the target language conversion component 360.

[0082] Next, the word vector of the first target language can be generated through the trained first target language word embedding component 240a, and the word vector of the second target language can be generated through the trained second target language word embedding component 240b. The generated word vectors of the first target language and the word vectors of the second target language are used to train the target language conversion component 360. During the training process, the word vector of the second target language is used as the input of the target language conversion component 360, and the word vector of the first target language is used as the output of the target language conversion component 360.

[0083] It should be noted that any suitable loss function and training algorithm in the art can be used to train the target language conversion component 360. For example, a function such as cross entropy or mean square error can be used as a loss function, and an optimization algorithm such as Adam (Adaptive Moment Estimation), AdaGrad (Adaptive Gradient Algorithm) and SGD (Stochastic Gradient Descent) can be used for training. The embodiment of the present invention does not limit the specific loss function and training algorithm.

[0084] After the first translation model 200a, the second translation model 200b, the source language conversion component 350, and the target language conversion component 360 are trained, in step S550, the third translation model 300 can be constructed based on the trained first translation model 200a, the second translation model 200b, the source language conversion component 350, and the target language conversion component 360. Specifically, the third translation model 300 can be constructed based on the second target language word embedding component 240b and the second target language word embedding component 240b in the second translation model 200b, the first encoder 220a and the first decoder 230a in the first translation model 200a, and the source language conversion component 350 and the target language conversion component 360.

[0085] refer to Figure 3In addition to directly adopting the source language conversion component 350 and the target language conversion component 360, the third translation model 300 can also use the first encoder 220a in the first translation model 200a as the third encoder 320 in the third translation model 300, use the first decoder 230a in the first translation model 200a as the third decoder 330 in the third translation model 300, use the second source language word embedding component 210b in the second translation model 200b as the third source language word embedding component 310 in the third translation model 300, use the second target language word embedding component 240b in the second translation model 200b as the third target language word embedding component 340 in the third translation model 300, use the second source language word embedding component 210b as the third source language word embedding component 310 in the third translation model 300, and use the second target language word embedding component 240b in the second translation model 200b as the third target language word embedding component 340 in the third translation model 300.

[0086] That is, the third translation model 300 may include the second source language word embedding component 210b and the second target language word embedding component 240b in the second translation model 200b, the first encoder 220a and the first decoder 230a in the first translation model 200a, and the source language conversion component 350 and the target language conversion component 360.

[0087] The third translation model 300 uses the second source language word embedding component 210b to receive the words of the second source language, obtains the word vector of the second source language, and then applies the nonlinear conversion relationship 1 corresponding to the source language conversion component 350 to the obtained word vector of the second source language to obtain the word vector of the first source language. The third translation model 300 uses the second target language word embedding component 240b to receive the words of the second target language output by the first decoder 230a, obtains the word vector of the second target language, and then applies the nonlinear conversion relationship 2 corresponding to the target language conversion component 360 to the obtained word vector of the second target language to obtain the word vector of the first target language. The third translation model 300 also uses the first encoder 220a to receive the word vector of the first source language and encode it, and uses the first decoder 230a to decode and output the current word of the second target language based on the state of the first encoder 220a and the word vector of the first target language.

[0088] The third translation model 300 obtained in this way has greatly improved the translation performance from the second source language to the second target language, and the translation quality is good.

[0089] In addition, in some embodiments, for better translation effect, the second bilingual corpus can be used to train the third translation model 300, so as to further fine-tune the parameters of the entire model 300. During the training process, the text in the second source language in the second bilingual corpus is used as the input of the second translation model 200b, and the translation text of the second source language input in the second bilingual corpus to the second target language is used as the output of the second translation model 200b.

[0090] Among them, any suitable loss function and training algorithm in the art can be used to train the third translation model 300. For example, a function such as cross entropy or negative log likelihood can be used as a loss function, and an optimization algorithm such as Adam (Adaptive Moment Estimation), AdaGrad (Adaptive Gradient Algorithm) and SGD (Stochastic Gradient Descent) can be used for training. The embodiment of the present invention does not limit the specific loss function and training algorithm.

[0091] Figure 6 FIG. 6 shows a structural block diagram of a model training device 600 according to an embodiment of the present invention. Figure 6 As shown, the model training device 600 may include a first model training module 610, a second model training module 620, a word vector conversion training module 630 and a third model training module 640.

[0092] The first model training module 610 can train the first translation model 200a. The first translation model 200a is used to translate the text of the first source language into the first target language. The second model training module 620 can train the second translation model 200b. The second translation model 200b is used to translate the text of the second source language into the second target language. The word vector conversion training module 630 can train the source language conversion component 350 using the word vector of the first source language and the word vector of the second source language, and can also train the target language conversion component 360 using the word vector of the first target language and the word vector of the second target language. The third model training module 640 can construct the third translation model 300 based on the trained first translation model 200a, the second translation model 200b, the source language conversion component 350 and the target language conversion component 360. The third translation model 300 is used to translate the text of the second source language into the second target language.

[0093] The above combined Figure 1 to Figure 5The specific description of the translation model 300 and the model training method 500 used by the translation device 120 has already explained the corresponding processing in each module of the model training device 600 in detail, and the repeated content will not be repeated here.

[0094] In summary, the model training method according to the embodiment of the present invention constructs a translation model of the desired language (e.g., the third translation model 300) by utilizing a translation model of other languages ​​with rich bilingual corpora (e.g., the first translation model 200a) and learning a word vector mapping relationship from other languages ​​to the desired language (e.g., the source language conversion component 350 and the target language conversion component 360), thereby achieving a great improvement in the translation performance and translation quality of the translation model of the desired language, and solving the problems of data sparsity and overfitting. Moreover, the whole process does not require repeated training, and the cost is relatively low.

[0095] Furthermore, by adopting a multi-task learning approach to train the word vectors of the desired language (e.g., the second source language word embedding component 210b and the second target language word embedding component 240b in the second translation model 200b), the model can learn the translation information while learning the word vectors of the desired language, which can improve the learning effect and make the word vectors finally learned more suitable for the translation task of the desired language. At the same time, it also reduces network overfitting and improves the generalization effect.

[0096] The embodiment of the present invention also provides a translation method, which uses the above-mentioned third translation model 300 to translate the text of the second source language into the second target language. Specifically, a word vector can be generated for the word of the second source language input into the third translation model 300, and the non-linear conversion relationship 1 is applied to the word vector of the second source language to obtain the word vector of the first source language. Then, the word vector of the first source language obtained by conversion is received by the encoder of the third translation model 300 and encoded. It is also possible to generate a word vector for the word of the second target language previously output by the decoder of the third translation model 300, and the non-linear conversion relationship 2 is applied to the word vector of the second target language to obtain the word vector of the first target language. Then, the decoder of the third translation model 300 decodes and outputs the current word of the second target language based on the state of the encoder of the third translation model 300 and the word vector of the first target language obtained by conversion.

[0097] It is understandable that the translation method according to an embodiment of the present invention can be applied to many application scenarios, including but not limited to web browsing scenarios, video watching scenarios, shopping scenarios, etc. For example, in a web browsing scenario, a web page in an unfamiliar language (especially a minority language) can be translated into a familiar language using a translation method according to an embodiment of the present invention. In a video watching scenario, a subtitle in an unfamiliar language (especially a minority language) can be translated into a familiar language using a translation method according to an embodiment of the present invention. In a shopping scenario, a product description in an unfamiliar language (especially a minority language) can be translated into a familiar language using a translation method according to an embodiment of the present invention. The present invention does not impose any restrictions on these application scenarios.

[0098] It should be understood that the various techniques described herein can be implemented in combination with hardware or software, or a combination thereof. Thus, the method and apparatus of the present invention, or some aspects or portions of the method and apparatus of the present invention, can take the form of program codes (i.e., instructions) embedded in a tangible medium, such as a floppy disk, a CD-ROM, a hard disk drive, or any other machine-readable storage medium, wherein when the program is loaded into a machine such as a computer and executed by the machine, the machine becomes a device for practicing the present invention.

[0099] In the case where the program code is executed on a programmable computer, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The memory is configured to store the program code; the processor is configured to execute various methods of the present invention according to the instructions in the program code stored in the memory.

[0100] By way of example and not limitation, computer readable media include computer storage media and communication media. Computer readable media include computer storage media and communication media. Computer storage media stores information such as computer readable instructions, data structures, program modules or other data. Communication media generally embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery medium. Any combination of the above is also included within the scope of computer readable media.

[0101] It should be understood that in order to streamline the present disclosure and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the following intention: that the claimed invention requires more features than those explicitly recited in each claim. More specifically, as reflected in the claims below, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Therefore, the claims that follow the specific embodiment are hereby expressly incorporated into the specific embodiment, with each claim itself serving as a separate embodiment of the present invention.

[0102] Those skilled in the art will appreciate that the modules or units or components of the devices in the examples disclosed herein may be arranged in the devices described in the embodiment, or alternatively may be located in one or more devices different from the devices in the examples. The modules in the foregoing examples may be combined into one module or may be divided into multiple submodules.

[0103] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0104] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.

[0105] In addition, some of the embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or by other devices that perform the functions. Therefore, a processor with necessary instructions for implementing the method or method elements forms a device for implementing the method or method elements. In addition, the elements described herein of the device embodiments are examples of devices for implementing the functions performed by the elements for the purpose of implementing the invention.

[0106] As used herein, unless otherwise specified, the use of ordinal numbers "first," "second," "third," etc. to describe common objects merely indicates that different instances of similar objects are involved, and is not intended to imply that the objects so described must have a given order in time, space, order, or in any other manner.

[0107] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments may be envisioned within the scope of the invention thus described. In addition, it should be noted that the language used in this specification is selected primarily for readability and teaching purposes, rather than for explaining or defining the subject matter of the present invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the present invention is illustrative, not restrictive, with respect to the scope of the present invention, which is defined by the appended claims.

Claims

1. A model training method, comprising: Training a first translation model, the first translation model being used to translate text in a first source language into a first target language; training a second translation model, the second translation model being used to translate text in a second source language into a second target language; Training a source language conversion component using the word vectors of the first source language and the word vectors of the second source language; Training a target language conversion component using the word vectors of the first target language and the word vectors of the second target language; as well as constructing a third translation model based on the trained first translation model, the second translation model, the source language conversion component, and the target language conversion component, wherein the third translation model is used to translate the text in the second source language into the second target language; Among them, the source language conversion component in the third translation model is used to map the word vector of the second source language to the word vector of the first source language, the third encoder in the third translation model is used to encode the converted word vector of the first source language, and the third encoder is constructed based on the first translation model. The target language conversion component in the third translation model is used to map the word vector of the second target language to the word vector of the first target language, and the converted word vector is input into the third decoder in the third translation model, and the third decoder is used to decode and output the words of the second target language, and the third decoder is constructed based on the second translation model.

2. The method of claim 1, wherein: The using the word vectors of the first source language and the word vectors of the second source language to train the source language conversion component includes: A nonlinear conversion of the word vector of the second source language to the word vector of the first source language is calculated to obtain a nonlinear conversion relationship one.

3. The method of claim 2, wherein: The calculating of the non-linear conversion of the word vector of the second source language to the word vector of the first source language includes: The nonlinear conversion relationship 1 is learned using at least one layer of feedforward neural network, and the feedforward neural network adopts a nonlinear activation function.

4. The method according to any one of claims 1 to 3, wherein: The step of training the target language conversion component using the word vector of the first target language and the word vector of the second target language includes: A nonlinear conversion of the word vector of the second target language to the word vector of the first target language is calculated to obtain a second nonlinear conversion relationship.

5. The method of claim 4, wherein: The calculating of the non-linear conversion from the word vector of the second target language to the word vector of the first target language includes: The second nonlinear conversion relationship is learned using at least one layer of feedforward neural network, wherein the feedforward neural network adopts a nonlinear activation function.

6. The method according to claim 5, wherein: The constructing a third translation model based on the trained first translation model, the second translation model, the source language conversion component and the target language conversion component includes: Applying the nonlinear conversion relationship 1 to the word vector of the second source language of the second translation model to obtain the word vector of the first source language of the third translation model; Applying the second nonlinear transformation relationship to the word vector of the second target language of the second translation model to obtain the word vector of the first target language of the third translation model; Using the encoder of the first translation model as the encoder of the third translation model; The decoder of the first translation model is used as the decoder of the third translation model.

7. The method of claim 6, wherein: The decoder of the first translation model and / or the encoder of the first translation model are constructed based on a long short-term memory neural network model or a Transformer model.

8. The method of claim 1, wherein: Training the first translation model includes: The first translation model is trained using a first bilingual corpus, where the first bilingual corpus includes texts written in the first source language and the first target language that have a translation relationship with each other.

9. The method of claim 1, wherein: Training the second translation model includes: The second translation model is trained in a multi-task learning manner.

10. The method of claim 9, wherein: The second translation model includes a second source language word embedding component and a second target language word embedding component, the second source language word embedding component is used to generate a word vector for a word in the second source language, and the second target language word embedding component is used to generate a word vector for a word in the second target language. Training the second translation model in a multi-task learning manner includes: Pre-training the second source language word embedding component using monolingual corpus in the second source language; Pre-training the second target language word embedding component using monolingual corpus in the second target language; The second translation model is trained using a second bilingual corpus, where the second bilingual corpus includes texts written in the second source language and the second target language that have a translation relationship with each other.

11. The method of claim 2, wherein: The first translation model includes a first source language word embedding component, the first source language word embedding component is used to generate word vectors for words in the first source language, and the training of the source language conversion component using the word vectors of the first source language and the word vectors of the second source language includes: Generate a word vector of the first source language via the trained first source language word embedding component; Generate a word vector of the second source language via the trained word embedding component of the second source language.

12. The method of claim 5, wherein: The first translation model includes a first target language word embedding component, the first target language word embedding component is used to generate word vectors for words in the first target language, and the training of the target language conversion component using the word vectors of the first target language and the word vectors of the second target language includes: Generate a word vector of the first target language via the trained word embedding component of the first target language; Generate a word vector of the second target language via the trained word embedding component of the second target language.

13. The method of claim 1, wherein: After constructing the third translation model, the method further includes: The third translation model is trained using the second bilingual corpus.

14. The method of claim 13, wherein: The second bilingual corpus is smaller in size than the first bilingual corpus.

15. A translation method suitable for using the third translation The third translation model translates the text in the second source language into the second target language, and the third translation model is trained by the model training method according to any one of claims 1 to 14, comprising: Generate word vectors for words in the second source language; Apply the nonlinear transformation relationship 1 to the word vector of the second source language to obtain the word vector of the first source language; Receiving the word vector of the first source language through the encoder of the third translation model and encoding it; Generate a word vector for the word in the second target language previously output by the decoder of the third translation model; Apply the nonlinear transformation relationship 2 to the word vector of the second target language to obtain the word vector of the first target language; as well as The decoder of the third translation model decodes and outputs the current word in the second target language based on the state of the encoder of the third translation model and the word vector of the first target language.

16. A model training device, comprising: A first model training module, adapted to train a first translation model, wherein the first translation model is used to translate text in a first source language into a first target language; A second model training module, adapted to train a second translation model, wherein the second translation model is used to translate text in a second source language into a second target language; a word vector conversion training module, adapted to train a source language conversion component using the word vectors of the first source language and the word vectors of the second source language; and adapted to train a target language conversion component using the word vectors of the first target language and the word vectors of the second target language; as well as a third model training module, adapted to construct a third translation model based on the trained first translation model, the second translation model, the source language conversion component and the target language conversion component, wherein the third translation model is used to translate the text in the second source language into the second target language; Among them, the source language conversion component in the third translation model is used to map the word vector of the second source language to the word vector of the first source language, the third encoder in the third translation model is used to encode the converted word vector of the first source language, and the third encoder is constructed based on the first translation model. The target language conversion component in the third translation model is used to map the word vector of the second target language to the word vector of the first target language, and the converted word vector is input into the third decoder in the third translation model, and the third decoder is used to decode and output the words of the second target language, and the third decoder is constructed based on the second translation model.

17. A translation device, adapted to translate a text in a second source language into a second target language using a third translation model, wherein the third translation model is trained using the model training method according to any one of claims 1 to 14, and comprises a third source language word embedding component, a third encoder, a third decoder, a third target language word embedding component, a source language conversion component, and a target language conversion component, wherein The third source language word embedding component is adapted to generate word vectors for words in the second source language; The source language conversion component is adapted to receive the word vector of the second source language and perform nonlinear conversion on the word vector of the second source language; The third target language word embedding component is adapted to generate word vectors for words in the second target language; The target language conversion component is adapted to perform nonlinear conversion on the word vectors of the second target language; The third encoder is adapted to receive the word vector converted by the source language conversion component and perform encoding; as well as The third decoder is adapted to decode and output a current word in the second target language based on a state of the third encoder and a word vector of a word previously output by the third decoder after conversion by the target language conversion component.

18. A computing device comprising: one or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the model training methods described in claims 1-14.

19. A readable storage medium storing a program, wherein the program includes instructions, and when the instructions are executed by a computing device, the computing device executes any one of the model training methods described in claims 1-14.

Citation Information

Patent Citations

  • A method and device for training a translation model

    CN109271644A

  • Machine translation method and device, electronic equipment, storage medium and translation model

    CN110427630A