Translation model fine-tuning method, translation method, device and electronic equipment
By updating and fine-tuning the vocabulary of the translation model and utilizing parallel corpora from both general and target domains, the problem of low translation accuracy and efficiency caused by insufficient parallel corpora is solved, achieving efficient translation in a specific domain.
Patent Information
- Application Number
- CN202410726174.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-06-05
AI Technical Summary
Because there are fewer parallel corpora for some languages, especially in specific domains, multilingual translation models or translation models do not pay much attention to some languages, particularly in terms of translation accuracy and efficiency between two languages in specific domains.
By acquiring the translation model to be fine-tuned and a set of parallel corpora in the general domain, and combining them with a set of monolingual corpora in the target domain, the vocabulary of the translation model is updated. The updated translation model is then fine-tuned using the parallel corpora in the general and target domains, thereby improving the vectorization accuracy of words in the target domain in the vocabulary.
It improves the accuracy and efficiency of translation models in specific domains, especially the accuracy of translating professional terms in the target domain.
Smart Images

Figure CN118536519B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of deep learning, natural language processing, and large models, and especially to a method for fine-tuning a translation model, a translation method, a device, and an electronic device. Background Technology
[0002] Currently, in translation between two languages, multilingual translation models or general-domain translation models between two languages are generally used to process the translation between the two languages.
[0003] In particular, due to the scarcity of parallel corpora for some languages, especially those in specific domains, multilingual translation models or translation models do not pay enough attention to certain languages, especially those in specific domains. This results in poor translation accuracy and efficiency between two languages in a specific domain. Summary of the Invention
[0004] This disclosure provides a method for fine-tuning a translation model, a translation method, an apparatus, and an electronic device.
[0005] According to one aspect of this disclosure, a method for fine-tuning a translation model is provided, the method comprising: obtaining a translation model to be fine-tuned, and a general domain parallel corpus set of the translation model; the general domain parallel corpus set includes parallel corpora between at least two target languages; obtaining a monolingual corpus set of the target language within the target domain; updating the vocabulary of the target language in the translation model according to the monolingual corpus set of the target language within the target domain to obtain a vocabulary-updated translation model; and fine-tuning the vocabulary-updated translation model according to the general domain parallel corpus set.
[0006] According to another aspect of this disclosure, a translation method is provided, the method comprising: obtaining a fine-tuned translation model; the fine-tuned translation model being determined based on a fine-tuning method for a translation model as described above; obtaining corpus of a first target language to be translated; and combining the corpus of the first target language and the fine-tuned translation model to determine corpus of a second target language corresponding to the corpus of the first target language.
[0007] According to another aspect of this disclosure, a fine-tuning apparatus for a translation model is provided, the apparatus comprising: a first acquisition module for acquiring a translation model to be fine-tuned, and a general domain parallel corpus set of the translation model; the general domain parallel corpus set includes parallel corpora between at least two target languages; a second acquisition module for acquiring a monolingual corpus set of the target language within the target domain; an update processing module for updating the vocabulary of the target language in the translation model according to the monolingual corpus set of the target language within the target domain, to obtain a vocabulary-updated translation model; and a fine-tuning processing module for fine-tuning the vocabulary-updated translation model according to the general domain parallel corpus set.
[0008] According to another aspect of this disclosure, a translation apparatus is provided, the apparatus comprising: a first acquisition module for acquiring a fine-tuned translation model; the fine-tuned translation model being determined based on a fine-tuning method for the translation model as described above; a second acquisition module for acquiring corpus of a first target language to be translated; and a determination module for combining the corpus of the first target language and the fine-tuned translation model to determine corpus of a second target language corresponding to the corpus of the first target language.
[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a fine-tuning method of the translation model proposed above in this disclosure; or to perform a translation method proposed above in this disclosure.
[0010] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to execute the fine-tuning method of the translation model proposed in this disclosure; or to execute the translation method proposed in this disclosure.
[0011] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the fine-tuning method for the translation model proposed above in this disclosure; or, implements the steps of the translation method proposed above in this disclosure.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0015] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0016] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;
[0017] Figure 4 This is a diagram illustrating the fine-tuning of the translation model;
[0018] Figure 5 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0019] Figure 6 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0020] Figure 7 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0021] Figure 8 This is a block diagram of an electronic device used to implement the fine-tuning method or translation method of the translation model in the embodiments of this disclosure. Detailed Implementation
[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0023] Currently, in translation between two languages, multilingual translation models or general-domain translation models between two languages are generally used to process the translation between the two languages.
[0024] In particular, due to the scarcity of parallel corpora for some languages, especially those in specific domains, multilingual translation models or translation models do not pay enough attention to certain languages, especially those in specific domains. This results in poor translation accuracy and efficiency between two languages in a specific domain.
[0025] To address the aforementioned issues, this disclosure proposes a method for fine-tuning a translation model, a translation method, an apparatus, and an electronic device.
[0026] Figure 1 The diagram is based on the first embodiment of this disclosure. It should be noted that the fine-tuning method for the translation model in this embodiment can be applied to a fine-tuning device for the translation model. This device can be configured in an electronic device so that the electronic device can perform the fine-tuning function of the translation model. The following embodiments use an electronic device as an example for illustration.
[0027] Among them, electronic devices can be any device with computing capabilities, such as personal computers (PCs), mobile terminals, servers, etc. Mobile terminals can be, for example, in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, smart speakers, servers, server clusters, and other hardware devices with various operating systems, touch screens and / or displays.
[0028] The fine-tuning device for the translation model can also be software within an electronic device, such as fine-tuning software for the translation model. In the following embodiments, an electronic device is used as an example for illustration.
[0029] like Figure 1 As shown, the fine-tuning method for this translation model may include the following steps:
[0030] Step 101: Obtain the translation model to be fine-tuned, and the general domain parallel corpus set of the translation model; the general domain parallel corpus set includes parallel corpora between at least two target languages.
[0031] It should be noted that the target languages can be any two languages. These two languages are not limited to two languages from different countries; they can also be different languages within the same country; or two languages from different regions within the same country; or two languages from different periods within the same country. Correspondingly, the translation model can be a translation model between these two target languages; or, the translation model can be a multilingual translation model with translation capabilities for both target languages.
[0032] In this embodiment of the disclosure, parallel corpora between at least two target languages refer to corpora of at least two target languages, and there is a translation relationship between the corpora of at least two target languages. The translation relationship means that corpora of one target language can be translated into corpora of another target language.
[0033] The general-domain parallel corpus set can include multiple general-domain corpora. A general-domain parallel corpus refers to parallel corpora from any domain. Domains include, for example, computer technology, information technology, communication technology, biotechnology, and medical technology, and can be set according to actual needs.
[0034] Step 102: Obtain a monolingual corpus of the target language within the target domain.
[0035] In this embodiment of the disclosure, the number of target domains can be at least one. Specifically, assuming the fine-tuned translation model is needed to translate corpora from a certain technical field, the target domain can be that technical field.
[0036] The monolingual corpus of the target language within the target domain may include multiple corpora within the target domain. The methods for acquiring these corpora may include at least one of the following: extracting from articles within the target domain; extracting from dialogue records within the target domain, etc., without specific limitations.
[0037] When acquiring corpora within the target domain by crawling articles from that domain, the process can involve, for example, segmenting and preprocessing the crawled articles to extract individual sentences; and then using these sentences as the corpus within the target domain. The preprocessing steps, such as formatting adjustments and removal of meaningless characters, can be configured according to specific needs.
[0038] Step 103: Based on the monolingual corpus of the target language in the target domain, update the vocabulary of the target language in the translation model to obtain the translation model with updated vocabulary.
[0039] In this embodiment of the disclosure, the electronic device may perform step 103 as follows: determine the target domain vocabulary of the target language based on the monolingual corpus of the target language in the target domain; add the words in the target domain vocabulary of the target language to the vocabulary of the target language in the translation model to obtain the translation model with updated vocabulary.
[0040] The target language monolingual corpus within the target domain contains a large number of words from that domain. Based on this corpus, a target domain vocabulary is determined, and the vocabulary of the target language in the translation model is then updated. This ensures that the target language vocabulary in the translation model includes a large number of words from the target domain, enabling the adjusted translation model to translate target language words more effectively and improving translation accuracy.
[0041] In one embodiment of this disclosure, the process by which an electronic device determines the target domain vocabulary of a target language can be, for example, training an initial word segmentation model using monolingual corpora from a monolingual corpus set to obtain a trained word segmentation model and a target domain vocabulary of the target language.
[0042] The monolingual corpus can be, for example, sentences. The open-source sub-word toolkit SentencePiece can treat sentences in a monolingual corpus of the target language within the target domain as a whole and then break them down into fragments. The word segmentation model can be, for example, a Byte Pair Encoding (BPE) model. This model can integrate the multiple fragments obtained from SentencePiece to obtain multiple words for the target domain. Then, based on the integration result, the parameters of the word segmentation model are adjusted, and the words are re-integrated to generate a target domain vocabulary for the target language.
[0043] Among them, the word segmentation model has a large computational load and high computational accuracy, which can improve the accuracy of the target domain vocabulary of the target language.
[0044] In this embodiment of the disclosure, there may be duplicate words between the target language's target domain vocabulary and the target language's vocabulary in the translation model. To address this, the electronic device can first determine the intersection between the target language's target domain vocabulary and the target language's vocabulary in the translation model; after adding words from the target language's target domain vocabulary to the target language's vocabulary in the translation model, one of each duplicate word in the target language's vocabulary in the translation model (i.e., words from the intersection) is deleted, ensuring that there are no duplicate words in the target language's vocabulary in the translation model.
[0045] Specifically, for newly added words in the target language vocabulary of the translation model, the electronic device can, in conjunction with the translation model, randomly assign a label to the word and perform random vectorization on it. It should be noted that the vector obtained after random vectorization can be adjusted during the fine-tuning process of the translation model, thereby improving the accuracy of the word's vector.
[0046] Step 104: Fine-tune the translation model after updating the vocabulary based on the general domain parallel corpus set.
[0047] In this embodiment of the disclosure, the electronic device performing step 104 may, for example, involve inputting the first target language corpus from each parallel corpus in the general domain parallel corpus set into the vocabulary-updated translation model; combining the updated vocabulary with the updated vocabulary, the translation model vectorizes each word in the first target language corpus to obtain a first vector; encoding and decoding the first vector to obtain a second vector; combining the second vector and the updated vocabulary to determine each predicted word in the second target language corpus, thereby determining the predicted corpus of the second target language; combining the predicted corpus of the second target language, the second target language corpus in the parallel corpus, and the loss function to determine the value of the loss function; and adjusting the parameters of the vocabulary-updated translation model based on the value of the loss function. During the parameter adjustment process of the vocabulary-updated translation model, the vectors corresponding to each word in the updated vocabulary are simultaneously adjusted; thereby improving the accuracy of word vectorization in each target domain, and thus improving the translation efficiency and accuracy of target domain terminology.
[0048] The fine-tuning method for the translation model in this embodiment of the present disclosure involves obtaining a translation model to be fine-tuned and a set of parallel corpora for a general domain of the translation model. The set of parallel corpora for the general domain includes parallel corpora between at least two target languages. A set of monolingual corpora for the target language within the target domain is obtained. Based on the set of monolingual corpora for the target language within the target domain, the vocabulary of the target language in the translation model is updated to obtain a translation model with an updated vocabulary. The updated translation model is then fine-tuned based on the set of parallel corpora for the general domain. The updating of the vocabulary of the target language in the translation model can improve the accuracy of vectorization of words in the target domain within the vocabulary during model fine-tuning, thereby improving the translation efficiency of the corpus in the target domain.
[0049] To further improve the translation efficiency of the translation model for corpora in the target domain, it is also possible to fine-tune the translation model by combining it with a set of parallel corpora in the target domain. For example... Figure 2 As shown, Figure 2 This is a schematic diagram based on the second embodiment of the present disclosure. Figure 2 The illustrated embodiment may include the following steps:
[0050] Step 201: Obtain the translation model to be fine-tuned, and the general domain parallel corpus set of the translation model; the general domain parallel corpus set includes parallel corpora between at least two target languages.
[0051] Step 202: Obtain a monolingual corpus of the target language within the target domain.
[0052] Step 203: Based on the monolingual corpus of the target language in the target domain, update the vocabulary of the target language in the translation model to obtain the translation model with updated vocabulary.
[0053] Step 204: Obtain a set of parallel corpora for the target domain.
[0054] In this embodiment of the disclosure, the process of the electronic device performing step 204 may be, for example, determining the parallel corpus of the target domain in the set of parallel corpus of the general domain; and determining the set of parallel corpus of the target domain based on the parallel corpus of the target domain.
[0055] Among them, obtaining parallel corpora of the target domain from the general domain parallel corpus set can reduce the cost of obtaining parallel corpora of the target domain. Moreover, the general domain parallel corpus set is the parallel corpus set used for pre-training of the translation model, and has high accuracy. Combining the general domain parallel corpus set to determine the parallel corpora of the target domain can further improve the accuracy of the determined target domain parallel corpus set.
[0056] In one example of this disclosure, the process by which an electronic device determines parallel corpora of a target domain within a set of parallel corpora of a general domain can be, for example, determining the domain to which the parallel corpora belong in the set of parallel corpora of a general domain; and, if the domain to which they belong is the target domain, determining the parallel corpora as parallel corpora of the target domain.
[0057] The electronic device can acquire a first domain recognition model, which can identify the domain to which the parallel corpus belongs. The first domain recognition model can have multiple domain categories. By processing the parallel corpus, it determines the probability that the parallel corpus belongs to each domain category; then, the domain category with the highest probability is taken as the domain to which the parallel corpus belongs.
[0058] Among them, the first domain identification model is combined to determine the domain of the parallel corpus. The first domain identification model has a large computational load and high computational accuracy, which can improve the accuracy of the parallel corpus of the target domain.
[0059] In another example, the process by which an electronic device determines parallel corpora of the target domain within a set of parallel corpora of the general domain can be, for example, determining whether the parallel corpora in the set of parallel corpora of the general domain belong to the target domain or not; and when the parallel corpora belong to the target domain, determining the parallel corpora as parallel corpora of the target domain.
[0060] The electronic device can acquire a second domain recognition model, which can identify whether parallel corpora belong to the target domain. The second domain recognition model can have two categories: one indicating that it belongs to the target domain, and the other indicating that it does not. By processing the parallel corpora, the category of the parallel corpora is determined; when the category indicates that it belongs to the target domain, the parallel corpora are determined to be parallel corpora within the target domain; when the category indicates that it does not belong to the target domain, the parallel corpora are determined to be parallel corpora within a domain other than the target domain.
[0061] Step 205: Fine-tune the translation model after updating the vocabulary based on the general domain parallel corpus set and the target domain parallel corpus set.
[0062] In one example of this disclosure, the electronic device performing step 205 may, for example, perform fine-tuning of the translation model after vocabulary update based on a general domain parallel corpus set to obtain a first translation model; perform fine-tuning of the translation model after vocabulary update based on a target domain parallel corpus set to obtain a second translation model; perform mean averaging on the parameters of the first and second translation models to obtain a processed model; and determine the processed model as the fine-tuned translation model.
[0063] In another example, the electronic device performing step 205 may be as follows: fine-tuning the translation model updated with the vocabulary based on a set of parallel corpora in a general domain to obtain a first translation model; fine-tuning the first translation model based on a set of parallel corpora in the target domain to obtain a second translation model; and identifying the second translation model as the fine-tuned translation model.
[0064] In this embodiment of the disclosure, to further improve the accuracy of the translation model after fine-tuning, data preprocessing and data cleaning can be performed on the general domain parallel corpus set and the target domain parallel corpus set. The data preprocessing includes at least one of the following: garbled character removal, non-printable character removal, format conversion, and link removal. The data cleaning includes at least one of the following: removal of duplicate parallel corpora, removal of parallel corpora with repeated words, removal based on parallel corpus length, removal based on the length ratio between two corpora in a parallel corpus, and removal based on the matching degree between two corpora in a parallel corpus.
[0065] This includes format conversion processing, such as converting traditional Chinese characters to simplified Chinese characters. It also includes link removal processing, such as removing webpage links from parallel corpora.
[0066] The removal based on parallel corpus length can refer to removing parallel corpora when their length is less than a first length threshold, or when their length is greater than a second length threshold. The removal based on the length ratio between two parallel corpora can refer to removing parallel corpora when the length ratio between two parallel corpora is less than a first ratio threshold, or when the length ratio between two parallel corpora is greater than a second ratio threshold.
[0067] The removal process based on the matching degree between two corpora in a parallel corpus can be described as follows: determining the matching degree between two corpora in a parallel corpus; and removing corpora from the parallel corpus if the matching degree is less than a matching degree threshold.
[0068] It should be noted that for details of steps 201 to 203, please refer to [the relevant documentation / reference]. Figure 1 Steps 101 to 103 in the illustrated embodiment will not be described in detail here.
[0069] The fine-tuning method for the translation model in this embodiment of the present disclosure involves obtaining a translation model to be fine-tuned and a set of parallel corpora for a general domain of the translation model. The set of parallel corpora for the general domain includes parallel corpora between at least two target languages. A set of monolingual corpora for the target language within the target domain is obtained. Based on the monolingual corpora for the target language within the target domain, the vocabulary of the target language in the translation model is updated to obtain a translation model with an updated vocabulary. A set of parallel corpora for the target domain is obtained. The translation model with an updated vocabulary is fine-tuned based on the set of parallel corpora for the general domain and the set of parallel corpora for the target domain. Updating the vocabulary of the target language in the translation model can improve the accuracy of vectorization of words in the target domain within the vocabulary during model fine-tuning. Fine-tuning the translation model with an updated vocabulary by combining the set of parallel corpora for the target domain can further improve the accuracy of the fine-tuned translation model.
[0070] To further improve the translation efficiency of the translation model for corpora in the target domain, fine-tuning of the model can be achieved by combining parallel corpora and pseudo-parallel corpora in the target domain. For example... Figure 3 As shown, Figure 3 This is a schematic diagram based on the third embodiment of the present disclosure. Figure 3 The illustrated embodiment may include the following steps:
[0071] Step 301: Obtain the translation model to be fine-tuned, and the general domain parallel corpus set of the translation model; the general domain parallel corpus set includes parallel corpora between at least two target languages.
[0072] Step 302: Obtain a monolingual corpus of the target language within the target domain.
[0073] Step 303: Based on the monolingual corpus of the target language in the target domain, update the vocabulary of the target language in the translation model to obtain the translation model with updated vocabulary.
[0074] Step 304: Obtain a set of parallel corpora for the target domain.
[0075] Step 305: Obtain a set of pseudo-parallel corpora for the target domain.
[0076] In one embodiment of this disclosure, at least two target languages include a first target language and a second target language; a set of pseudo-parallel corpora in the target domain is determined by combining a reverse translation model; and a translation model is used to translate the corpus of the first target language to obtain the corpus of the second target language. Correspondingly, the electronic device performing step 305 may, for example, involve: acquiring the corpus of the second target language in the target domain; translating the corpus of the second target language in the target domain using a reverse translation model to obtain the corpus of the first target language in the target domain; determining pseudo-parallel corpora between at least two target languages in the target domain based on the corpus of the second target language and the corpus of the first target language in the target domain; and determining a set of pseudo-parallel corpora in the target domain based on the pseudo-parallel corpora between at least two target languages in the target domain.
[0077] The translation model is used to translate the corpus of the first target language to obtain the corpus of the second target language. The reverse translation model is used to translate the corpus of the second target language to obtain the corpus of the first target language.
[0078] In another example of this disclosure embodiment, at least two target languages include a first target language and a second target language. Correspondingly, the electronic device performing step 305 may, for example, involve: acquiring parallel corpora between the second target language and non-target languages within the target domain; translating the monolingual corpora of the non-target languages in the parallel corpora according to a translation model to obtain monolingual corpora of the first target language; determining pseudo-parallel corpora between at least two target languages within the target domain based on the monolingual corpora of the first target language and the monolingual corpora of the second target language in the parallel corpora; and determining a set of pseudo-parallel corpora in the target domain based on the pseudo-parallel corpora between at least two target languages within the target domain.
[0079] Step 306: Fine-tune the translation model after updating the vocabulary based on the general domain parallel corpus set, the target domain parallel corpus set, and the target domain pseudo-parallel corpus set.
[0080] In this embodiment of the disclosure, the electronic device performing step 306 may, for example, involve fine-tuning the translation model updated with a vocabulary based on a general domain parallel corpus set to obtain a first translation model; fine-tuning the translation model updated with a vocabulary based on a general domain parallel corpus set, a target domain parallel corpus set, and a target domain pseudo-parallel corpus set to obtain a second translation model; fine-tuning the first translation model based on a target domain parallel corpus set to obtain a third translation model; fine-tuning the second translation model based on a target domain parallel corpus set to obtain a fourth translation model; and determining the fine-tuned translation model based on at least two of the first, second, third, and fourth translation models.
[0081] In this process, the electronic device uses different sets or combinations of different sets to fine-tune the translation model after the vocabulary update, resulting in multiple translation models. Then, by combining the multiple translation models, the fine-tuned translation model is determined. This allows for the comprehensive analysis of the fine-tuning results of multiple translation models to determine the fine-tuned translation model, thereby further improving the translation accuracy of the determined fine-tuned translation model.
[0082] The process by which the electronic device determines the fine-tuned translation model based on at least two of the first, second, third, and fourth translation models can be, for example, by performing parameter averaging on at least two of the first, second, third, and fourth translation models to obtain the parameter-averaged translation model; and then determining the parameter-averaged translation model as the fine-tuned translation model.
[0083] The electronic device performs parameter averaging on at least two of the first, second, third, and fourth translation models, so that the fine-tuned translation model can comprehensively consider the fine-tuning results of at least two translation models, thereby further improving the translation accuracy of the determined fine-tuned translation model.
[0084] In this embodiment of the disclosure, the number of translation models to be fine-tuned can be at least two; the at least two translation models have different parameter counts and / or structures; and the number of fine-tuned translation models is at least two. The setting of at least two fine-tuned translation models allows the translation result to be determined by combining these two models during the translation process, thereby improving the accuracy of the translation result.
[0085] In this embodiment of the disclosure, as an alternative to step 306, the electronic device may perform the following process: fine-tuning the translation model after vocabulary update based on a general domain parallel corpus set and a target domain pseudo-parallel corpus set.
[0086] The specific process of fine-tuning the translation model after updating the vocabulary based on the general domain parallel corpus set and the target domain pseudo-parallel corpus set can be found in the detailed description of step 306.
[0087] It should be noted that for details of steps 301 to 303, please refer to [the relevant documentation / reference]. Figure 1 Steps 101 to 103 in the illustrated embodiment will not be described in detail here.
[0088] The fine-tuning method for the translation model in this embodiment of the present disclosure involves obtaining a translation model to be fine-tuned and a set of parallel corpora for a general domain of the translation model. The set of parallel corpora for the general domain includes parallel corpora between at least two target languages. A set of monolingual corpora for the target language within the target domain is obtained. Based on the monolingual corpora for the target language within the target domain, the vocabulary of the target language in the translation model is updated to obtain a translation model with an updated vocabulary. A set of parallel corpora for the target domain is obtained. A set of pseudo-parallel corpora for the target domain is obtained. The translation model with an updated vocabulary is fine-tuned based on the set of parallel corpora for the general domain, the set of parallel corpora for the target domain, and the set of pseudo-parallel corpora for the target domain. Updating the vocabulary of the target language in the translation model can improve the accuracy of vectorization of words in the target domain within the vocabulary during model fine-tuning. Fine-tuning the translation model with an updated vocabulary by combining the set of parallel corpora for the target domain and the set of pseudo-parallel corpora for the target domain can further improve the accuracy of the fine-tuned translation model.
[0089] The following example illustrates this. For example... Figure 4 The image shown is a schematic diagram illustrating the fine-tuning of the translation model. Figure 4 The process may include the following steps: Step 401, obtaining general parallel corpora (a set of general domain parallel corpora), domain-specific parallel corpora (a set of target domain parallel corpora), and pseudo-parallel corpora (a set of target domain pseudo-parallel corpora). Step 402, performing data preprocessing and data cleaning on the above three types of parallel corpora. Step 403, vocabulary expansion, that is, expanding the vocabulary of the target language in the translation model by combining the monolingual corpus set of the target language within the target domain. Step 404, model fine-tuning, that is, fine-tuning the translation model to be fine-tuned based on the three sets after data cleaning, obtaining Model 1 to Model n. Step 405, model ensemble, that is, performing parameter mean averaging on Model 1 to Model n, obtaining the fine-tuned translation model.
[0090] Figure 5 This is a schematic diagram based on the fourth embodiment of this disclosure. It should be noted that the translation method of this embodiment can be applied to a translation device, which can be configured in an electronic device to enable the electronic device to perform translation functions. The following embodiments use an electronic device as an example for illustration.
[0091] Among them, electronic devices can be any device with computing capabilities, such as personal computers (PCs), mobile terminals, servers, etc. Mobile terminals can be, for example, in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, smart speakers, servers, server clusters, and other hardware devices with various operating systems, touch screens and / or displays.
[0092] The translation device can also be software within an electronic device, such as translation software. In the following embodiments, an electronic device is used as an example for illustration.
[0093] like Figure 5 As shown, the fine-tuning method for this translation model may include the following steps:
[0094] Step 501: Obtain the fine-tuned translation model; the fine-tuned translation model is determined based on the fine-tuning method of the above translation model.
[0095] In this embodiment of the disclosure, the method for determining the fine-tuned translation model can be referred to... Figures 1 to 4 The descriptions in the embodiments will not be elaborated further here.
[0096] Step 502: Obtain the corpus of the first target language to be translated.
[0097] Step 503: Combining the corpus of the first target language and the fine-tuned translation model, determine the corpus of the second target language corresponding to the corpus of the first target language.
[0098] In this embodiment of the disclosure, the number of fine-tuned translation models is at least two; the number of parameters of the at least two fine-tuned translation models may be the same or different; the structure of the at least two fine-tuned translation models may be the same or different. Correspondingly, the electronic device performing step 503 may, for example, input the corpus of the first target language into the at least two fine-tuned translation models respectively, obtain at least two translation results; and determine the corpus of the second target language corresponding to the corpus of the first target language based on the at least two translation results.
[0099] In one example, the process by which the electronic device determines the corpus of the second target language corresponding to the corpus of the first target language based on at least two translation results can be as follows: for each character position, determine the character at the character position in at least two translation results and the probability corresponding to the character; select the character with the highest probability and determine it as the target character at the character position; and determine the corpus of the second target language corresponding to the corpus of the first target language based on the target characters at each character position.
[0100] In another example, the process by which the electronic device determines the corpus of the second target language corresponding to the corpus of the first target language based on at least two translation results can be as follows: for each character position, determine the character at that position and the probability corresponding to the character in at least two translation results; determine the quantity of each character at that position; select the character with the largest quantity and determine it as the target character at that position; if the characters at that position are all different in at least two translation results, select the character with the largest probability and determine it as the target character at that position; and determine the corpus of the second target language corresponding to the corpus of the first target language based on the target characters at each character position.
[0101] The electronic device determines the corpus of the second target language corresponding to the corpus of the first target language based on the characters at each character position in at least two translation results and the probabilities of the characters; it can integrate the translation results of the corpus of the first target language to be translated from at least two fine-tuned models, thereby further improving the accuracy of the determined corpus of the second target language.
[0102] The translation method of this disclosure includes: obtaining a fine-tuned translation model; determining the fine-tuned translation model based on the aforementioned fine-tuning method; obtaining corpus of a first target language to be translated; and determining corpus of a second target language corresponding to the first target language corpus by combining the corpus of the first target language and the fine-tuned translation model. The fine-tuned translation model updates the vocabulary of the target language, which improves the accuracy of vectorization of words in the target domain during model fine-tuning, thereby improving the translation efficiency of corpus in the target domain and the translation accuracy of target domain terminology.
[0103] To achieve the above embodiments, this disclosure also provides a fine-tuning device for a translation model. For example... Figure 6 As shown, Figure 6 This is a schematic diagram according to the fifth embodiment of the present disclosure. The fine-tuning device 60 of the translation model may include: a first acquisition module 601, a second acquisition module 602, an update processing module 603, and a fine-tuning processing module 604.
[0104] The system includes a first acquisition module 601, used to acquire the translation model to be fine-tuned and a set of parallel corpora for the general domain of the translation model; the set of parallel corpora for the general domain includes parallel corpora between at least two target languages; a second acquisition module 602, used to acquire a set of monolingual corpora for the target language within the target domain; an update processing module 603, used to update the vocabulary of the target language in the translation model according to the set of monolingual corpora for the target language within the target domain, to obtain a translation model with an updated vocabulary; and a fine-tuning processing module 604, used to fine-tune the translation model with an updated vocabulary according to the set of parallel corpora for the general domain.
[0105] As one possible implementation of this disclosure, the update processing module 603 includes a first determining unit and an adding unit; the first determining unit is used to determine the target domain vocabulary of the target language based on the monolingual corpus set of the target language in the target domain; the adding unit is used to add words from the target domain vocabulary of the target language to the vocabulary of the target language in the translation model to obtain a translation model with an updated vocabulary.
[0106] As one possible implementation of this disclosure, the first determining unit is specifically used to train an initial word segmentation model using the monolingual corpus in the monolingual corpus set, to obtain the trained word segmentation model and the target domain vocabulary of the target language.
[0107] As one possible implementation of this disclosure, the apparatus further includes: a third acquisition module, configured to acquire a target domain parallel corpus set and / or a target domain pseudo-parallel corpus set; the target domain pseudo-parallel corpus set is determined by combining a reverse translation model or parallel corpus between the target language and a non-target language; correspondingly, the fine-tuning processing module 604 is specifically configured to perform joint fine-tuning processing on the vocabulary-updated translation model based on the general domain parallel corpus set, the target domain parallel corpus set, and / or the target domain pseudo-parallel corpus set.
[0108] As one possible implementation of this disclosure, the third acquisition module includes a second determining unit and a third determining unit; the second determining unit is used to determine the parallel corpus of the target domain in the general domain parallel corpus set; the third determining unit is used to determine the target domain parallel corpus set based on the parallel corpus of the target domain.
[0109] As one possible implementation of this disclosure, the second determining unit is specifically used to determine the domain to which the parallel corpus in the general domain parallel corpus set belongs; and if the domain to which it belongs is the target domain, determine the parallel corpus as the parallel corpus of the target domain.
[0110] As one possible implementation of this disclosure, the at least two target languages include a first target language and a second target language; the translation model is used to translate the corpus of the first target language to obtain the corpus of the second target language; the third acquisition module is specifically used to: acquire the corpus of the second target language in the target domain; combine a reverse translation model to translate the corpus of the second target language in the target domain to obtain the corpus of the first target language in the target domain; determine pseudo-parallel corpus between the at least two target languages in the target domain based on the corpus of the second target language in the target domain and the corpus of the first target language in the target domain; and determine the pseudo-parallel corpus set in the target domain based on the pseudo-parallel corpus between the at least two target languages in the target domain.
[0111] As one possible implementation of this disclosure, the at least two target languages include a first target language and a second target language; the third acquisition module is specifically used to: acquire parallel corpora between the second target language and non-target languages within the target domain; translate the monolingual corpora of the non-target languages in the parallel corpora according to the translation model to obtain monolingual corpora of the first target language; determine pseudo-parallel corpora between the at least two target languages within the target domain based on the monolingual corpora of the first target language and the monolingual corpora of the second target language in the parallel corpora; and determine the target domain pseudo-parallel corpora set based on the pseudo-parallel corpora between the at least two target languages within the target domain.
[0112] As one possible implementation of this disclosure, the fine-tuning processing module 604 includes a first processing unit, a second processing unit, a third processing unit, a fourth processing unit, and a fourth determining unit. The first processing unit is used to fine-tune the vocabulary-updated translation model based on the general domain parallel corpus set to obtain a first translation model. The second processing unit is used to fine-tune the vocabulary-updated translation model based on the general domain parallel corpus set, the target domain parallel corpus set, and the target domain pseudo-parallel corpus set to obtain a second translation model. The third processing unit is used to fine-tune the first translation model based on the target domain parallel corpus set to obtain a third translation model. The fourth processing unit is used to fine-tune the second translation model based on the target domain parallel corpus set to obtain a fourth translation model. The fourth determining unit is used to determine the fine-tuned translation model based on at least two of the first, second, third, and fourth translation models.
[0113] As one possible implementation of this disclosure, the fourth determining unit is specifically used to perform parameter averaging on at least two of the first translation model, the second translation model, the third translation model, and the fourth translation model to obtain a parameter-averaged translation model; and to determine the parameter-averaged translation model as the fine-tuned translation model.
[0114] As one possible implementation of this disclosure, the number of translation models to be fine-tuned is at least two; the number and / or structure of the at least two translation models are different; and the number of translation models after fine-tuning is at least two.
[0115] As one possible implementation of this disclosure, the apparatus further includes: a data processing module, used to perform data preprocessing and data cleaning on the target domain parallel corpus set and the target domain pseudo-parallel corpus set; the data preprocessing includes at least one of the following: garbled character removal processing, non-printable character removal processing, format conversion processing, and link removal processing; the data cleaning processing includes at least one of the following: duplicate parallel corpus removal processing, parallel corpus removal processing with repeated words, removal processing based on the length of parallel corpus, removal processing based on the length ratio between two corpora in the parallel corpus, and removal processing based on the matching degree between two corpora in the parallel corpus.
[0116] The fine-tuning device for the translation model in this embodiment of the present disclosure acquires a translation model to be fine-tuned and a set of parallel corpora for a general domain of the translation model; the set of parallel corpora for the general domain includes parallel corpora between at least two target languages; acquires a set of monolingual corpora for the target language within the target domain; updates the vocabulary of the target language in the translation model according to the set of monolingual corpora for the target language within the target domain to obtain a translation model with an updated vocabulary; and fine-tunes the translation model with an updated vocabulary according to the set of parallel corpora for the general domain. The updating of the vocabulary of the target language in the translation model can improve the accuracy of vectorization of words in the target domain within the vocabulary during model fine-tuning, thereby improving the translation efficiency of the corpus in the target domain and the translation accuracy of professional terms in the target domain.
[0117] To achieve the above embodiments, this disclosure also provides a translation apparatus. For example... Figure 7 As shown, Figure 7 This is a schematic diagram according to the sixth embodiment of the present disclosure. The translation device 70 may include: a first acquisition module 701, a second acquisition module 702, and a determination module 703.
[0118] The first acquisition module 701 is used to acquire the fine-tuned translation model; the fine-tuned translation model is determined based on the fine-tuning method of the above-mentioned translation model; the second acquisition module 702 is used to acquire the corpus of the first target language to be translated; the determination module 703 is used to combine the corpus of the first target language and the fine-tuned translation model to determine the corpus of the second target language corresponding to the corpus of the first target language.
[0119] As one possible implementation of this disclosure, the number of fine-tuned translation models is at least two; the determining module 703 is specifically used to input the corpus of the first target language into at least two of the fine-tuned translation models respectively, and obtain at least two translation results; and determine the corpus of the second target language corresponding to the corpus of the first target language based on the at least two translation results.
[0120] As one possible implementation of this disclosure, the determining module 703 is further configured to: for each character position, determine the character at the character position in the at least two translation results and the probability corresponding to the character; select the character with the highest probability and determine it as the target character at the character position; and determine the corpus of the second target language corresponding to the corpus of the first target language based on the target characters at each character position.
[0121] The translation apparatus of this embodiment acquires a fine-tuned translation model; the fine-tuned translation model is determined based on the fine-tuning method described above; corpus of a first target language to be translated is acquired; and corpus of a second target language corresponding to the corpus of the first target language is determined by combining the corpus of the first target language and the fine-tuned translation model; wherein, the vocabulary of the target language in the fine-tuned translation model is updated, which can improve the accuracy of vectorization of words in the target domain in the vocabulary during model fine-tuning, thereby improving the translation efficiency of the corpus in the target domain.
[0122] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information are all carried out with the consent of the users, and all comply with the provisions of relevant laws and regulations, and do not violate public order and good morals.
[0123] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0124] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0125] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0126] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0127] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a translation model fine-tuning method or a translation method. For example, in some embodiments, the translation model fine-tuning method or the translation method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the translation model fine-tuning method or the translation method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to perform a fine-tuning method for the translation model or a translation method.
[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0129] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0130] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0133] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0134] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for fine-tuning a translation model, the method comprising: obtaining a translation model to be fine-tuned, and a general-domain parallel corpus set of the translation model; the general-domain parallel corpus set comprising parallel corpora between at least two target languages, the parallel corpora between the at least two target languages referring to corpora of the at least two target languages and having a translation relationship therebetween; obtaining a monolingual corpus set of the target language in a target domain; determining a target-domain vocabulary of the target language according to the monolingual corpus set of the target language in the target domain; adding words in the target-domain vocabulary of the target language to a vocabulary of the target language in the translation model to obtain a vocabulary-updated translation model; obtaining a target-domain parallel corpus set and / or a target-domain pseudo-parallel corpus set; the target-domain pseudo-parallel corpus set being determined in combination with a back-translation model or parallel corpora between the target language and a non-target language; fine-tuning the vocabulary-updated translation model according to the general-domain parallel corpus set to obtain a first translation model; fine-tuning the vocabulary-updated translation model according to the general-domain parallel corpus set, the target-domain parallel corpus set and the target-domain pseudo-parallel corpus set to obtain a second translation model; fine-tuning the first translation model according to the target-domain parallel corpus set to obtain a third translation model; fine-tuning the second translation model according to the target-domain parallel corpus set to obtain a fourth translation model; determining a fine-tuned translation model according to at least two of the first translation model, the second translation model, the third translation model and the fourth translation model.
2. The method of claim 1, wherein, The determining of the target-domain vocabulary of the target language according to the monolingual corpus set of the target language in the target domain comprises: training an initial word segmentation model using monolingual corpora in the monolingual corpus set to obtain a trained word segmentation model and the target-domain vocabulary of the target language.
3. The method of claim 1, wherein, The obtaining of the target-domain parallel corpus set comprises: determining parallel corpora of the target domain in the general-domain parallel corpus set; determining the target-domain parallel corpus set according to the parallel corpora of the target domain.
4. The method of claim 3, wherein, The determining of the parallel corpora of the target domain in the general-domain parallel corpus set comprises: determining a domain to which parallel corpora in the general-domain parallel corpus set belong; in a case where the domain is the target domain, determining the parallel corpora as parallel corpora of the target domain.
5. The method of claim 1, wherein, The at least two target languages comprise a first target language and a second target language; the translation model is used to translate corpora of the first target language to obtain corpora of the second target language; The obtaining of the target-domain pseudo-parallel corpus set comprises: obtaining corpora of the second target language in the target domain; performing translation processing on the corpora of the second target language in the target domain in combination with a back-translation model to obtain corpora of the first target language in the target domain; determine, according to the parallel corpus between the second target language and the non-target language in the target domain, the pseudo-parallel corpus between the at least two target languages in the target domain; determine, according to the pseudo-parallel corpus between the at least two target languages in the target domain, the target domain pseudo-parallel corpus set.
6. The method of claim 1 or 5, wherein, The at least two target languages include a first target language and a second target language; obtaining the target domain pseudo-parallel corpus set includes: obtaining parallel corpus between the second target language and a non-target language in the target domain; performing translation processing on the monolingual corpus of the non-target language in the parallel corpus according to the translation model to obtain monolingual corpus of the first target language; determine, according to the monolingual corpus of the first target language and the monolingual corpus of the second target language in the parallel corpus, the pseudo-parallel corpus between the at least two target languages in the target domain; determine, according to the pseudo-parallel corpus between the at least two target languages in the target domain, the target domain pseudo-parallel corpus set.
7. The method of claim 1, wherein, The determining, according to at least two of the first translation model, the second translation model, the third translation model and the fourth translation model, of the fine-tuned translation model includes: performing parameter mean value processing on at least two of the first translation model, the second translation model, the third translation model and the fourth translation model to obtain a parameter mean value processed translation model; determining the parameter mean value processed translation model as the fine-tuned translation model.
8. The method of claim 1, wherein, The number of the translation models to be fine-tuned is at least two; the parameter amount and / or structure of the at least two translation models are different; The number of the fine-tuned translation models is at least two.
9. The method of claim 1, wherein, The method further includes: performing data preprocessing and data cleaning processing on the target domain parallel corpus set and the target domain pseudo-parallel corpus set; The data preprocessing includes at least one of the following: removing processing of garbled code, removing processing of non-printing characters, format conversion processing, link removal processing; The data cleaning processing includes at least one of the following: removing processing of duplicate parallel corpus, removing processing of parallel corpus with repeated words, removing processing based on the length of parallel corpus, removing processing based on the length ratio between two corpora in parallel corpus, removing processing based on the matching degree between two corpora in parallel corpus.
10. A translation method, the method comprising: obtaining a fine-tuned translation model; the fine-tuned translation model is determined based on the fine-tuning method of the translation model in any one of claims 1 to 9; obtaining a first target language corpus to be translated; determining the second target language corpus corresponding to the first target language corpus in combination with the first target language corpus and the fine-tuned translation model.
11. The method of claim 10, wherein, The number of the fine-tuned translation models is at least two; The combination of the first target language corpus and the fine-tuned translation model determines the second target language corpus corresponding to the first target language corpus, which includes: Input the first target language corpus into at least two fine-tuned translation models respectively, and obtain at least two translation results; According to the at least two translation results, determine the second target language corpus corresponding to the first target language corpus.
12. The method of claim 11, wherein, According to the at least two translation results, determine the second target language corpus corresponding to the first target language corpus, which includes: For each character position, determine the character at the character position in the at least two translation results and the corresponding probability of the character; Select the character with the maximum corresponding probability as the target character at the character position; According to the target character at each character position, determine the second target language corpus corresponding to the first target language corpus.
13. A fine-tuning device for a translation model, the device comprising: A first acquisition module for acquiring a translation model to be fine-tuned and a general domain parallel corpus set of the translation model; the general domain parallel corpus set includes parallel corpora between at least two target languages, and the parallel corpora between the at least two target languages refer to corpora of at least two target languages, and there is a translation relationship between the corpora of at least two target languages; A second acquisition module for acquiring a monolingual corpus set of the target language in the target domain; An update processing module for updating the vocabulary of the target language in the translation model according to the monolingual corpus set of the target language in the target domain to obtain a vocabulary-updated translation model; A fine-tuning processing module for fine-tuning the vocabulary-updated translation model according to the general domain parallel corpus set; The update processing module includes a first determination unit and an adding unit; The first determination unit is configured to determine the target domain vocabulary of the target language according to the monolingual corpus set of the target language in the target domain; The adding unit is configured to add the words in the target domain vocabulary of the target language to the vocabulary of the target language in the translation model to obtain a vocabulary-updated translation model; A third acquisition module for acquiring a target domain parallel corpus set and / or a target domain pseudo-parallel corpus set; the target domain pseudo-parallel corpus set is determined by combining a back-translation model or parallel corpora between the target language and a non-target language; The fine-tuning processing module includes a first processing unit, a second processing unit, a third processing unit, a fourth processing unit, and a fourth determination unit; The first processing unit is configured to fine-tune the vocabulary-updated translation model according to the general domain parallel corpus set to obtain a first translation model; The second processing unit is configured to fine-tune the vocabulary-updated translation model according to the general domain parallel corpus set, the target domain parallel corpus set, and the target domain pseudo-parallel corpus set to obtain a second translation model; The third processing unit is configured to fine-tune the first translation model according to the target-domain parallel corpus set, to obtain a third translation model. The fourth processing unit is configured to fine-tune the second translation model according to the target-domain parallel corpus set, to obtain a fourth translation model. The fourth determination unit is configured to determine a fine-tuned translation model according to at least two translation models from among the first translation model, the second translation model, the third translation model, and the fourth translation model.
14. The apparatus of claim 13, wherein, The first determination unit is specifically configured to, train an initial segmentation model using monolingual corpus in the monolingual corpus set, to obtain a trained segmentation model and a target-domain vocabulary of the target language.
15. The apparatus of claim 13, wherein, The third obtaining module includes a second determination unit and a third determination unit. The second determination unit is configured to determine a target-domain parallel corpus in the general-domain parallel corpus set. The third determination unit is configured to determine a target-domain parallel corpus set according to the target-domain parallel corpus.
16. The apparatus of claim 15, wherein, The second determination unit is specifically configured to, determine a domain to which a parallel corpus in the general-domain parallel corpus set belongs; in a case where the domain is the target domain, determine the parallel corpus as a target-domain parallel corpus.
17. The apparatus of claim 13, wherein, The at least two target languages include a first target language and a second target language; the translation model is configured to translate corpus of the first target language to obtain corpus of the second target language; and the third obtaining module is specifically configured to, obtain corpus of the second target language in the target domain; translate the corpus of the second target language in the target domain by using a back-translation model, to obtain corpus of the first target language in the target domain; determine pseudo-parallel corpus between the at least two target languages in the target domain according to the corpus of the second target language in the target domain and the corpus of the first target language in the target domain; determine a target-domain pseudo-parallel corpus set according to the pseudo-parallel corpus between the at least two target languages in the target domain.
18. The apparatus of claim 13 or 17, wherein, The at least two target languages include a first target language and a second target language; and the third obtaining module is specifically configured to, obtain parallel corpus between the second target language and a non-target language in the target domain; translate monolingual corpus of the non-target language in the parallel corpus by using the translation model, to obtain monolingual corpus of the first target language; determine pseudo-parallel corpus between the at least two target languages in the target domain according to the monolingual corpus of the first target language and monolingual corpus of the second target language in the parallel corpus; determine a target-domain pseudo-parallel corpus set according to the pseudo-parallel corpus between the at least two target languages in the target domain.
19. The apparatus of claim 13, wherein, The fourth determination unit is specifically configured to, perform parameter mean value processing on at least two translation models from among the first translation model, the second translation model, the third translation model, and the fourth translation model, to obtain a translation model after parameter mean value processing. The parameter-averaged translation model is determined as the fine-tuned translation model.
20. The apparatus of claim 13, wherein, The number of the translation models to be fine-tuned is at least two; the number of parameters and / or the structure of the at least two translation models are different; The number of the fine-tuned translation models is at least two.
21. The apparatus of claim 13, wherein, The apparatus further comprises: a data processing module configured to perform data preprocessing and data cleaning processing on the target domain parallel corpus set and the target domain pseudo-parallel corpus set; The data preprocessing comprises at least one of the following: removing non-printable characters, format conversion processing, link removal processing; The data cleaning processing comprises at least one of the following: removing duplicate parallel corpora, removing parallel corpora with repeated words, removing parallel corpora based on the length of the parallel corpora, removing parallel corpora based on the length ratio between the two corpora in the parallel corpora, and removing parallel corpora based on the matching degree between the two corpora in the parallel corpora.
22. A translation apparatus, comprising: a first obtaining module configured to obtain fine-tuned translation models; the fine-tuned translation models are determined based on the fine-tuning method of the translation model in any one of claims 1 to 9; a second obtaining module configured to obtain a first target language corpus to be translated; a determining module configured to determine a second target language corpus corresponding to the first target language corpus based on the first target language corpus and the fine-tuned translation models.
23. The apparatus of claim 22, wherein, The number of the fine-tuned translation models is at least two; the determining module is specifically configured to, input the first target language corpus into at least two fine-tuned translation models respectively, and obtain at least two translation results; determine the second target language corpus corresponding to the first target language corpus based on the at least two translation results.
24. The apparatus of claim 23, wherein, The determining module is specifically further configured to, for each character position, determine the character at the character position in the at least two translation results and the probability corresponding to the character; select the character with the maximum corresponding probability as the target character at the character position; and determine the second target language corpus corresponding to the first target language corpus based on the target characters at the respective character positions.
25. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 9; or perform the method of any one of claims 10 to 12.
26. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1 to 9; or perform the method of any one of claims 10 to 12.
27. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9; or, implements the method according to any one of claims 10 to 12.
Citation Information
Patent Citations
Machine translation model obtaining method and device, text translation method and device and storage medium
CN111859994A
Text information translation method and device, electronic equipment and storage medium
CN112270200A