A Cross-Lingual Fine-Tuning Method, Device, Equipment and Storage Medium

By fine-tuning the pre-trained model across languages, and using source synonyms and target synonyms to replace the translation sample pair data, the problem of insufficient cross-language alignment ability of the pre-trained model is solved, and the fine-tuning effect of the translation task is improved.

CN115034316BActive Publication Date: 2025-07-18JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210685991.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-07-18
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

The existing pre-trained models trained on monolingual data lead to insufficient cross-lingual alignment capabilities, resulting in poor fine-tuning effects in downstream translation tasks.

Method used

By replacing target synonyms for the source language text and replacing source synonyms for the target language text, constructing translation samples to fine-tune across languages, and introducing cross-language alignment information.

Benefits of technology

The cross-lingual differences between upstream self-supervised tasks and downstream translation tasks are narrowed, and the fine-tuning effect of pre-trained models in downstream translation tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115034316B_ABST
    Figure CN115034316B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a cross-lingual fine-tuning method, device, equipment, and storage medium. The method includes: obtaining translation sample pair data, where the translation sample pair data includes a first source text corresponding to a source language and a first target text corresponding to a target language; performing target near-synonym replacement on some source words in the first source text to determine a second source text after replacement, where the target near-synonym refers to a target word with a semantic similar to the source word; performing source near-synonym replacement on some target words in the first target text to determine a second target text after replacement, where the source near-synonym refers to a source word with a semantic similar to the target word; based on the first source text, the second source text, the first target text, and the second target text, performing cross-lingual fine-tuning on a pre-trained model to determine a pre-fine-tuned model. Through the technical solution of the embodiment of the present invention, the cross-lingual difference between the upstream self-supervised task and the downstream cross-lingual task can be reduced, thereby improving the fine-tuning effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of super deep learning technology, and in particular, to a cross-lingual fine-tuning method, device, equipment, and storage medium. Background Art

[0002] With the rapid development of computer technology, pre-trained models have gradually become a research hotspot in the industry due to their large parameter scale, strong general ability, and good comprehensive performance.

[0003] Currently, pre-trained models are upstream general models trained in a self-supervised manner on large-scale monolingual data. Since their parameters contain rich monolingual knowledge, through the direct fine-tuning paradigm of continued training in downstream tasks, they can be applied to application scenarios such as text classification, dialogue, and summarization tasks.

[0004] However, in the process of implementing the present invention, the inventors found that there are at least the following problems in the existing technology:

[0005] The upstream self-supervised task is only trained on monolingual data. Although it can effectively enhance the pre-trained model's mastery of monolingual knowledge, it does not train the cross-lingual alignment ability of the pre-trained model. Therefore, when the downstream task is a cross-lingual task such as translation, directly fine-tuning the pre-trained model will have obvious cross-lingual differences, thereby reducing the fine-tuning effect. Summary of the Invention

[0006] The embodiments of the present invention provide a cross-lingual fine-tuning method, device, equipment, and storage medium to narrow the cross-lingual differences between the upstream self-supervised task and the downstream translation task, thereby improving the fine-tuning effect.

[0007] In a first aspect, the embodiments of the present invention provide a cross-lingual fine-tuning method, including:

[0008] Obtain translation sample pair data, where the translation sample pair data includes a first source text corresponding to a source language and a first target text corresponding to a target language;

[0009] Replace some source words in the first source text with target near-synonyms to determine a second source text after replacement, where the target near-synonym refers to a target word with a similar semantic meaning to the source word;

[0010] Replace some target words in the first target text with source near-synonyms to determine a second target text after replacement, where the source near-synonym refers to a source word with a similar semantic meaning to the target word;

[0011] Based on the first source text, the second source text, the first target text, and the second target text, perform cross-lingual fine-tuning on the pre-trained model to determine a pre-fine-tuned model.

[0012] In a second aspect, an embodiment of the present invention further provides a cross-lingual fine-tuning device, including:

[0013] A translation sample pair data acquisition module, configured to acquire translation sample pair data, where the translation sample pair data includes a first source text corresponding to a source language and a first target text corresponding to a target language;

[0014] A source word replacement module, configured to perform target near-synonym replacement on some source words in the first source text to determine a second source text after replacement, where the target near-synonym refers to a target word with a semantic meaning similar to that of the source word;

[0015] A target word replacement module, configured to perform source near-synonym replacement on some target words in the first target text to determine a second target text after replacement, where the source near-synonym refers to a source word with a semantic meaning similar to that of the target word;

[0016] A pre-fine-tuning model determination module, configured to perform cross-lingual fine-tuning on a pre-trained model based on the first source text, the second source text, the first target text, and the second target text to determine a pre-fine-tuning model.

[0017] In a third aspect, an embodiment of the present invention further provides an electronic device, where the electronic device includes:

[0018] One or more processors;

[0019] A memory, configured to store one or more programs;

[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the cross-lingual fine-tuning method provided in any embodiment of the present invention.

[0021] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the cross-lingual fine-tuning method provided in any embodiment of the present invention.

[0022] The embodiments in the above-mentioned invention have the following advantages or beneficial effects:

[0023] By performing target near-synonym replacement on some source words in the first source text corresponding to the source language, a second source text containing both source words and target words is obtained, and by performing source near-synonym replacement on some target words in the first target text corresponding to the target language, a second target text containing both source words and target words is obtained, thus achieving code-switching of near-synonyms. The first source text, the second source text, the first target text, and the second target text are used as training data to perform cross-lingual fine-tuning on the pre-trained model. Therefore, cross-lingual alignment information can be introduced into the obtained pre-fine-tuned model, narrowing the cross-lingual differences in the upstream self-supervised task and the downstream translation task, achieving a soft landing of the pre-trained model in the downstream translation task, and further improving the fine-tuning effect of the model in the downstream translation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a flowchart of a cross-lingual fine-tuning method provided by an embodiment of the present invention;

[0026] Figure 2 is an example of code-switching involved in an embodiment of the present invention;

[0027] Figure 3 is an example of cross-lingual fine-tuning of a pre-trained model involved in an embodiment of the present invention;

[0028] Figure 4 is a flowchart of another cross-lingual fine-tuning method provided by an embodiment of the present invention;

[0029] Figure 5 is an example of a word vector alignment process involved in an embodiment of the present invention;

[0030] Figure 6 is a schematic structural diagram of a cross-lingual fine-tuning device provided by an embodiment of the present invention;

[0031] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0033] Figure 1 FIG. is a flowchart of a cross-lingual fine-tuning method provided by an embodiment of the present invention. This embodiment is applicable to the situation of cross-lingual fine-tuning of a pre-trained model, especially applicable to the fine-tuning scenario when the downstream task is a cross-lingual task such as a translation task. This method can be executed by a cross-lingual fine-tuning device, and this device can be implemented in a software and / or hardware manner and integrated into an electronic device. As Figure 1 shown, the method specifically includes the following steps:

[0034] S110. Obtain translation sample pair data, where the translation sample pair data includes a first source text corresponding to the source language and a first target text corresponding to the target language.

[0035] Among them, the translation sample pair data may refer to the parallel corpus used in the downstream translation task, that is, a data set S = { {X}, {Y}} composed of text pairs expressing the same meaning in two languages. The source language may refer to the language to be translated. The target language refers to the language after translation. The first source text may refer to a sentence expressed in the source language, that is, the sentence to be translated. The first target text may refer to a sentence expressed in the target language with the same meaning as the first source text, that is, the translated sentence.

[0036] S120. Replace some source words in the first source text with target near-synonyms to determine the second source text after replacement, where the target near-synonym refers to a target word with a similar semantic meaning to the source word.

[0037] Among them, the source word may refer to a vocabulary expressed in the source language. The target word may refer to a vocabulary expressed in the target language. Some source words may refer to at least one source word in the first source text, but not all source words. The second source text may be a text after code conversion of the first source text and containing at least one source word and at least one target word.

[0038] Specifically, based on an unsupervised word vector alignment method, target words with similar semantic meanings to each source word can be determined from the target language thesaurus, so as to obtain the target near-synonyms corresponding to each source word. For each first source text X i in the translation sample pair data, based on the Poisson distribution, some fragments of the first source text X i can be randomly selected, and each source word in this part of the fragment is replaced with the corresponding target near-synonym to obtain the second source text X' i after code conversion.Figure 2 An example of code-switching is given. For example, Figure 2 as shown, a source word z2 in the first source text is replaced with a corresponding target near-synonym e2, so as to obtain the second source text after code-switching, and the first source text can be used as the supervision data for the second source text.

[0039] S130. Replace some target words in the first target text with source near-synonyms to determine the second target text after replacement, where the source near-synonym refers to a source word with a semantic meaning similar to that of the target word.

[0040] Among them, some target words may refer to at least one target word in the first target text, but not all target words. The second target text may be a text after code-switching of the first target text, which contains at least one source word and at least one target word.

[0041] Specifically, based on the word vector alignment method, the source words with semantic meanings similar to each target word can be determined from the source language word library, and the source near-synonyms corresponding to each target word can be obtained. For each first target text Y in the translation sample pair data i , based on the Poisson distribution, some fragments in the first target text Y i can be randomly selected, and each target word in these fragments is replaced with the corresponding source near-synonym to obtain the second target text Y' after code-switching i ,. For example Figure 2 as shown, two target words e1 and e2 in the first target text are respectively replaced with the corresponding target near-synonyms z1 and z2, so as to obtain the second target text after code-switching, and the first target text can be used as the supervision data for the second target text.

[0042] It should be noted that the execution order of step S130 is not limited here. For example, step S130 can be executed sequentially after step S120, can also be executed before step S120, or can be executed simultaneously with S120.

[0043] S140. Based on the first source text, the second source text, the first target text, and the second target text, perform cross-lingual fine-tuning on the pre-trained model to determine the pre-fine-tuned model.

[0044] Among them, the pre-trained model may refer to an upstream general model trained on a large-scale monolingual data, such as an upstream general model trained on source language data and target language data. The pre-fine-tuned model may refer to the model obtained after the first fine-tuning of the pre-trained model and introducing cross-lingual alignment information, so as to reduce the cross-lingual difference between the upstream self-supervised task and the downstream translation task.

[0045] Specifically, the second source text and the second target text that contain both the source word and the target word can be used as training data, and the corresponding first source text and first target text can be used as label data to construct a self-supervised code-switching restoration task for the first training of the pre-trained model, that is, the first cross-lingual fine-tuning, so that cross-lingual alignment information can be introduced into the pre-fine-tuning model, narrowing the cross-lingual differences between the pre-training stage and the translation fine-tuning stage, achieving a soft landing of the pre-trained model in the downstream translation task, and at the same time, the required resource consumption is much less than the way of continuing pre-training.

[0046] It should be noted that by using the second source text and the second target text that contain both the source word and the target word as training data, and the corresponding first source text and first target text as label data to train the pre-trained model to predict the original text, the pre-trained model can utilize text segments in the same language to train the reconstruction ability of monolingual data in the first fine-tuning stage, and can also utilize text segments in different languages to train the cross-lingual translation ability, thus playing a bridging role between the pre-training stage and the translation fine-tuning stage.

[0047] Exemplarily, S140 may include: inputting the second source text into the pre-trained model and obtaining a first output text based on the output of the pre-trained model; inputting the second target text into the pre-trained model and obtaining a second output text based on the output of the pre-trained model; determining a training error according to the first source text, the first output text, the first target text, and the second output text based on a preset training function, and backpropagating the training error into the pre-trained model to adjust the network parameters in the pre-trained model; when a preset convergence condition is met, determining that the fine-tuning of the pre-trained model ends and obtaining a pre-fine-tuning model.

[0048] Specifically, the second source text X′ i and the second target text Y′ i can be respectively input into the pre-trained model for language restoration to obtain the first output text X * i and the second output text Y * i restored by the pre-trained model, and based on the training function, according to the cross-entropy between the first output text X * i and the first source text X i and the cross-entropy between the second output text Y * i and the first target text Y iThe cross-entropy between them determines the training error, and the training error is backpropagated into the pre-trained model to adjust the network parameters in the pre-trained model until the preset convergence condition is met. For example, when the number of iterations reaches the preset number or the training error converges, it is determined that the first training of the pre-trained model ends, and a pre-fine-tuned model is obtained. Among them, the training function L can be expressed as follows:

[0049]

[0050] It should be noted that the number of cross-lingual training steps in this embodiment is 1 / 7 of the total number of steps, so the resource consumption is much smaller than the way of continuing pre-training, greatly improving the resource utilization rate.

[0051] The technical solution of this embodiment replaces some source words in the first source text corresponding to the source language with target synonyms to obtain a second source text containing both source words and target words, and replaces some target words in the first target text corresponding to the target language with source synonyms to obtain a second target text containing both source words and target words, thus realizing the code-switching of synonyms. And the first source text, the second source text, the first target text and the second target text are used as training data to perform cross-lingual fine-tuning on the pre-trained model. Therefore, cross-lingual alignment information can be introduced into the obtained pre-fine-tuned model, narrowing the cross-lingual differences in the upstream self-supervised task and the downstream translation task, realizing the soft landing of the pre-trained model in the downstream translation task, and further improving the fine-tuning effect of the model in the downstream translation task.

[0052] Based on the above technical solution, after S140, it may further include: performing cross-lingual fine-tuning on the pre-fine-tuned model based on the translation sample pair data to determine the target translation model.

[0053] Among them, the target translation model may refer to a machine translation model trained based on the downstream translation task. Specifically, Figure 3 A cross-lingual fine-tuning example of a pre-trained model is given, as Figure 3 shown. By using the sample pairs composed of the first source text and the second source text, and the sample pairs composed of the first target text and the second target text, the pre-trained model is first trained to determine the pre-fine-tuned model after training, and then the pre-fine-tuned model is retrained based on the translation sample pair data, that is, cross-lingual fine-tuning again, so that the final target translation model can be obtained. By using the pre-fine-tuned model, the cross-lingual differences in the upstream self-supervised task and the downstream translation task can be narrowed, realizing the soft landing of the upstream general model in the downstream translation task, and further improving the performance of the model in the cross-lingual translation task and the training effect and model performance of the target translation model.

[0054] Figure 4The figure is a flowchart of another cross - language fine - tuning method provided by an embodiment of the present invention. Based on the above - mentioned embodiments, this embodiment details the specific process of replacing source words in the first source text with target near - synonyms and replacing target words in the first target text with source near - synonyms. The explanations of the same or corresponding terms in the above - mentioned embodiments will not be repeated here.

[0055] Refer to Figure 4 , another cross - language fine - tuning method provided by this embodiment specifically includes the following steps:

[0056] S410. Obtain translation sample pair data, where the translation sample pair data includes a first source text corresponding to the source language and a first target text corresponding to the target language.

[0057] S420. Extract some source words from the first source text to obtain each extracted first source word.

[0058] Specifically, based on the Poisson distribution, randomly extract some fragments from the first source text X i and use each source word in this part of the fragment as the first source word. Alternatively, the first source text can be segmented to obtain each source word in the first source text, and some source words in each source word can be extracted in order or randomly to obtain each extracted first source word.

[0059] S430. Based on the word - vector alignment method, determine at least one target near - synonym corresponding to each first source word.

[0060] Among them, the word - vector alignment method can refer to the method of semantically aligning source words and target words based on word vectors to obtain each target near - synonym after alignment of each first source word.

[0061] Exemplarily, S430 may include: for each first source word, obtain the first aligned word - vector corresponding to the first source word after word - vector alignment and the second aligned word - vector corresponding to each target word; based on the first aligned vector corresponding to the first source word and the second aligned word - vectors corresponding to each target word, determine the distance between the first source word and each target word; based on each distance, determine at least one target near - synonym corresponding to the first source word from each target word.

[0062] Among them, the distance between the first source word and the target word can be used to characterize the degree of correlation between these two words in the aligned word vector space. The closer the distance, the higher the degree of correlation, that is, the more similar the semantics. Specifically, in this embodiment, based on the word vector alignment method, the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word can be determined in real time, or based on the word vector alignment method, the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word can be determined in advance, so that during code conversion, based on the pre-determined aligned word vectors, the first aligned word vector corresponding to each first source word and the second aligned word vector corresponding to each target word can be directly obtained, further improving the cross-language fine-tuning efficiency. For each first source word, based on the cosine distance calculation method, according to the first aligned vector corresponding to the first source word and the second aligned word vectors corresponding to each target word, the cosine distance between the first source word and each target word can be determined, and the cosine distances are compared. The k target words with the closest cosine distance can be determined as the k target near-synonyms {y′ i corresponding to the first source word x′ i , y′ i+1 , …, y′ i+k-1}, that is, the candidate target near-synonyms of the first source word.

[0063] Exemplarily, for the case of pre-determining the first aligned word vector corresponding to each first source word and the second aligned word vector corresponding to each target word, "obtaining the first aligned word vector corresponding to the first source word and the second aligned word vector corresponding to each target word after word vector alignment" may include: based on a preset language processing model, determining the source word vector corresponding to each source word in the source language word library and the target word vector corresponding to each target word in the target source language word library; performing alignment processing on the source word vectors and the target word vectors to determine the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word; based on the first aligned word vector corresponding to each source word, obtaining the first aligned word vector corresponding to the first source word.

[0064] Among them, the preset language processing model can be a language model that encodes words to generate word vectors. For example, the preset language processing model can be, but is not limited to, the word2vec model. The preset language processing model in this embodiment is obtained by pre-training data based on translation samples.

[0065] Specifically, each source word in the source language word library can be input into the trained preset language processing model, and based on the output of the preset language processing model, the source word vector containing semantic relationships corresponding to each source word in the source language word vector space can be obtained, that is, (x1,.., x n)。Similarly, each target word in the target language vocabulary is input into the trained preset language processing model, and based on the output of the preset language processing model, the target word vector corresponding to each target word in the target language vector space and including semantic relationships can be obtained, that is, (y1,..., y n )。Semantic alignment processing can be performed on each source word vector and each target word vector based on the preset word vector alignment method to obtain the first aligned word vector (x′1,.., x′ n ) corresponding to each source word after semantic alignment and the second aligned word vector (y′1,..., y′ n ) corresponding to each target word, so that the first aligned word vector corresponding to each first source word and the second aligned word vector corresponding to each target word can be directly obtained, further improving the fine-tuning efficiency.

[0066] Exemplarily, performing alignment processing on each source word vector and each target word vector to determine the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word may include: mapping each source word vector and each target word vector into an aligned word vector space based on the preset unsupervised alignment method, and determining the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word after mapping.

[0067] Specifically, Figure 5 An example of a word vector alignment process is given. As Figure 5 shown, each source word vector (x1,.., x n ) in the source language vector space and each target word vector (y1,..., y n ) in the target language vector space can be mapped into the same feature space, that is, the aligned word vector space, so that in the aligned word vector space, the first aligned word vector (x′1,.., x′ n ) corresponding to each source word and the second aligned word vector (y′1,..., y′ n ) corresponding to each target word can be obtained.

[0068] S440. Replace each first source word in the first source text with a target near-synonym based on at least one target near-synonym, and determine the second source text after replacement.

[0069] Specifically, if the first source word corresponds to one target synonym, i.e., k = 1, each first source word in the first source text can be directly replaced with the corresponding target synonym to obtain the second source text after replacement. If the first source word corresponds to at least two target synonyms, i.e., k ≥ 2, one target synonym can be randomly selected from each of the target synonyms, and the first source word in the first source text can be replaced with the selected target synonym. Thus, by randomly selecting one target synonym from multiple target synonyms for final replacement, the alignment diversity, i.e., translation diversity, can be ensured, further improving the translation effect of the target translation model.

[0070] S450. Extract some target words from the first target text to obtain each extracted first target word.

[0071] Specifically, based on the Poisson distribution, some fragments of the first target text Y i can be randomly extracted, and each target word in this part of the fragment can be used as the first target word. Alternatively, the first target text can be segmented to obtain each target word in the first target text, and some target words in each target word can be extracted in order or randomly to obtain each extracted first target word.

[0072] S460. Based on the word vector alignment method, determine at least one source synonym corresponding to each first target word.

[0073] Among them, the word vector alignment method may refer to a method of semantically aligning source words and target words based on word vectors to obtain each source synonym after alignment of each first target word.

[0074] Exemplarily, S460 may include: for each first target word, obtain the first alignment word vector corresponding to each source word after word vector alignment and the second alignment word vector corresponding to each first target word; based on the first alignment vector corresponding to each source word and the second alignment word vector corresponding to this first target word, determine the distance between each source word and this first target word; based on each distance, determine at least one source synonym corresponding to this first target word from each source word.

[0075] Among them, the distance between the source word and the first target word can be used to characterize the degree of correlation between these two words in the aligned word vector space. The closer the distance, the higher the degree of correlation, that is, the more similar the semantics. Specifically, in this embodiment, based on the word vector alignment method, the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word can be determined in real time, or based on the word vector alignment method, the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word can be determined in advance, so that during code-switching, based on the pre-determined aligned word vectors, the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each first target word can be directly obtained, further improving the cross-lingual fine-tuning efficiency. Among them, the determination method of each aligned word vector can refer to the above relevant content and will not be elaborated here.

[0076] For each first target word, based on the cosine distance calculation method, according to the first aligned vector corresponding to each source word and the second aligned word vector corresponding to this first target word, the cosine distance between each source word and this first target word can be determined, and the cosine distances are compared. The k source words with the closest cosine distance can be determined as the k source near-synonyms {x' i , x' i , …, x' i+1 , …, x' i+k-1} of this first target word y', that is, the candidate source near-synonyms of the first target word.

[0077] S470. Based on at least one source near-synonym, perform source near-synonym replacement on each first target word in the first target text to determine the replaced second target text.

[0078] Specifically, if the first target word corresponds to one source near-synonym, that is, k = 1, then each first target word in the first target text can be directly replaced with the corresponding source near-synonym to obtain the replaced second target text. If the first target word corresponds to at least two source near-synonyms, that is, k ≥ 2, then one source near-synonym can be randomly selected from each source near-synonym, and this first target word in the first target text can be replaced with the selected source near-synonym. Thus, by randomly selecting one source near-synonym from multiple source near-synonyms for final replacement, the alignment diversity, that is, the translation diversity, can be ensured, and the translation effect of the target translation model is further improved.

[0079] S480. Based on the first source text, the second source text, the first target text, and the second target text, perform cross-lingual fine-tuning on the pre-trained model to determine the pre-fine-tuned model.

[0080] The technical solution of this embodiment determines at least one target near-synonym corresponding to each first source word extracted from the first source text based on the word vector alignment method, and performs target near-synonym replacement on each first source word in the first source text based on the at least one target near-synonym to determine the second source text after replacement, thereby realizing the code-switching of the first source text. Based on the word vector alignment method, at least one source near-synonym corresponding to each first target word extracted from the first target text is determined, and source near-synonym replacement is performed on each first target word in the first target text based on the at least one source near-synonym to determine the second target text after replacement, thereby realizing the code-switching of the first target text, and further enabling the more accurate construction of sample pairs required for the code-switching restoration task.

[0081] The following is an embodiment of a cross-lingual fine-tuning device provided by an embodiment of the present invention. This device and the cross-lingual fine-tuning methods of the above embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the cross-lingual fine-tuning device, reference may be made to the embodiments of the above cross-lingual fine-tuning methods.

[0082] Figure 6 FIG. is a schematic structural diagram of a cross-lingual fine-tuning device provided by an embodiment of the present invention. This embodiment is applicable to the situation of cross-lingual fine-tuning of a pre-trained model, especially applicable to the fine-tuning scenario when the downstream task is a cross-lingual task such as a translation task. As Figure 6 shown, the device specifically includes: a translation sample pair data acquisition module 610, a source word replacement module 620, a target word replacement module 630, and a pre-fine-tuning model determination module 640.

[0083] Among them, the translation sample pair data acquisition module 610 is used to acquire translation sample pair data, and the translation sample pair data includes a first source text corresponding to the source language and a first target text corresponding to the target language; the source word replacement module 620 is used to perform target near-synonym replacement on some source words in the first source text to determine the second source text after replacement, where the target near-synonym refers to a target word with a similar semantic meaning to the source word; the target word replacement module 630 is used to perform source near-synonym replacement on some target words in the first target text to determine the second target text after replacement, where the source near-synonym refers to a source word with a similar semantic meaning to the target word; the pre-fine-tuning model determination module 640 is used to perform cross-lingual fine-tuning on the pre-trained model based on the first source text, the second source text, the first target text, and the second target text to determine the pre-fine-tuning model.

[0084] In the technical solution of this embodiment, by performing target near-synonym replacement on some source words in the first source text corresponding to the source language, a second source text containing both source words and target words is obtained, and by performing source near-synonym replacement on some target words in the first target text corresponding to the target language, a second target text containing both source words and target words is obtained, thereby realizing the code conversion of near-synonyms. The first source text, the second source text, the first target text, and the second target text are used as training data to perform cross-lingual fine-tuning on the pre-trained model. Thus, cross-lingual alignment information can be introduced into the obtained pre-fine-tuned model, reducing the cross-lingual differences in the upstream self-supervised task and the downstream translation task, realizing the soft landing of the pre-trained model in the downstream translation task, and further improving the fine-tuning effect of the model in the downstream translation task.

[0085] Optionally, the source word replacement module 620 includes:

[0086] A source word extraction unit for extracting some source words from the first source text to obtain each extracted first source word;

[0087] A target near-synonym determination unit for determining at least one target near-synonym corresponding to each first source word based on the word vector alignment method;

[0088] A second source text determination unit for performing target near-synonym replacement on each first source word in the first source text based on at least one target near-synonym to determine the replaced second source text.

[0089] Optionally, the target near-synonym determination unit includes:

[0090] An aligned word vector acquisition sub-unit for, for each first source word, acquiring the first aligned word vector corresponding to the first source word after word vector alignment and the second aligned word vector corresponding to each target word;

[0091] A distance determination sub-unit for determining the distance between the first source word and each target word based on the first aligned vector corresponding to the first source word and the second aligned word vector corresponding to each target word;

[0092] A target near-synonym determination sub-unit for determining at least one target near-synonym corresponding to the first source word from each target word based on each distance.

[0093] Optionally, the aligned word vector acquisition sub-unit is specifically used for:

[0094] Based on a preset language processing model, determine the source word vectors corresponding to each source word in the source language thesaurus, and the target word vectors corresponding to each target word in the target source language thesaurus; perform alignment processing on each source word vector and each target word vector to determine the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word; based on the first aligned word vector corresponding to each source word, obtain the first aligned word vector corresponding to the first source word.

[0095] Optionally, the aligned word vector obtaining subunit is further specifically configured to:

[0096] Based on a preset unsupervised alignment method, map each source word vector and each target word vector into an aligned word vector space, and determine the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word after mapping.

[0097] Optionally, the second source text determining unit is specifically configured to:

[0098] If the first source word corresponds to at least two target near-synonyms, extract one target near-synonym from the at least two target near-synonyms, and replace the first source word in the first source text with the extracted target near-synonym.

[0099] Optionally, the target word replacement module 630 includes:

[0100] A target word extraction unit for extracting some target words in the first target text to obtain each extracted first target word;

[0101] A source near-synonym determining unit for determining at least one source near-synonym corresponding to each first target word based on a word vector alignment method;

[0102] A second target text determining unit for performing source near-synonym replacement on each first target word in the first target text based on the at least one source near-synonym to determine the replaced second target text.

[0103] Optionally, the pre-fine-tuning model determining module 640 is specifically configured to:

[0104] Input the second source text into the pre-trained model, and obtain a first output text based on the output of the pre-trained model; input the second target text into the pre-trained model, and obtain a second output text based on the output of the pre-trained model; based on a preset training function, determine a training error according to the first source text, the first output text, the first target text, and the second output text, and backpropagate the training error to the pre-trained model to adjust the network parameters in the pre-trained model; when a preset convergence condition is satisfied, determine that the fine-tuning of the pre-trained model ends, and obtain a pre-fine-tuning model.

[0105] Optionally, the device further includes:

[0106] A target translation model determination module, configured to, after determining a pre-fine-tuning model, perform cross-lingual fine-tuning on the pre-fine-tuning model based on translation sample pairs of data to determine a target translation model.

[0107] The cross-lingual fine-tuning device provided by an embodiment of the present invention can execute the cross-lingual fine-tuning method provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to executing the cross-lingual fine-tuning method.

[0108] It should be noted that, in the embodiments of the above cross-lingual fine-tuning device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and are not used to limit the protection scope of the present invention.

[0109] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Figure 7 It shows a block diagram of an exemplary electronic device 12 suitable for implementing the embodiments of the present invention. Figure 7 The shown electronic device 12 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0110] As Figure 7 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0111] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0112] The electronic device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0113] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 7 not shown, typically referred to as a "hard disk drive"). Although Figure 7 not shown in, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 through one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0114] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.

[0115] The electronic device 12 may also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 22. Further, the electronic device 12 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through the bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0116] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, for example, implementing the steps of a cross-lingual fine-tuning method provided by the present embodiment. The method includes:

[0117] Obtain translation sample pair data, where the translation sample pair data includes a first source text corresponding to the source language and a first target text corresponding to the target language;

[0118] Perform target near-synonym replacement on some source words in the first source text to determine the second source text after replacement, where the target near-synonym refers to a target word with a similar semantic meaning to the source word;

[0119] Perform source near-synonym replacement on some target words in the first target text to determine the second target text after replacement, where the source near-synonym refers to a source word with a similar semantic meaning to the target word;

[0120] Based on the first source text, the second source text, the first target text, and the second target text, perform cross-lingual fine-tuning on the pre-trained model to determine the pre-fine-tuned model.

[0121] Of course, those skilled in the art can understand that the processor can also implement the technical solutions of the cross-lingual fine-tuning method provided by any embodiment of the present invention.

[0122] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the cross-lingual fine-tuning method provided by any embodiment of the present invention. The method includes:

[0123] Obtain translation sample pair data, where the translation sample pair data includes a first source text corresponding to the source language and a first target text corresponding to the target language;

[0124] Perform target near-synonym replacement on some source words in the first source text to determine the second source text after replacement, where the target near-synonym refers to a target word with a similar semantic meaning to the source word;

[0125] Perform source near-synonym replacement on some target words in the first target text to determine the second target text after replacement, where the source near-synonym refers to a source word with a similar semantic meaning to the target word;

[0126] Based on the first source text, the second source text, the first target text, and the second target text, perform cross-lingual fine-tuning on the pre-trained model to determine the pre-fine-tuned model.

[0127] The computer storage medium of an embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0128] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0129] The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0130] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0131] Those of ordinary skill in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented with program codes executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0132] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A cross-lingual fine-tuning method, characterized in that, Including: Obtain translation sample pair data, where the translation sample pair data includes a first source text corresponding to the source language and a first target text corresponding to the target language; Perform target near-synonym replacement on some source words in the first source text to determine the second source text after replacement, where the target near-synonym refers to a word that is semantically similar to the source word and corresponds to the target word in the target language; Perform source near-synonym replacement on some target words in the first target text to determine the second target text after replacement, where the source near-synonym refers to a word that is semantically similar to the target word and corresponds to the source word in the source language; Based on the first source text, the second source text, the first target text, and the second target text, perform cross-lingual fine-tuning on the pre-trained model to determine the pre-fine-tuned model; Among them, the pre-trained model is trained on monolingual data, and the languages corresponding to the monolingual data include the source language and the target language. The performing cross-lingual fine-tuning on the pre-trained model based on the first source text, the second source text, the first target text, and the second target text includes: Use the second source text and the second target text as training data, and use the first source text and the first target text as label data to construct a self-supervised code-switching restoration task to perform cross-lingual fine-tuning on the pre-trained model.

2. The method according to claim 1, wherein The performing target near-synonym replacement on some source words in the first source text to determine the second source text after replacement includes: Extract some source words from the first source text to obtain each extracted first source word; Based on the word vector alignment method, determine at least one target near-synonym corresponding to each first source word; Based on the at least one target near-synonym, perform target near-synonym replacement on each first source word in the first source text to determine the second source text after replacement.

3. The method according to claim 2, wherein The determining at least one target near-synonym corresponding to each first source word based on the word vector alignment method includes: For each first source word, obtain the first aligned word vector corresponding to the first source word after word vector alignment and the second aligned word vector corresponding to each target word; Based on the first aligned vector corresponding to the first source word and the second aligned word vectors corresponding to each target word, determine the distance between the first source word and each target word; Based on each of the distances, determine at least one target near-synonym corresponding to the first source word from each target word.

4. The method according to claim 3, wherein The obtaining the first aligned word vector corresponding to the first source word after word vector alignment and the second aligned word vector corresponding to each target word includes: Based on a preset language processing model, determine the source word vector corresponding to each source word in the source language word library and the target word vector corresponding to each target word in the target source language word library; Perform alignment processing on each of the source word vectors and each of the target word vectors to determine the first aligned word vector corresponding to each source word and the second aligned word vector corresponding to each target word; Based on the first aligned word vector corresponding to each source word, obtain the first aligned word vector corresponding to the first source word.

5. The method according to claim 4, wherein Performing alignment processing on each of the source word vectors and each of the target word vectors to determine a first aligned word vector corresponding to each source word and a second aligned word vector corresponding to each target word, includes: Based on a preset unsupervised alignment method, mapping each of the source word vectors and each of the target word vectors into an aligned word vector space, and determining a first aligned word vector corresponding to each source word and a second aligned word vector corresponding to each target word after mapping.

6. The method according to claim 2, wherein Based on the at least one target near-synonym, performing target near-synonym replacement on each of the first source words in the first source text to determine a second source text after replacement, includes: If the first source word corresponds to at least two target near-synonyms, extracting one target near-synonym from the at least two target near-synonyms, and replacing the first source word in the first source text with the extracted target near-synonym.

7. The method according to claim 1, wherein Performing source near-synonym replacement on some of the target words in the first target text to determine a second target text after replacement, includes: Extracting some of the target words in the first target text to obtain each extracted first target word; Based on a word vector alignment method, determining at least one source near-synonym corresponding to each of the first target words; Based on the at least one source near-synonym, performing source near-synonym replacement on each of the first target words in the first target text to determine a second target text after replacement.

8. The method according to claim 1, wherein Based on the first source text, the second source text, the first target text, and the second target text, performing cross-lingual fine-tuning on a pre-trained model to determine a pre-fine-tuned model, further includes: Inputting the second source text into the pre-trained model, and obtaining a first output text based on the output of the pre-trained model; Inputting the second target text into the pre-trained model, and obtaining a second output text based on the output of the pre-trained model; Based on a preset training function, determining a training error according to the first source text, the first output text, the first target text, and the second output text, and backpropagating the training error into the pre-trained model to adjust network parameters in the pre-trained model; When a preset convergence condition is satisfied, determining that the cross-lingual fine-tuning of the pre-trained model ends, and obtaining a pre-fine-tuned model.

9. The method according to any one of claims 1-8, characterized in that, After determining the pre-fine-tuned model, further includes: Based on the translation sample pair data, performing cross-lingual fine-tuning on the pre-fine-tuned model to determine a target translation model.

10. A cross-lingual fine-tuning device, characterized in that, Includes: A translation sample pair data acquisition module, configured to acquire translation sample pair data, where the translation sample pair data includes a first source text corresponding to a source language and a first target text corresponding to a target language; A source word replacement module, configured to perform target near-synonym replacement on some of the source words in the first source text to determine a second source text after replacement, where the target near-synonym refers to a target word that is semantically similar to the source word and corresponds to the target language; A target word replacement module, configured to perform source near-synonym replacement on some of the target words in the first target text to determine a second target text after replacement, where the source near-synonym refers to a source word that is semantically similar to the target word and corresponds to the source language; A pre-fine-tuning model determination module, configured to use the second source text and the second target text as training data, and use the first source text and the first target text as label data to construct a self-supervised code-switching restoration task for cross-lingual fine-tuning of a pre-trained model to determine a pre-fine-tuning model, wherein the pre-trained model is trained on monolingual data, and the languages corresponding to the monolingual data include the source language and the target language.

11. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the cross-lingual fine-tuning method according to any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the cross-lingual fine-tuning method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Machine translation method, training method, corresponding device and electronic equipment

    CN110956045A

  • Translation model training method, device and equipment and storage medium

    CN112560510A