Cross-language migration method based on hybrid expert and code conversion data

Through data distillation, the code conversion data set is synthesized and a hybrid expert model is constructed, which solves the problem of word-level synthesis difficulties and English forgetting in cross-language migration, and achieves efficient and flexible cross-language migration effects.

CN120296420APending Publication Date: 2025-07-11NANJING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510380429.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has difficulties in low-resource language processing, especially in cross-language migration, which lacks effective word-level code conversion synthesis methods, and has the problem of English forgetting, resulting in limited cross-language migration capabilities of the model.

Method used

The code conversion data set is synthesized using a data distillation method and a hybrid expert model is built. Through the hybrid expert module and language router, the hybrid expert model is trained to achieve cross-language migration and avoid English forgetting.

Benefits of technology

It realizes low-cost and efficient code conversion data synthesis, improves cross-language migration capabilities, promotes multi-language application of cross-language migration models, and solves the problem of English forgetting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296420A_ABST
    Figure CN120296420A_ABST
Patent Text Reader

Abstract

The invention provides a cross-language migration method based on mixed experts and code conversion data, which comprises the following steps of: 1, synthesizing the code conversion data based on a data distillation method to obtain a code conversion data set; step 2, constructing a hybrid expert model; 3, training the hybrid expert model constructed in the step 2 by using the code conversion data set obtained in the step 1 to obtain a trained hybrid expert model; and step 4, realizing cross-language migration by using the trained hybrid expert model. According to the invention, the hybrid expert structure can ensure that the English ability is unchanged in the code conversion data training process, so that the cross-language enhancement effect of the code conversion data can be further stimulated; the method can be applied to all open source large models in an unlimited manner, and the English cross-language migration capability can be transferred to any language, so that the multi-language capability of the model is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross - language transfer method, in particular to a cross - language transfer method based on a mixture of experts and code - switched data. Background Art

[0002] The information provided in this section is only background information related to the present disclosure, and it is not necessarily prior art.

[0003] With the advent of the deep - learning era, large language models (LLMs), such as GPT - 4 (reference: Achiam J, Adler S, Agarwal S, et al. Gpt - 4 technical report[J]. arXiv preprint arXiv:2303.08774, 2023.) and Qwen (reference: Yang A, Yang B, Zhang B, et al. Qwen2.5 technical report[J]. arXiv preprint arXiv:2412.15115, 2024.), rely on large - scale multilingual datasets for pre - training to master the ability to understand, generate, and translate various human languages. Although these models perform well in multiple languages, they still face difficulties in processing some low - resource languages. Due to the imbalance in the amount of pre - training data for different languages, academia and industry have explored various methods to improve the cross - language transfer ability of models, by transferring capabilities from high - resource languages such as Chinese or English to low - resource languages, thereby enhancing the capabilities in low - resource languages. Among these methods, training with a mixture of code - switched data is a method with remarkable effects. Code - switching refers to the phenomenon where multiple languages appear in the same context. This data form places related concepts and semantics in different languages in the same context, enabling large models based on context learning to obtain powerful cross - language transfer capabilities. Code - switching can be divided into sentence - level and word - level according to the granularity of different language texts. Given a monolingual document, its sentence - level code - switching can directly translate some sentences through a translation model. However, for word - level code - switching, there is currently a lack of effective synthesis means, which limits the large - scale application of code - switched data in cross - language transfer learning.

[0004] The existing technical solution of code-switching curriculum learning (CSCL) (Reference: Yoo H, Park C, Yun S, et al. Code-Switching Curriculum Learning for Multilingual Transfer in LLMs[J]. arXiv preprint arXiv:2411.02460, 2024.) uses GPT-4o to generate code-switching data. When such data needs to be applied on a large scale, using GPT-4o to generate it will incur huge costs. At the same time, CSCL requires parallel sentence pairs (mutually translated sentences) as input. Usually, the sentences in the pre-training data only contain one language. An intuitive approach is to first translate these sentences and then synthesize code-switching data from the obtained parallel sentence pairs according to the method of CSCL. However, this approach not only incurs translation costs but also has the problem of error accumulation. Therefore, how to synthesize code-switching in any language from monolingual sentences at low cost and quickly is an important issue.

[0005] The problem of catastrophic forgetting in English: Although CSCL uses replay techniques to alleviate the problem of catastrophic forgetting in English, its model still has a large degree of forgetting in English. The continuous forgetting of English capabilities during the training process will further limit the cross-lingual transfer effect of code-switching data, resulting in its inability to achieve the maximum effect. Therefore, there is currently a lack of a framework without forgetting to further stimulate the effectiveness of code-switching data.

[0006] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0007] Object of the Invention: The technical problem to be solved by the present invention is to provide a cross-lingual transfer method based on mixture of experts and code-switching data in view of the deficiencies of the prior art.

[0008] To solve the above technical problem, the present invention discloses a cross-lingual transfer method based on mixture of experts and code-switching data, including the following steps:

[0009] Step 1, synthesize code-switching data based on the data distillation method to obtain a code-switching data set;

[0010] Step 2, construct a mixture of experts model;

[0011] Step 3, use the code-switching data set obtained in Step 1 to train the mixture of experts model constructed in Step 2 to obtain a trained mixture of experts model;

[0012] Step 4: Use the trained mixture-of-experts model to achieve cross-lingual transfer.

[0013] Furthermore, the code conversion data synthesis based on the data distillation method described in Step 1 includes:

[0014] Step 1-1: Collect bilingual translation parallel corpora, specifically including:

[0015] Let S l be the sentence in language l. Then, the bilingual translation parallel corpus includes sentence pairs that are translations of each other between language A and language B (S A , S B );

[0016] Step 1-2: Use a large model, i.e., the teacher model, to generate the annotation type code conversion CS A and the replacement type code conversion CS B for sentences S A and S B in the sentence pair (S Annotation , S Replacement ), respectively, to obtain the intermediate dataset D CS- . Specifically, it is represented as follows:

[0017]

[0018] where t represents the code conversion type, i.e., annotation and replacement, Prompt t represents the instruction template for the code conversion of type t, and D CS-SFT is the intermediate dataset, where all samples only have a single language original input text;

[0019] Step 1-3: Use the intermediate dataset D CS-SFT to perform supervised instruction fine-tuning on the small model, i.e., the student model, to obtain the fine-tuned student model;

[0020] Step 1-4: Use the fine-tuned student model to synthesize code conversion data for the collected pre-training data to obtain the code conversion dataset D Cs .

[0021] Furthermore, the construction of the mixture-of-experts model described in Step 2 includes:

[0022] Step 2-1: Select a pre-trained language model with a dense structure as the basic architecture of the mixture-of-experts model;

[0023] Step 2-2: Add a mixture-of-experts module to the above basic architecture.

[0024] Further, the pre-trained language model with a dense structure described in step 2-1 is a standard Transformer structure.

[0025] Further, in step 2-1, the pre-trained language model with a dense structure specifically includes:

[0026] An input word embedding module, an output word embedding module, and a multi-layer neural network; among them,

[0027] The input word embedding module is a representation matrix that converts the input word into a corresponding vector representation;

[0028] The output word embedding module converts the input vector representation into a corresponding word for output.

[0029] Further, each layer of the multi-layer neural network has the same structure, including: a self-attention module and a feed-forward neural network module.

[0030] Further, in the multi-layer neural network, the self-attention module adds information of other words in the context to the vector representation of each word through the self-attention mechanism; the feed-forward neural network module is a non-linear transformation matrix for learning.

[0031] Further, the addition of the mixture-of-experts module described in step 2-2 is as follows:

[0032] Step 2-2-1: Duplicate the feed-forward neural network module to obtain a first feed-forward neural network module and a second feed-forward neural network module, which are used as the source-language expert and the target-language expert respectively;

[0033] Step 2-2-2: Construct a language router before the source-language expert and the target-language expert;

[0034] Step 2-2-3: The source-language expert and the target-language expert respectively process the input representations assigned by the language router;

[0035] Step 2-2-4: Concatenate and aggregate the outputs of the source-language expert and the target-language expert to obtain the final output of the mixture-of-experts model.

[0036] Further, the language router described in step 2-2-2 routes according to the language corresponding to the input representation, routes the input representation of the source language to the source-language expert, and routes the input representation of the target language to the target-language expert.

[0037] Further, the training described in step 3 includes:

[0038] Step 3-1: Freeze all the original parameters of the mixture-of-experts model and do not update them;

[0039] Step 3-2: Input the code conversion dataset obtained in Step 1 and update the parameters of the target language expert.

[0040] Beneficial effects:

[0041] 1. The code conversion synthesis model in the present invention does not require bilingual parallel corpora of the input text. In the prior art, parallel corpora of the input text are required to synthesize code conversion data. For each pair of languages, the monolingual documents need to be translated first before use. Therefore, the present invention is more flexible and easier to use. The hybrid expert structure of the present invention can ensure the invariance of English ability during the training process of code conversion data, which can further stimulate the cross-language enhancement effect of code conversion data.

[0042] 2. The present invention trains a code conversion data synthesis model with a smaller parameter scale through a data distillation-based method. Its inference cost is reduced significantly, and at the same time, the synthesis speed is faster, which can effectively promote the large-scale application of code conversion synthetic data in the field of cross-language transfer. The cross-language transfer framework of the present invention based on the hybrid expert model and code conversion data can be applied to all open-source large models without limitation, and can transfer the English cross-language ability to any language, thereby effectively improving the multilingual ability of the model. Description of the drawings

[0043] The following further describes the present invention in detail with reference to the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0044] Figure 1 Schematic diagram for constructing a code conversion dataset in the embodiment.

[0045] Figure 2 Schematic diagram for constructing a hybrid expert model in the embodiment. Specific embodiments

[0046] The overall idea of the present invention is as follows: propose a more flexible and lower-cost synthesis method to promote the application and research of code conversion synthetic data in cross-language alignment tasks; at the same time, propose a model structure without forgetting to further promote the role of code conversion data in cross-language transfer.

[0047] The present invention mainly proposes a cross-language transfer framework based on hybrid experts and code conversion data. At the same time, since this framework requires large-scale synthesis of code conversion data, in order to reduce costs and improve speed, the present invention also includes a low-cost code conversion data synthesis method based on data distillation.

[0048] Table 1 Example table of Chinese-English code conversion data

[0049]

[0050] As shown in Table 1, existing work simply classifies code conversion according to granularity into sentence level and word level. In the present invention, the code conversion data of each granularity is further subdivided into annotation and replacement types:

[0051] Annotation: This type refers to using the translation in another language to annotate after the original text, usually appearing within parentheses;

[0052] Replacement: This type uses the translation to directly replace the original text.

[0053] The solution of the present invention first aims to synthesize the 4 types of Chinese code conversion data shown in Table 1 based on English pre-training data, while the synthesis method of existing work has problems of high cost and slow speed.

[0054] In order to reduce the synthesis cost of word-level code conversion data and improve the synthesis speed, the present invention proposes a synthesis solution based on data distillation.

[0055] Given languages A and B, where language A is the language that the model is more proficient in, this solution expects to transfer the ability from language A to language B:

[0056] Step 1: Synthesize code conversion data based on the data distillation method and construct a code conversion data set

[0057] This solution first collects a small batch of high-quality bilingual translation parallel corpora. To ensure data quality and diversity, this solution collects the validation set and test set data of multiple translation data sets. Furthermore, this solution uses a stronger large model to synthesize code conversion data according to the synthesis method of CSCL. Finally, a supervised fine-tuning data set for code conversion data synthesis is constructed based on this batch of data.

[0058] This solution defines the large model as having a parameter count exceeding 70 billion, and at the same time its performance is a current leading-edge model, which is called the teacher model hereinafter; the small model has a parameter count in the tens of billions scale, and at the same time has good generation ability and low inference cost, which is called the student model hereinafter.

[0059] Specifically, let S l be the sentence of language l. Given the sentence pair (S A , S B ) that are translations of each other between languages A and B, this solution first uses the teacher model to generate the annotation-type code conversion CS A and the replacement-type code conversion CS B for (S Annotation and S Replacement, in order to construct a supervised fine-tuning dataset for single-source input to enable code conversion data to be generated by only inputting text in a single language to the model, this solution constructs a dataset D in the following form CS-SFT :

[0060]

[0061] where l and t represent language and code conversion type, and Prompt t represents the instruction template for the supervised fine-tuning data of code conversion of type t. Type t includes the annotation Annotation and replacement Replacement described above. In D CS-SFT , all samples only have a single-language original input text.

[0062] Finally, the solution of the present invention uses D CS-SF to perform supervised instruction fine-tuning on the student model (Reference: Longpre S, Hou L, Vu T, et al. The flan collection: Designing data and methods for effective instruction tuning[C] / / International Conference on Machine Learning. PMLR, 2023: 22631-22648.), enabling it to have the ability to synthesize high-quality code conversion data. Since the model has a small number of parameters, it has the advantages of low cost and fast speed. Moreover, from the above method of constructing the fine-tuning data, it can be seen that the model can directly synthesize the pre-trained document data in a single language without parallel corpora.

[0063] Furthermore, this solution uses the model trained above to synthesize code conversion data for a large amount of pre-trained data, obtaining the code conversion dataset D CS .

[0064] Step 2: Construct a mixture-of-experts model:

[0065] Given a pre-trained language model with a dense structure (hereinafter referred to as the dense model), this solution aims to enhance the cross-lingual transfer ability of the dense model. The dense model is a standard Transformer structure (Reference: Grattafiori A, Dubey A, Jauhri A, et al. The llama 3herd of models[J]. arXiv e-prints, 2024: arXiv:2407.21783.), mainly composed of input word embeddings, output word embeddings, and multiple layers of neural networks. Among them, the input word embedding is a representation matrix that outputs the corresponding vector representation for the input word, and the output word embedding matrix receives the output representation of the final layer of the model and converts it into the corresponding word. For the intermediate layers of the model, each layer has the same structure, mainly including an attention module and a feed-forward neural network module. The self-attention module adds information of other words in the context to the vector representation of each word through the self-attention mechanism, and the feed-forward neural network module is a non-linear transformation matrix with a large number of parameters, enabling the model to have strong learning ability.

[0066] To solve the problem of English forgetting during the process of using code-switching data to enhance cross-lingual transfer, the present invention extends the dense model to a mixture-of-experts model. Specifically, the present invention replicates the feed-forward neural network module (i.e., the expert in the figure) in the pre-trained language model with a dense structure, and combines the new expert 1 and the original expert 0 module through a language router to form a mixture-of-experts module. The remaining modules of the dense model, including the word embedding module and the attention module of each layer, remain unchanged. In this mixture-of-experts module, expert 0 is the source language expert, and expert 1 will serve as the target language expert. The language router routes according to the language corresponding to the input representation, routes the representation of the source language to expert 0, routes the representation of the target language to expert 1 for separate processing, and finally concatenates and aggregates the outputs of the two.

[0067] First, define the language router Route lang , for the word T in a certain sample in the code-switching dataset D obtained in step 1 CS in:

[0068]

[0069] Given the input x ∈ R of the vector representation of a token of the above mixture-of-experts module h , where h is the dimension of the vector. Given the feed-forward neural network parameters FFN0 and FFN1 of the source language expert and the target language expert, the calculation result of the mixture-of-experts module is defined as y:

[0070] y = Router lang×FFN0(x)+(1 - Router lang )×FFN1(x)

[0071] Step 3: Train the mixture-of-experts model:

[0072] During the training process, all the original parameters of the mixture-of-experts model are frozen and not updated, so as to ensure zero forgetting of the source language ability. The parameters of the target language expert FFN1 are updated to enhance the target language by transferring the source language ability. When using code-switching data for training, the transfer learning of the source language and the target language mainly occurs in the attention module. The words of the target language can transfer the required knowledge and ability from the words of the source language through the attention mechanism. At the same time, the mixture-of-experts module and the frozen training strategy of the present invention ensure that the source language ability will not be forgotten, thus improving the upper limit of cross-lingual transfer learning.

[0073] Example:

[0074] According to the above scheme introduction, taking the target language as Chinese and expecting to transfer the ability from the source language English to the target language Chinese as an example, as Figure 1 shown, in the code-switching data synthesis stage, the present invention uses the TowerInstruct large model translator to synthesize sentence-level code-switching data. For word-level code-switching data, the present invention first collects the validation set and test set data of the Chinese-English parallel corpus in the Flores200 dataset. These parallel corpora are created by artificial experts and have high quality. Then, the API of the GPT-4o-mini model is used to collect the answers to the code-switching data synthesis questions composed of these translation corpora. Next, the fine-tuning data is constructed following the supervised fine-tuning data construction process of the present invention. Finally, the present invention selects Qwen2.5-3B-Instruct for fine-tuning. This model is the best model under the current parameter scale. The final model has the advantages of high-quality code-switching synthesis ability, low cost, and fast speed. Finally, the present invention performs English-to-Chinese code-switching data enhancement on the high-quality English pre-trained document data of FineWeb-Edu to obtain the final code-switching dataset.

[0075] Using this batch of data, this scheme enhances cross-lingual transfer using the Qwen2.5-7B base large model. First, Qwen2.5-7B is transformed according to the above architecture, as Figure 2 shown, by adding a target language expert module. Finally, the model is continuously pre-trained on the code-switching dataset, following the above parameter training settings, to achieve code-switching cross-lingual transfer without forgetting.

[0076] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the inventive content of a cross-language migration method based on a mixture of experts and code-converted data and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0077] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the essence of the technical solutions in the embodiments of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in a storage medium and includes several instructions for causing a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MCU, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present invention.

[0078] The present invention provides an idea and method for a cross-language migration method based on a mixture of experts and code-converted data. There are many methods and ways to specifically implement this technical solution. The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.

Claims

1. A cross - language migration method based on a mixture of experts and code - converted data, characterized in that It includes the following steps: Step 1, perform code conversion data synthesis based on the data distillation method to obtain a code conversion data set; Step 2, construct a mixture-of-experts model; Step 3, use the code conversion data set obtained in Step 1 to train the mixture-of-experts model constructed in Step 2 to obtain a trained mixture-of-experts model; Step 4, use the trained mixture-of-experts model to achieve cross-language transfer.

2. A cross - language transfer method based on a mixture of experts and code - converted data according to claim 1, characterized in that, The code conversion data synthesis based on the data distillation method described in Step 1 includes: Step 1-1, collect bilingual translation parallel corpora, specifically including; Let S l be a sentence of language l. Then the bilingual translation parallel corpus includes sentence pairs (S A , S B ) that are translations of each other between language A and language B; Step 1-2, use the large model, i.e., the teacher model, to train the sentence pairs (S A ,S B ) generate sentences S respectively A and sentence S B Annotation type code conversion CS Annotation and replace type code conversion CS Replacement , get the intermediate data set D CS-SF , specifically expressed as follows: Among them, t represents the code conversion type, namely annotation and replacement, Prompt t represents the instruction template for code conversion of type t, D CS- is an intermediate data set, in which all samples only have a single-language original input text; Step 1-3, use the intermediate dataset D CS-SF Perform supervised instruction fine-tuning on the small model, i.e., the student model, to obtain the fine-tuned student model; Steps 1-4, using the fine-tuned student model, synthesize code conversion data from the collected pre-training data to obtain the code conversion dataset D CS .

3. A cross-language migration method based on a mixture of experts and code-converted data according to claim 2, characterized in that The construction of the mixture-of-experts model described in Step 2 includes: Step 2-1, select a pre-trained language model with a dense structure as the basic architecture of the mixture-of-experts model; Step 2-2, add a mixture-of-experts module to the above basic architecture.

4. A cross - language migration method based on a mixture of experts and code - converted data according to claim 3, characterized in that, The pre-trained language model with a dense structure described in Step 2-1 is a standard Transformer structure.

5. A cross - language migration method based on a mixture of experts and code - converted data according to claim 4, characterized in that, In Step 2-1, the pre-trained language model with a dense structure specifically includes: An input word embedding module, an output word embedding module, and a multi-layer neural network; among them, The input word embedding module is a representation matrix that converts the input word into a corresponding vector representation; The output word embedding module converts the input vector representation into a corresponding word for output.

6. The cross - language migration method based on a mixture of experts and code - converted data according to claim 5, characterized in that, The multi-layer neural network has the same structure for each layer, including: a self-attention module and a feed-forward neural network module.

7. A cross-language migration method based on a mixture of experts and code-converted data according to claim 6, characterized in that In the multi-layer neural network, the self-attention module adds information of other words in the context to the vector representation of each word through the self-attention mechanism; the feed-forward neural network module is a non-linear transformation matrix for learning.

8. A cross-language transfer method based on a mixture of experts and code-switching data according to claim 7, characterized in that The addition of the mixture-of-experts module described in Step 2-2 is specifically as follows: Step 2-2-1, copy the feed-forward neural network module to obtain a first feed-forward neural network module and a second feed-forward neural network module, which are used as the source language expert and the target language expert respectively; Step 2-2-2, construct a language router before the source language expert and the target language expert; Step 2-2-3, the source language expert and the target language expert respectively process the input representation assigned by the language router; Step 2-2-4, splice and aggregate the outputs of the source language expert and the target language expert to obtain the final output of the mixture-of-experts model.

9. A cross - language transfer method based on a mixture of experts and code - converted data according to claim 8, characterized in that, The language router described in Step 2-2-2 routes according to the language corresponding to the input representation, routes the input representation of the source language to the source language expert, and routes the input representation of the target language to the target language expert.

10. A cross - language transfer method based on a mixture of experts and code - converted data according to claim 9, characterized in that, The training described in Step 3 includes: Step 3-1, freeze all the original parameters of the mixture-of-experts model without updating; Step 3-2, input the code conversion data set obtained in Step 1 and update the parameters of the target language expert.

Citation Information

Cited By

  • Sound duplicating method and related device

    CN122157639A