Knowledge migration method for multi-language machine translation

By evaluating language perception and extracting parameters of the teacher model, and selective updates are performed using the shared low-rank adaptation module and the language-specific low-rank adaptation module, the parameter interference and knowledge forgetting problems in multilingual machine translation are solved, and the translation performance is improved.

CN120354865AActive Publication Date: 2025-07-22TIANJIN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510828368.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-22
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The prior art has problems such as parameter interference, knowledge forgetting, insufficient knowledge transfer efficiency and insufficient language-aware information utilization in multilingual machine translation, which affects translation performance.

Method used

By evaluating the language perception of associated neurons in the teacher model, classifying them into language-general and language-specific neurons, the weight sub-matrix is extracted, and selective parameter updates are used to use the shared low-rank adaptation module and the language-specific low-rank adaptation module to achieve knowledge transfer.

Benefits of technology

It improves the performance of multilingual machine translation, alleviates parameter interference, maintains efficient parameter utilization, and significantly improves translation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354865A_ABST
    Figure CN120354865A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge migration method for multi-language machine translation, which can be applied to the technical field of natural language processing and large language models. The method comprises the following steps: performing language perceptibility evaluation on associated neurons related to a multi-language machine translation task in a teacher model to obtain a language perceptibility score set; classifying associated neurons in the teacher model into language general neurons and language specific neurons according to the language perceptibility score set, and extracting a weight sub-matrix from a weight matrix of the teacher model based on a classification result; and based on a machine translation task of a specific language, performing selective parameter updating on a shared low-rank adaptation module and a multi-language specific low-rank adaptation module in the student model by using the weight sub-matrix so as to complete knowledge migration from the teacher model to the student model. The method provided by the invention can solve the technical problems of parameter interference, knowledge forgetting, insufficient knowledge migration efficiency, insufficient language perception information utilization and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing and large language models, and particularly to a knowledge transfer method for multilingual machine translation. Background Art

[0002] Natural Language Processing (NLP) refers to the use of technical means to enable a computer to understand, generate, and interact with human language. The core tasks of natural language processing are machine translation, text classification, sentiment analysis, etc. Large Language Models (LLMs) have demonstrated powerful capabilities in the field of natural language processing. Especially in the task of Multilingual Machine Translation (MMT), they can handle translations between multiple languages.

[0003] When applying large language models to multilingual machine translation, the existing technical solutions still have the following technical problems: parameter interference. During the fine-tuning process of using a single model to handle multiple language pairs, the parameter optimization objectives for different language pairs may conflict with each other, resulting in parameter interference and affecting the overall translation performance; knowledge forgetting. The fine-tuning process may overwrite or damage the valuable knowledge learned by the model during the pre-training stage, including general language understanding capabilities and effective representations of specific languages; insufficient knowledge transfer efficiency. Existing knowledge transfer methods (such as standard knowledge distillation) mainly focus on the consistency of the output layer and fail to fully utilize and transfer the deeper, structured knowledge contained within the teacher model's parameters, especially language-specific knowledge; insufficient utilization of language perception information. Existing model adaptation techniques have not been able to effectively distinguish and utilize the language-general parameters and language-specific parameters within the model, limiting the optimization potential of the model's performance in a multilingual environment. Summary of the Invention

[0004] In view of the above problems, the present invention provides a knowledge transfer method for multilingual machine translation.

[0005] According to a first aspect of the present invention, there is provided a knowledge transfer method for multilingual machine translation, comprising:

[0006] Evaluating the language perception of the associated neurons related to the multilingual machine translation task in the teacher model to obtain a set of language perception scores;

[0007] Classifying the associated neurons in the teacher model into language-general neurons and language-specific neurons according to the set of language perception scores, and extracting weight sub-matrices from the weight matrix of the teacher model based on the classification results;

[0008] For a machine translation task based on a specific language, a weight sub-matrix is used to selectively update the parameters of the shared low-rank adaptation module and multiple language-specific low-rank adaptation modules in the student model, thereby completing the knowledge transfer from the teacher model to the student model.

[0009] According to an embodiment of the present invention, the above-mentioned evaluation of the language perception of the associated neurons related to the multi-language machine translation task in the teacher model to obtain the language perception score set includes:

[0010] Based on a representation analysis method, calculate the correlation score of each neural network layer in the teacher model for a specific language pair;

[0011] Based on the correlation score, determine the neural network layers related to the multi-language machine translation task from the teacher model to obtain the associated neural network layers;

[0012] Based on a preset perception measurement method, use the loss function used in the model training process to calculate the gradient output by each associated neuron in the associated neural network layer for a specific language;

[0013] Forward propagate and backward propagate the gradient output by each associated neuron for a specific language to obtain the specific language perception score of each associated neuron;

[0014] Repeat the calculation operation of the specific language perception score to obtain the language perception score set of each associated neuron for all languages.

[0015] According to an embodiment of the present invention, the above-mentioned preset perception measurement method includes a sensitivity-based perception measurement method, an activation value size-based perception measurement method, an adaptive gradient descent-based perception measurement method, a gradient pruning and model compression-based perception measurement method, a Brenner gradient-based perception measurement method, a Tenengrad gradient-based perception measurement method, a Laplace gradient-based perception measurement method, and a probe classifier-based perception measurement method.

[0016] According to an embodiment of the present invention, the above-mentioned classification of the associated neurons in the teacher model into language-general neurons and language-specific neurons according to the language perception score set includes:

[0017] In the case where the highest language perception score in the language perception score set of the associated neuron is lower than the preset threshold, classify the associated neuron as a language-general neuron;

[0018] In the case where the highest language perception score in the language perception score set of the associated neuron is greater than or equal to the preset threshold and the highest language perception score corresponds to a specific language, classify the associated neuron as a language-specific neuron;

[0019] Based on the classification results, define the associated neurons in the associated neural network layer as the language-general neuron index set and the language-specific neuron index set for each specific language pair.

[0020] According to an embodiment of the present invention, the above-mentioned extraction of the weight sub-matrix from the weight matrix of the teacher model based on the classification results includes:

[0021] Based on the language-general neuron index set and the dimension information of the weight matrix of the teacher model, use a preset parameter extraction method to extract a language-general weight sub-matrix compatible with the student model from the weight matrix of the teacher model;

[0022] Based on the language-specific neuron index set and the dimension information of the weight matrix of the teacher model, use a preset parameter extraction method to extract a language-specific weight sub-matrix compatible with the student model from the weight matrix of the teacher model.

[0023] According to an embodiment of the present invention, the above-mentioned preset parameter extraction method includes a parameter extraction method based on score Top-K selection, a parameter extraction method based on sparsification, a parameter extraction method based on gradient screening, a parameter extraction method based on linear structure, a parameter extraction method based on tree structure, and a parameter extraction method based on graph structure.

[0024] According to an embodiment of the present invention, for the above-mentioned machine translation task based on a specific language, use the weight sub-matrix to perform selective parameter updates on the shared low-rank adaptation module and multiple language-specific low-rank adaptation modules in the student model, and then complete the knowledge transfer from the teacher model to the student model, including:

[0025] Complete the initialization operation of multiple language-specific low-rank adaptation modules by performing a singular value decomposition operation on the language-general weight sub-matrix;

[0026] Complete the initialization operation of the shared low-rank adaptation module by performing a singular value decomposition operation on the language-specific weight sub-matrix;

[0027] Determine the parameter update granularity according to the requirements of the multi-language machine translation task or the computing resources, and perform parameter updates on the initialized multiple language-specific low-rank adaptation modules and the initialized shared low-rank adaptation module based on the parameter update granularity and a preset selective fine-tuning mechanism.

[0028] According to an embodiment of the present invention, the above-mentioned preset selective fine-tuning mechanism includes a selective fine-tuning mechanism based on a customized Adapter and a selective fine-tuning mechanism based on Prefix Tuning.

[0029] According to an embodiment of the present invention, the parameter update of the initialized multiple language-specific low-rank adaptation modules and the initialized shared low-rank adaptation module based on the parameter update granularity and the preset selective fine-tuning mechanism includes:

[0030] Freeze the original parameters of the student model;

[0031] In the case of processing a specific language pair, activate the initialized shared low-rank adaptation module, activate the initialized language-specific low-rank adaptation module corresponding to the specific language pair, freeze other initialized language-specific low-rank adaptation modules, calculate the original parameters of the student model, the parameters of the initialized shared low-rank adaptation module, and the parameters of the activated language-specific low-rank adaptation module to obtain the modified weights, and backpropagate the modified weights to update the parameters of the initialized shared low-rank adaptation module and the activated language-specific low-rank adaptation module.

[0032] According to an embodiment of the present invention, the above knowledge transfer method for multi-language machine translation further includes:

[0033] Replace the shared low-rank adaptation module with a parameter sharing module based on a regularization term, a parameter dynamic sharing module based on a gating mechanism, or a parameter sharing module based on pruning quantization.

[0034] The above knowledge transfer method for multi-language machine translation tasks provided by the present invention realizes efficient and accurate knowledge transfer from the teacher model to the student model by performing language-aware evaluation and parameter extraction on the teacher model and using the extracted parameters to update the parameters of the student model, and avoids the problem of parameter interference that may occur in multi-language machine translation tasks; by using the shared low-rank adaptation module and language-specific low-rank adaptation module in the student model, it can significantly improve the performance of multi-language machine translation, effectively alleviate parameter interference, and maintain a high parameter efficiency. Description of the Drawings

[0035] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0036] Figure 1 is an application scenario diagram of the knowledge transfer method for multi-language machine translation according to an embodiment of the present invention;

[0037] Figure 2 is a flowchart of the knowledge transfer method for multi-language machine translation according to an embodiment of the present invention;

[0038] Figure 3 is an architecture diagram of the knowledge transfer method for multi-language machine translation tasks according to an embodiment of the present invention;

[0039] Figure 4 is a structural block diagram of a knowledge transfer device for multi - language machine translation according to an embodiment of the present invention;

[0040] Figure 5 is a block diagram of an electronic device suitable for implementing a knowledge transfer method for multi - language machine translation according to an embodiment of the present invention. Detailed implementation manners

[0041] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present invention. However, obviously, one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well - known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0042] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. as used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0043] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0044] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0045] In the field of natural language processing technology, in order to enable large pre-trained models to efficiently adapt to downstream tasks (such as machine translation) while reducing computational and storage overheads, relevant researchers have developed parameter-efficient fine-tuning (PEFT) techniques. For example, Low-Rank Adaptation (LoRA) performs fine-tuning by introducing low-rank matrices into the original model layers. At the same time, Knowledge Distillation (KD) is used to transfer the knowledge possessed by a large and superior-performing "teacher" model to a smaller and more efficient "student" model, enabling the student model to achieve performance close to that of the teacher model. Meanwhile, research on the internal structure of large language models by relevant researchers has shown that there are specific neurons or parameters in the model that exhibit different activation patterns when processing different languages. Some parameters tend to handle cross-lingual general language features (language-general parameters), while others may be specifically responsible for handling the characteristics of specific languages or language pairs (language-specific parameters).

[0046] However, in the field of natural language processing technology, existing technical solutions with at least one of parameter-efficient fine-tuning technology, knowledge distillation technology, and knowledge transfer technology still have one of the following technical problems: parameter interference, knowledge forgetting, insufficient knowledge transfer efficiency, insufficient utilization of language perception information, and other technical problems.

[0047] To solve one of the existing technical problems, the present invention provides a knowledge transfer method for multilingual machine translation tasks, which includes a Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine Translation (MLAS-LoRA). This method realizes efficient and low-interference fine-tuning by identifying and extracting language-aware parameters in the teacher model and injecting this knowledge into the student model using a new multiple LoRA structure.

[0048] The above knowledge transfer method provided by the present invention will be further described in detail below through specific embodiments, specific implementation manners, and specific experiments.

[0049] Figure 1 It is an application scenario diagram of the knowledge transfer method for multilingual machine translation according to an embodiment of the present invention.

[0050] As Figure 1As shown, the application scenario 100 according to this embodiment may include natural language processing and large language models. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0051] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0052] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.

[0053] The server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background management server can analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0054] It should be noted that the knowledge transfer method for multilingual machine translation provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the knowledge transfer device for multilingual machine translation provided by the embodiments of the present invention can generally be set in the server 105. The knowledge transfer method for multilingual machine translation provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the knowledge transfer device for multilingual machine translation provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0055] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers therein are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0056] Based on the Figure 1 scenario described below, through Figures 2 to 3 a knowledge transfer method for multilingual machine translation of the disclosed embodiments will be described in detail.

[0057] Figure 2 is a flowchart of a knowledge transfer method for multilingual machine translation according to an embodiment of the present invention.

[0058] As Figure 2 shown, the knowledge transfer method for multilingual machine translation of this embodiment includes operations S210 to S230.

[0059] In operation S210, the associated neurons related to the multilingual machine translation task in the teacher model are evaluated for language perception to obtain a set of language perception scores.

[0060] During the process of evaluating language perception, first determine the associated neural network layers related to the multilingual machine translation task in the teacher model. For example, if the teacher model is a Transformer-based model, then find the Transformer layer most relevant to multilingual machine translation in the teacher model, and then determine the associated neurons from these most relevant Transformer layers.

[0061] In operation S220, the associated neurons in the teacher model are classified into language-general neurons and language-specific neurons according to the set of language perception scores, and a weight submatrix is extracted from the weight matrix of the teacher model based on the classification result.

[0062] In operation S230, based on the machine translation task of a specific language, the weight submatrix is used to selectively update the parameters of the shared low-rank adaptation module and multiple language-specific low-rank adaptation modules in the student model, thereby completing the knowledge transfer from the teacher model to the student model.

[0063] According to an embodiment of the present invention, the above knowledge transfer method for multilingual machine translation further includes: replacing the shared low-rank adaptation module with a parameter sharing module based on a regularization term, a parameter dynamic sharing module based on a gating mechanism, or a parameter sharing module based on pruning quantization.

[0064] The above-mentioned knowledge transfer method for multilingual machine translation tasks provided by the present invention realizes efficient and accurate knowledge transfer from the teacher model to the student model by performing language perception evaluation and parameter extraction on the teacher model, and using the extracted parameters to update the parameters of the student model, and avoids the problem of parameter interference that may occur in multilingual machine translation tasks; by using the shared low-rank adaptation module and language-specific low-rank adaptation module in the student model, it can significantly improve the performance of multilingual machine translation, effectively alleviate parameter interference, and maintain high parameter efficiency.

[0065] Figure 3 It is an architecture diagram of a knowledge transfer method for multilingual machine translation tasks according to an embodiment of the present invention.

[0066] As Figure 3 shown, the above-mentioned knowledge transfer method for multilingual machine translation tasks provided by the present invention mainly includes steps such as language perception evaluation and parameter extraction, and knowledge injection and fine-tuning of multi-language perception LoRA. Among them, FFN (Feed-Forward Network) is a feed-forward neural network; Add+Norm is the combination of two operations: residual connection (Add) and layer normalization (Layer Normalization).

[0067] In the process of language perception evaluation and parameter extraction, evaluate the perception of neurons in the teacher model for different language pairs; classify the neurons as language-general or language-specific according to the perception; extract the corresponding parameter subsets from the teacher model based on the classification results; construct a student model architecture including a shared language-general LoRA module and multiple language-specific LoRA modules; initialize the LoRA modules with the extracted parameter subsets; and when fine-tuning for a specific language pair, selectively update only the shared general LoRA module and the corresponding specific LoRA module. The extraction of the parameter subsets is guided by the neuron language perception scores and takes into account the dimensions of the student model. The initialization of the LoRA modules is based on the singular value decomposition (SVD) of the extracted parameter subsets.

[0068] In the process of the selective fine-tuning mechanism, when processing data of a specific language pair, freeze the original parameters of the student model and the specific LoRA modules of non-corresponding language pairs, and only activate and update the general LoRA module and the specific LoRA module of the current language pair.

[0069] According to an embodiment of the present invention, the above-mentioned evaluation of the language perception of the associated neurons related to the multilingual machine translation task in the teacher model, and the obtained language perception score set includes: calculating the correlation score of each neural network layer in the teacher model for a specific language pair based on a representation analysis method; determining the neural network layers related to the multilingual machine translation task from the teacher model based on the correlation score to obtain associated neural network layers; calculating the gradient output by each associated neuron in the associated neural network layer for a specific language using the loss function used in the model training process based on a preset perception metric method; performing forward propagation and backward propagation on the gradient output by each associated neuron for a specific language to obtain the specific language perception score of each associated neuron; repeating the calculation operation of the specific language perception score to obtain the language perception score set of each associated neuron for all languages.

[0070] The above technical solution for obtaining the language perception score set of the associated neurons for all languages can accurately identify the translation sensitivity of the teacher model to different languages, improve the translation performance and effect for specific languages; at the same time, the perception score calculation based on gradient propagation provides a quantitative basis for the adjustment of multilingual model parameters.

[0071] According to an embodiment of the present invention, the above-mentioned preset perception metric methods include a perception metric method based on sensitivity, a perception metric method based on activation value magnitude, a perception metric method based on adaptive gradient descent, a perception metric method based on gradient pruning and model compression, a perception metric method based on Brenner gradient, a perception metric method based on Tenengrad gradient, a perception metric method based on Laplace gradient, and a perception metric method based on a probe classifier.

[0072] Those skilled in the art can select one or more of the above perception metric methods according to actual needs, or select other neuron perception metric methods.

[0073] According to an embodiment of the present invention, the above-mentioned classification of the associated neurons in the teacher model into language-general neurons and language-specific neurons according to the language perception score set includes: in the case where the highest language perception score in the language perception score set of the associated neurons is lower than a preset threshold, classifying the associated neurons as language-general neurons; in the case where the highest language perception score in the language perception score set of the associated neurons is greater than or equal to the preset threshold and the highest language perception score corresponds to a specific language, classifying the associated neurons as language-specific neurons; based on the classification result, defining the associated neurons in the associated neural network layer as a language-general neuron index set and a language-specific neuron index set for each specific language pair.

[0074] Through the dynamic partitioning mechanism of the preset threshold, the above embodiments achieve the physical isolation of the general language features and the language-specific features, improving the utilization rate of the model parameters; meanwhile, the establishment of the general language neuron index set provides a stable carrier for cross-lingual knowledge transfer and accelerates the convergence speed of the low-resource language translation task.

[0075] According to an embodiment of the present invention, the above-mentioned extraction of the weight sub-matrix from the weight matrix of the teacher model based on the classification result includes: based on the general language neuron index set and the dimension information of the weight matrix of the teacher model, using a preset parameter extraction method to extract a general language weight sub-matrix compatible with the student model from the weight matrix of the teacher model; based on the language-specific neuron index set and the dimension information of the weight matrix of the teacher model, using a preset parameter extraction method to extract a language-specific weight sub-matrix compatible with the student model from the weight matrix of the teacher model.

[0076] According to an embodiment of the present invention, the above-mentioned preset parameter extraction methods include a parameter extraction method based on score Top-K selection, a parameter extraction method based on sparsification, a parameter extraction method based on gradient screening, a parameter extraction method based on linear structure, a parameter extraction method based on tree structure, and a parameter extraction method based on graph structure.

[0077] Those skilled in the art can select one or more of the above parameter extraction methods according to the actual translation task requirements, or adopt other parameter extraction methods.

[0078] The above embodiments achieve the efficient transfer of the teacher model knowledge through preset parameter extraction methods (such as Top-K, sparsification, graph structure, etc.); using the dynamic adaptation mechanism of the weight sub-matrix, it supports the knowledge transfer across model architectures (CNN, Transformer, etc.).

[0079] The following further elaborates on the language perception evaluation and parameter extraction process involved in the present invention through specific embodiments.

[0080] The language perception evaluation and parameter extraction process includes the teacher model related layer selection operation, neuron language perception scoring operation, neuron classification operation, and parameter extraction operation.

[0081] Among them, for the teacher model related layer selection operation: first, it is necessary to determine the Transformer layer in the teacher model that is most relevant to the machine translation task. A method based on representation analysis is adopted to calculate the correlation score for each layer relative to a specific language pair . This score is measured by calculating the L1 norm of the difference between the activation vectors of consecutive layers, as shown in formula (1):

[0082] (1),

[0083] where N represents the total number of forward propagations, represents the first forward propagation, represents the th forward propagation during the th layer of the neural network's activation vector, represents the th forward propagation during the th layer of the neural network's activation vector, represents the th forward propagation's input vector; represents the L1 norm; The present invention selects the layers with the highest scores, where matches the number of layers of the student model.

[0084] Among them, the neuron language perception scoring operation: within the selected layers, further evaluate each neuron 's "perception" intensity for a specific language pair . Adopt a sensitivity-based method (inspired by Taylor expansion) to define the neuron language perception score , as shown in formula (2):

[0085] (2),

[0086] where is the loss function, is the output of neuron in layer , is the gradient of the loss with respect to the output of this neuron, represents the absolute value. This score approximately represents the change in the loss function if the output of neuron is set to zero. This score is calculated by performing forward and backward propagations on the seed sentence of a specific language pair.

[0087] Among them, the neuron classification operation: analyze the set of language perception scores of each neuron on all specific language pairs, as shown in formula (3):

[0088] (3),

[0089] where represents the th specific language pair 's perception score.

[0090] Set a threshold (e.g., set to 0.2 according to experience). If the neuron 's highest language perception score is less than , then it is classified as a language-general neuron. Such neurons capture general linguistic knowledge across languages. If the highest language perception score is greater than or equal to , and this highest score corresponds to the th specific language pair , then it is classified as a language-specific neuron (for the th specific language pair ). Such neurons specifically handle the nuances of a specific language pair and help alleviate parameter interference. According to the classification results, define the set of language-general neuron indices for each selected layer and the set of language-specific neuron indices for each specific language pair

[0091] Among them, parameter extraction: Considering that the dimension of the weight matrix of the teacher model is usually larger than the dimension of the weight matrix of the student model , it is necessary to extract a submatrix compatible with the student model from the weight matrix of the teacher model . Using the neuron index set and language perception scores obtained in the previous step, through a perception-based extraction function Extract(.), extract the language-general submatrix and the specific submatrix for each specific language pair respectively, where gen is the abbreviation of general and spe is the abbreviation of specific, as shown in formulas (4) and (5):

[0092] (4),

[0093] (5).

[0094] The above extraction process aims to preserve the structural integrity of the most relevant parts in the teacher model.

[0095] According to an embodiment of the present invention, for the above-mentioned machine translation task based on a specific language, the knowledge transfer from the teacher model to the student model is completed by selectively updating the parameters of the shared low-rank adaptation module and multiple language-specific low-rank adaptation modules in the student model using weight submatrices, including: initializing multiple language-specific low-rank adaptation modules by performing singular value decomposition on the language-general weight submatrix; initializing the shared low-rank adaptation module by performing singular value decomposition on the language-specific weight submatrix; determining the parameter update granularity according to the requirements of the multilingual machine translation task or computing resources, and updating the parameters of the initialized multiple language-specific low-rank adaptation modules and the initialized shared low-rank adaptation module based on the parameter update granularity and a preset selective fine-tuning mechanism.

[0096] According to an embodiment of the present invention, the above-mentioned preset selective fine-tuning mechanism includes a selective fine-tuning mechanism based on a customized Adapter and a selective fine-tuning mechanism based on Prefix Tuning.

[0097] According to an embodiment of the present invention, the updating of the parameters of the initialized multiple language-specific low-rank adaptation modules and the initialized shared low-rank adaptation module based on the parameter update granularity and the preset selective fine-tuning mechanism includes: freezing the original parameters of the student model; in the case of processing a specific language pair, activating the initialized shared low-rank adaptation module, activating the initialized language-specific low-rank adaptation module corresponding to the specific language pair, freezing other initialized language-specific low-rank adaptation modules, calculating the original parameters of the student model, the parameters of the initialized shared low-rank adaptation module, and the parameters of the activated language-specific low-rank adaptation module to obtain modified weights, and performing backpropagation on the modified weights to update the parameters of the initialized shared low-rank adaptation module and the activated language-specific low-rank adaptation module.

[0098] The knowledge injection and fine-tuning process of the above-mentioned multilingual-aware LoRA involved in the above embodiment will be further described in detail through specific embodiments below.

[0099] The knowledge injection and fine-tuning of multilingual-aware LoRA include multi-LoRA architecture design, LoRA module initialization, and parameter selective fine-tuning.

[0100] Among them, the multi-LoRA architecture design: Based on the standard LoRA, the present invention modifies the pre-trained weights through low-rank decomposition BA , but expands it, where represents the pre-trained weights, represents the modified weights, represents the first low-rank matrix in the LoRA module and the second low-rank matrix Perform matrix multiplication. In each (or selected) layer of the student model, introduce a dual LoRA structure: a shared language-general LoRA module: , where represents the first general low-rank matrix of the shared language-general LoRA module, represents the second general low-rank matrix of the shared language-general LoRA module, which is shared among all language pairs and is used to capture and transfer general translation knowledge. Multiple language-specific LoRA modules: , represents the th language-specific LoRA module for the specific language pair , where represents the first specific low-rank matrix of the language-specific LoRA module, represents the second specific low-rank matrix of the language-specific LoRA module, and each module corresponds to the th specific language pair , which is used to capture and transfer the knowledge unique to this language pair and mitigate parameter interference by isolating language-specific adjustments.

[0101] Among them, LoRA module initialization: Use the sub-matrices extracted in the first stage to initialize these LoRA modules. Perform singular value decomposition (SVD) on the extracted language-general sub-matrix as shown in formula (6):

[0102] (6),

[0103] where represents the left singular matrix, which is an orthogonal matrix; represents the singular value matrix, which is a diagonal matrix; represents the transpose of the right singular matrix, which is an orthogonal matrix.

[0104] Then, initialize the general LoRA module according to the first preset rank as shown in formulas (7) and (8):

[0105] (7),

[0106] (8),

[0107] where represents the left singular matrix, which is an orthogonal matrix; represents the singular value matrix, which is a diagonal matrix; represents the transpose of the right singular matrix, which is an orthogonal matrix.

[0108] Similarly, for each extracted language-specific submatrix perform SVD as shown in Equation (9):

[0109] (9),

[0110] where, represents the left singular matrix, which is an orthogonal matrix; represents the singular value matrix, which is a diagonal matrix; represents the transpose of the right singular matrix, which is an orthogonal matrix.

[0111] Then, initialize the corresponding specific LoRA module according to the second preset rank as shown in Equations (10) and (11):

[0112] (10),

[0113] (11),

[0114] where, represents the left singular matrix, which is an orthogonal matrix; represents the singular value matrix, which is a diagonal matrix; represents the transpose of the right singular matrix, which is an orthogonal matrix.

[0115] Among them, parameter selective fine-tuning: During the fine-tuning process, freeze the original parameters of the student model . When processing the input data of the th specific language pair : Activate the shared language-general LoRA module . Activate the language-specific LoRA module corresponding to the current specific language pair . Freeze the specific LoRA modules of all other specific language pairs. Calculate the modified weights as shown in Equation (12):

[0116] (12),

[0117] where, represents the first specific low-rank matrix in the LoRA module and the second specific low-rank matrix perform matrix multiplication.

[0118] During the backpropagation process, only update the parameters of the activated LoRA modules (i.e., ). This selective update strategy further reduces the number of trainable parameters and effectively prevents interference between different language pairs.

[0119] Those skilled in the art can adjust the granularity of selective update according to task requirements or computing resources. For example, some partial parameters of certain layers can be selectively updated instead of the entire LoRA module.

[0120] Those skilled in the art can use other parameter-efficient fine-tuning techniques (such as Adapter, Prefix Tuning, etc.) according to task requirements or computing resources, and transform them not limited to LoRA, to achieve a similar dual-module (general + specific) structure and selective update mechanism. For example, shared general Adapters and multiple language-specific Adapters can be designed.

[0121] The above knowledge transfer method for multilingual machine translation provided by the present invention will be further described in detail through specific experiments below.

[0122] In terms of model selection: Gemma-2-9b-it is selected as the teacher model, and Gemma-2-2b-it is selected as the student model.

[0123] Dataset: 14 language pairs (cs-en, en-cs, de-en, en-de, et-en, en-et, fi-en, en-fi, ru-en, en-ru, tr-en, en-tr, zh-en, en-zh) in the WMT18 dataset are used. 200,000 sentence pairs are randomly selected for fine-tuning in each translation direction, and 1,000 sentence pairs are used to calculate the neuron language perception scores.

[0124] Data processing: Parallel data is processed using a unified translation instruction template ("Translate from [SRC] to [TGT]:").

[0125] Specific experimental steps (taking the en-de translation task as an example):

[0126] Perception evaluation and extraction: Using 1,000 seed sentence pairs of en-de, through the forward and backward propagation of the teacher model (Gemma-2-9b-it), calculate the language perception scores of neurons for en-de in each relevant layer. At the same time, calculate the language perception scores of the neurons for the other 13 language pairs. According to all the scores and the threshold λ = 0.2, classify the neurons as "general" or "en-de specific" (or other specific language pairs). According to the classification results, extract the general submatrix and the specific submatrix from the weights of the teacher model to match the dimensions of the student model (Gemma-2-2b-it).

[0127] LoRA Initialization: Perform SVD on the general sub-matrix to initialize the shared general LoRA module. Perform SVD on the specific sub-matrix to initialize the en-de specific LoRA module (repeat the extraction and initialization process of this specific LoRA for the other 13 specific language pairs).

[0128] Selective Fine-tuning: Fine-tune the student model using 200,000 sentence pairs of training data for en-de. Set the fine-tuning epoch to 3, batch size to 64, use the AdamW optimizer, learning rate to 1e-4, and gradient accumulation steps to 8. In each training step, freeze the original parameters of the student model and all non-en-de specific LoRA modules. Only update the parameters of the general LoRA module and the en-de specific LoRA module.

[0129] Evaluation: Use the WMT18 test set and BLEU metric to evaluate the translation performance of the fine-tuned model on en-de and other language pairs, as shown in Table 1.

[0130] Table 1 shows the comparison results of BLEU scores between the method of the present invention (MLAS-LORA) and multiple baseline methods on 10 language pairs of the WMT18 dataset (including xx-to-English and English-to-xx translation directions) (the experiment is based on the Gemma-2-2b-it student model and Gemma-2-9b-it teacher model).

[0131] Table 1 Comparison Results of BLEU Scores

[0132]

[0133] As shown in Table 1, the MLAS-LORA method proposed by the present invention has consistent and significantly better BLEU scores than all the compared baseline methods in all tested language pairs and translation directions. The average BLEU score is about 1.7 points higher than that of the strong baseline. This fully demonstrates the effectiveness and superiority of the technical solution of the present invention in improving the performance of multi-language machine translation, that is, compared with the prior art, the method provided by the present invention can significantly improve the performance of multi-language machine translation, effectively alleviate parameter interference, and maintain high parameter efficiency.

[0134] Based on the above knowledge transfer method for multi-language machine translation, the present invention also provides a knowledge transfer device for multi-language machine translation. The following will be combined with Figure 4 Describe this device in detail.

[0135] Figure 4 is the structural block diagram of the knowledge transfer device for multi-language machine translation according to an embodiment of the present invention.

[0136] As Figure 4As shown in the figure, the knowledge transfer device 400 for multilingual machine translation in this embodiment includes a language perception evaluation module 410, a neuron classification and parameter extraction module 420, and a knowledge transfer module 430.

[0137] The language perception evaluation module 410 is configured to evaluate the language perception of the associated neurons related to the multilingual machine translation task in the teacher model to obtain a set of language perception scores. In one embodiment, the language perception evaluation module 410 can be used to perform the operation S210 described above, which will not be elaborated here.

[0138] The neuron classification and parameter extraction module 420 is configured to classify the associated neurons in the teacher model into language-general neurons and language-specific neurons according to the set of language perception scores, and extract a weight sub-matrix from the weight matrix of the teacher model based on the classification result. In one embodiment, the neuron classification and parameter extraction module 420 can be used to perform the operation S220 described above, which will not be elaborated here.

[0139] The knowledge transfer module 430 is configured to perform selective parameter update on the shared low-rank adaptation module and multiple language-specific low-rank adaptation modules in the student model based on the machine translation task of a specific language, so as to complete the knowledge transfer from the teacher model to the student model. In one embodiment, the knowledge transfer module 430 can be used to perform the operation S230 described above, which will not be elaborated here.

[0140] According to an embodiment of the present invention, any multiple of the language perception evaluation module 410, the neuron classification and parameter extraction module 420, and the knowledge transfer module 430 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the language perception evaluation module 410, the neuron classification and parameter extraction module 420, and the knowledge transfer module 430 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the language perception evaluation module 410, the neuron classification and parameter extraction module 420, and the knowledge transfer module 430 can be at least partially implemented as a computer program module, and when the computer program module is run, it can perform the corresponding functions.

[0141] Figure 5 A block diagram of an electronic device suitable for implementing a knowledge transfer method for multilingual machine translation according to an embodiment of the present invention.

[0142] As Figure 5 shown, the electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage section 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include on-board memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0143] In the RAM 503, various programs and data required for the operation of the electronic device 500 are stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The processor 501 performs various operations of the method flow according to an embodiment of the present invention by executing the program in the ROM 502 and / or the RAM 503. It should be noted that the program may also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 may also perform various operations of the method flow according to an embodiment of the present invention by executing the program stored in the one or more memories.

[0144] According to an embodiment of the present invention, the electronic device 500 may further include an input / output (I / O) interface 505, and the input / output (I / O) interface 505 is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read from it can be installed into the storage section 508 as needed.

[0145] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present invention are implemented.

[0146] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than the ROM 502 and RAM 503.

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0148] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0149] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A knowledge transfer method for multilingual machine translation, characterized in that, The method includes: Evaluating the language perception of the associated neurons related to the multilingual machine translation task in the teacher model to obtain a set of language perception scores; Classifying the associated neurons in the teacher model into language-general neurons and language-specific neurons according to the set of language perception scores, and extracting a weight submatrix from the weight matrix of the teacher model based on the classification result; Based on a machine translation task for a specific language, selectively updating the parameters of the shared low-rank adaptation module and multiple language-specific low-rank adaptation modules in the student model using the weight submatrix, thereby completing the knowledge transfer from the teacher model to the student model.

2. The method according to claim 1, wherein Evaluating the language perception of the associated neurons related to the multilingual machine translation task in the teacher model to obtain a set of language perception scores, including: Calculating the correlation score of each neural network layer in the teacher model for a specific language pair based on a representation analysis method; Determining the neural network layers related to the multilingual machine translation task from the teacher model based on the correlation score to obtain associated neural network layers; Based on a preset perception metric method, calculating the gradient of each associated neuron in the associated neural network layer for a specific language output using the loss function used in the model training process; Performing forward propagation and backward propagation on the gradient of each associated neuron for a specific language output to obtain the specific language perception score of each associated neuron; Repeating the calculation operation of the specific language perception score to obtain a set of language perception scores of each associated neuron for all languages.

3. The method according to claim 2, characterized in that The preset perception metric method includes a sensitivity-based perception metric method, an activation value size-based perception metric method, an adaptive gradient descent-based perception metric method, a gradient pruning and model compression-based perception metric method, a Brenner gradient-based perception metric method, a Tenengrad gradient-based perception metric method, a Laplace gradient-based perception metric method, and a probe classifier-based perception metric method.

4. The method according to claim 2, wherein Classifying the associated neurons in the teacher model into language-general neurons and language-specific neurons according to the set of language perception scores includes: In the case where the highest language perception score in the set of language perception scores of the associated neurons is lower than a preset threshold, classifying the associated neurons as language-general neurons; In the case where the highest language perception score in the set of language perception scores of the associated neurons is greater than or equal to the preset threshold and the highest language perception score corresponds to the specific language, classifying the associated neurons as language-specific neurons; Based on the classification result, defining the associated neurons in the associated neural network layer as a language-general neuron index set and a language-specific neuron index set for each specific language pair.

5. The method according to claim 4, characterized in that Extracting a weight submatrix from the weight matrix of the teacher model based on the classification result includes: Based on the dimension information of the language-general neuron index set and the weight matrix of the teacher model, use a preset parameter extraction method to extract a language-general weight sub-matrix compatible with the student model from the weight matrix of the teacher model; Based on the language-specific neuron index set and the dimension information of the weight matrix of the teacher model, use the preset parameter extraction method to extract a language-specific weight sub-matrix compatible with the student model from the weight matrix of the teacher model.

6. The method according to claim 5, characterized in that, The preset parameter extraction method includes a parameter extraction method based on score Top-K selection, a parameter extraction method based on sparsification, a parameter extraction method based on gradient screening, a parameter extraction method based on linear structure, a parameter extraction method based on tree structure, and a parameter extraction method based on graph structure.

7. The method according to claim 5, characterized in that, Based on a machine translation task for a specific language, using the weight sub-matrix to perform selective parameter updates on the shared low-rank adaptation module and multiple language-specific low-rank adaptation modules in the student model, thereby completing the knowledge transfer from the teacher model to the student model includes: Complete the initialization operation of multiple language-specific low-rank adaptation modules by performing a singular value decomposition operation on the language-general weight sub-matrix; Complete the initialization operation of the shared low-rank adaptation module by performing a singular value decomposition operation on the language-specific weight sub-matrix; Determine the parameter update granularity according to the requirements of the multilingual machine translation task or computing resources, and perform parameter updates on the initialized multiple language-specific low-rank adaptation modules and the initialized shared low-rank adaptation module based on the parameter update granularity and a preset selective fine-tuning mechanism.

8. The method according to claim 7, wherein The preset selective fine-tuning mechanism includes a selective fine-tuning mechanism based on a customized Adapter and a selective fine-tuning mechanism based on Prefix Tuning.

9. The method according to claim 7, wherein Performing parameter updates on the initialized multiple language-specific low-rank adaptation modules and the initialized shared low-rank adaptation module based on the parameter update granularity and the preset selective fine-tuning mechanism includes: Freeze the original parameters of the student model; In the case of processing the specific language pair, activate the initialized shared low-rank adaptation module, activate the initialized language-specific low-rank adaptation module corresponding to the specific language pair, freeze other initialized language-specific low-rank adaptation modules, calculate the original parameters of the student model, the parameters of the initialized shared low-rank adaptation module, and the parameters of the activated language-specific low-rank adaptation module to obtain modified weights, and backpropagate the modified weights to perform parameter updates on the initialized shared low-rank adaptation module and the activated language-specific low-rank adaptation module.

10. The method according to claim 1, wherein It also includes: Replace the shared low-rank adaptation module with a parameter sharing module based on a regularization term, a parameter dynamic sharing module based on a gating mechanism, or a parameter sharing module based on pruning and quantization.

Citation Information

Patent Citations

  • Output regularization method based on teacher model classification layer weight

    CN114782742A

  • Language model training method, language task processing method and system

    CN119721297A

  • Method and apparatus for training student model for image processing

    WO2022077646A1

  • KR20230015675A