Multi-expert parameter efficient fine-tuning method based on expert perception difference initialization

CN122713360APending Publication Date: 2026-09-08HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610791274.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种基于专家感知差分初始化的多专家参数高效微调方法,以至少解决相关技术中基于低秩适配的专家模型训练耗时长的问题

Benefits of technology

[0041] Compared to related technologies, the efficient fine-tuning method for multi-expert parameters based on expert-aware differential initialization provided in this application involves inputting training data into a pre-trained model. Utilizing the routing mechanism of the pre-trained model, it determines the LoRA expert corresponding to the hidden state vector of the training data and constructs a data sample set corresponding to the LoRA expert based on the hidden state vector. The change patterns of the data sample set are extracted to obtain feature vectors, and a down-projection matrix of the LoRA expert is constructed based on these feature vectors. This down-projection matrix represents the hidden state of the LoRA expert, thus guiding the LoRA expert to learn the intrinsic distribution characteristics of different data subspaces before training. Subsequently, the parameters of the LoRA expert are unfrozen, and the pre-trained model is trained. This allows the pre-trained model to achieve higher accuracy in the early stages of training, significantly improving training efficiency. This overcomes the long training time of traditional low-rank adaptation expert models and reduces the memory overhead required for training low-rank adaptation expert models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122713360A_ABST
    Figure CN122713360A_ABST
Patent Text Reader

Abstract

This application relates to an efficient fine-tuning method for multi-expert parameters based on expert-perceptual differential initialization. The method includes: acquiring a pre-trained model including LoRA experts and training data, and freezing the parameters of the LoRA experts in the pre-trained model; inputting the training data into the pre-trained model to obtain the corresponding hidden state vectors generated by the pre-trained model based on the training data, as well as the LoRA experts corresponding to each training data route, and constructing a data sample set corresponding to the LoRA experts based on the hidden state vectors; extracting the change patterns of the data sample set to obtain the feature vectors corresponding to the LoRA experts, and obtaining the downprojection matrix of the corresponding LoRA experts based on the feature vectors; unfreezing the parameters of the LoRA experts, and training the pre-trained model. This application solves the problem of long training time for expert models based on low-rank adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer science, and in particular to an efficient method for fine-tuning multi-expert parameters based on expert-perceptual differential initialization. Background Technology

[0002] With the widespread application of large-scale pre-trained models in fields such as natural language processing and computer vision, the number of model parameters and computational overhead continue to increase. Expert models based on low-rank adaptation can achieve model capability expansion and task adaptation under limited computing resources, and have become an important technical approach for lightweight fine-tuning of large models and multi-task collaborative optimization.

[0003] In related technologies, expert models based on low-rank adaptation randomly initialize low-rank matrices for LoRA (Low-Rank Adaptation) experts during the training phase, and only update the parameters of these low-rank matrices during training. However, during this training and optimization process, each LoRA expert is prone to getting stuck in similar parameter space regions, resulting in a longer optimization and convergence path for the expert model and increasing training time.

[0004] Currently, no effective solution has been proposed to address the issue of long training time for expert models based on low-rank adaptation in related technologies. Summary of the Invention

[0005] This application provides an efficient fine-tuning method for multi-expert parameters based on expert-perceptual differential initialization, which at least solves the problem of long training time for expert models based on low-rank adaptation in related technologies.

[0006] In a first aspect, embodiments of this application provide an efficient fine-tuning method for multi-expert parameters based on expert-perceptual differential initialization, the method comprising:

[0007] Obtain a pre-trained model including LoRA experts and training data, and freeze the parameters of the LoRA experts in the pre-trained model;

[0008] The training data is input into the pre-trained model to obtain the corresponding hidden state vector generated by the pre-trained model based on the training data, as well as the LoRA expert corresponding to each training data, and a data sample set corresponding to the LoRA expert is constructed based on the hidden state vector.

[0009] Extract the variation patterns of the data sample set to obtain the feature vectors corresponding to the LoRA expert, and obtain the down projection matrix of the corresponding LoRA expert based on the feature vectors;

[0010] Unfreeze the parameters of the LoRA expert and train the pre-trained model.

[0011] In some embodiments, the variation patterns of the data sample set are extracted to obtain feature vectors corresponding to the LoRA expert, and the downprojection matrix of the corresponding LoRA expert is obtained based on the feature vectors, including:

[0012] Extract the variation patterns of the data sample set to obtain the principal component matrix of the data sample set, wherein the principal component matrix includes multiple feature vectors corresponding to LoRA experts;

[0013] In the principal component matrix, a specified number of target feature vectors are extracted according to the eigenvalues ​​sorted from largest to smallest, and the specified number is the same as the rank parameter of the LoRA expert;

[0014] Construct the lower projection matrix of the corresponding LoRA expert based on a specified number of target feature vectors.

[0015] In some embodiments, obtaining the lower projection matrix corresponding to the LoRA expert based on the feature vector includes:

[0016] Obtain a first initialization matrix composed of the feature vectors, and a randomly generated second initialization matrix;

[0017] Based on the mixing intensity parameter, the first initialization matrix and the second initialization matrix are fused to obtain the fusion matrix;

[0018] Noise is added to the fusion matrix to obtain the lower projection matrix of the corresponding LoRA expert.

[0019] In some embodiments, the pre-trained model further includes a backbone network and a gating network. The step of inputting the training data into the pre-trained model to obtain the corresponding hidden state vectors generated by the pre-trained model based on the training data, and the LoRA experts for each route corresponding to the training data, and constructing a data sample set corresponding to the LoRA experts based on the hidden state vectors, includes:

[0020] The training data is divided into token sequences, and the token sequences are input into the backbone network to obtain the input hidden state vector of each token output by the backbone network.

[0021] Based on the routing decisions of the gated network for each token in the pre-trained model, the LoRA expert for the route corresponding to the token and the output hidden state vector output by the LoRA expert based on the input hidden state vector are obtained.

[0022] The input hidden state vectors and output hidden state vectors of each token are integrated according to the corresponding LoRA expert to obtain the data sample set of the LoRA expert.

[0023] In some embodiments, the step of integrating the input hidden state vectors and output hidden state vectors of each token according to the corresponding LoRA expert to obtain the data sample set of the LoRA expert includes:

[0024] Obtain the gating value corresponding to the LoRA expert, calculated by the gating network in the pre-trained model based on the input hidden state vector of the token;

[0025] Among the multiple tokens, obtain the target token whose gating value is greater than or equal to a preset threshold;

[0026] The input hidden state vector and output hidden state vector of the target token are integrated according to the corresponding LoRA expert to obtain the data sample set of the LoRA expert.

[0027] In some embodiments, the method further includes:

[0028] Obtain the frequency at which the LoRA expert is activated by the target token;

[0029] If the frequency is lower than the preset frequency, adjust the routing strategy of the LoRA expert in the pre-trained model.

[0030] In some embodiments, after obtaining the lower projection matrix of the corresponding LoRA expert based on the eigenvectors, the method further includes:

[0031] The lower projection matrix is ​​normalized using norm normalization.

[0032] Obtain the variance scaling factor, and perform variance scaling processing on the lower projection matrix according to the variance scaling factor.

[0033] In some embodiments, the method further includes:

[0034] If the storage usage of the LoRA expert's data sample set exceeds a specified threshold, execute at least one of the following storage strategies: store the LoRA expert's data sample set in the central processing unit; convert the data format of the LoRA expert's data sample set to a specified format; and downsample the LoRA expert's data sample set.

[0035] Secondly, embodiments of this application provide a highly efficient fine-tuning device for expert parameters based on expert-perceptual differential initialization, comprising:

[0036] The acquisition module is used to acquire a pre-trained model including LoRA experts and training data, and freeze the parameters of the LoRA experts in the pre-trained model.

[0037] The sample construction module is used to input the training data into the pre-trained model, obtain the corresponding hidden state vector generated by the pre-trained model based on the training data, and the LoRA expert corresponding to each training data route, and construct a data sample set corresponding to the LoRA expert based on the hidden state vector;

[0038] The matrix construction module is used to extract the change patterns of the data sample set, obtain the feature vectors corresponding to the LoRA expert, and obtain the down projection matrix of the corresponding LoRA expert based on the feature vectors.

[0039] The training module is used to unfreeze the parameters of the LoRA expert and train the pre-trained model.

[0040] Thirdly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements efficient fine-tuning of expert parameters based on expert-perceptual differential initialization as described in the first aspect above.

[0041] Compared to related technologies, the efficient fine-tuning method for multi-expert parameters based on expert-aware differential initialization provided in this application involves inputting training data into a pre-trained model. Utilizing the routing mechanism of the pre-trained model, it determines the LoRA expert corresponding to the hidden state vector of the training data and constructs a data sample set corresponding to the LoRA expert based on the hidden state vector. The change patterns of the data sample set are extracted to obtain feature vectors, and a down-projection matrix of the LoRA expert is constructed based on these feature vectors. This down-projection matrix represents the hidden state of the LoRA expert, thus guiding the LoRA expert to learn the intrinsic distribution characteristics of different data subspaces before training. Subsequently, the parameters of the LoRA expert are unfrozen, and the pre-trained model is trained. This allows the pre-trained model to achieve higher accuracy in the early stages of training, significantly improving training efficiency. This overcomes the long training time of traditional low-rank adaptation expert models and reduces the memory overhead required for training low-rank adaptation expert models.

[0042] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0043] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0044] Figure 1This is a hardware structure block diagram of a terminal based on an efficient fine-tuning method for multi-expert parameters using expert-perceptual differential initialization, according to an embodiment of this application.

[0045] Figure 2 This is a flowchart of an efficient fine-tuning method for multi-expert parameters based on expert-perceptual differential initialization according to an embodiment of this application;

[0046] Figure 3 This is a flowchart of a downward projection matrix mixing strategy according to an embodiment of this application;

[0047] Figure 4 This is a schematic diagram of an EAD-LoRA model preheating according to an embodiment of this application;

[0048] Figure 5 This is a structural block diagram of a multi-expert parameter high-efficiency fine-tuning device based on expert-perception differential initialization according to an embodiment of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0050] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0051] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application means two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0052] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of a terminal based on an efficient multi-expert parameter fine-tuning method using expert-perceptual differential initialization, according to an embodiment of this application. Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.

[0053] The memory 104 can be used to store computer programs, deploy the pre-trained model and computer program corresponding to the multi-expert parameter efficient fine-tuning method based on expert-perceptive differential initialization in this embodiment, and the processor 102 performs model training by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0054] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0055] based on Figure 1 The application scenario diagram shown illustrates that this embodiment provides an efficient fine-tuning method for multi-expert parameters based on expert-perceptual differential initialization. Figure 2 This is a flowchart of an efficient multi-expert parameter fine-tuning method based on expert-perceptual differential initialization according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0056] Step S210: Obtain the pre-trained model including the LoRA expert and the training data, and freeze the parameters of the LoRA expert in the pre-trained model.

[0057] The pre-trained model is a low-rank adapter-based expert model, comprising a backbone network and at least one LoRA (low-rank adapter) expert. The backbone network is the shared underlying network architecture in the pre-trained model. Each LoRA expert corresponds to an independent low-rank adapter, meaning each LoRA expert contains an independent low-rank decomposition matrix.

[0058] Training data is used for "warm-up" training of LoRA experts. The training data for LoRA experts can be selected from data that is consistent with or highly correlated with the fine-tuning task performed by the LoRA expert in terms of semantic domain, input / output format, or language style, so that the training data is adapted to the fine-tuning task performed by the LoRA expert. For example, the data distribution corresponding to the fine-tuning task to be performed by the LoRA expert is determined, and training data consistent with the data distribution of that fine-tuning task is obtained.

[0059] Step S220: Input the training data into the pre-trained model to obtain the corresponding hidden state vector generated by the pre-trained model based on the training data, as well as the LoRA expert corresponding to each training data, and construct the data sample set corresponding to the LoRA expert based on the hidden state vector.

[0060] Different LoRA experts have their corresponding data sample sets. Hidden state vectors are vectors obtained by encoding the training data into a pre-trained model. They are used to represent the semantic information, contextual dependencies, or other features related to the LoRA expert's fine-tuning task in the training data. Hidden state vectors include the input hidden state vector generated by the pre-trained backbone model encoding the training data, and the output hidden state vector generated by the LoRA expert based on the input hidden state.

[0061] Pre-trained models can route training data to corresponding LoRA experts through gating networks or other routing mechanisms, such as routing based on input keywords, task identifiers, or prompt templates, thereby obtaining the correspondence between training data and LoRA experts. Taking pre-trained models using gating networks for routing as an example, optionally, training data is input into the pre-trained model, and the backbone network of the pre-trained model encodes the training data, generating input hidden state vectors corresponding to each training data point; the gating network routes each training data point to the corresponding LoRA expert based on the input hidden state vectors; the LoRA expert receives the input hidden state vectors, and its output vector is the aforementioned output hidden state vector. The hidden state vectors of training data corresponding to the same LoRA expert are integrated to obtain a set of hidden state vectors corresponding to each LoRA expert, and this set of hidden state vectors is used as the data sample set for that LoRA expert. The specific implementation principles and methods of other model routing mechanisms can be found in the descriptions in related technologies, and are not limited here.

[0062] Step S230: Extract the variation pattern of the data sample set to obtain the feature vector corresponding to the LoRA expert, and obtain the down projection matrix of the corresponding LoRA expert based on the feature vector.

[0063] The change pattern refers to the evolutionary pattern or structural characteristic of the data in the data sample set along any ordered dimension. Optionally, the change pattern of the data sample set can be extracted using methods such as principal component analysis, variance calculation, and sequence analysis to obtain feature vectors. The LoRA projection matrix of the expert is initialized as a matrix composed of r feature vectors; where r is the pre-configured LoRA rank.

[0064] Step S240: Unfreeze the parameters of the LoRA expert and train the pre-trained model.

[0065] Specifically, LoRA technology is used for efficient fine-tuning of pre-trained model parameters. For example, the standard MoE-LoRA training process can be used for model fine-tuning. Specific implementation methods for model training using LoRA technology can be found in relevant technical documentation and will not be elaborated upon here.

[0066] In the aforementioned efficient fine-tuning method for multi-expert parameters based on expert-perceptual differential initialization, the routing mechanism of the pre-trained model is used to determine the LoRA expert corresponding to the hidden state vector of the training data, and a data sample set corresponding to the LoRA expert is constructed based on the hidden state vector. The change patterns of the data sample set are extracted, and the downprojection matrix of the LoRA expert is encoded into a feature vector to represent its hidden state change pattern. Thus, in the initial stage of training, the LoRA expert is guided to learn the intrinsic mechanism of different data subspaces through the downprojection matrix, realizing the functional differentiation of the LoRA expert. This enables the pre-trained model to achieve higher accuracy and faster loss reduction in the early stage of training, greatly improving training efficiency, overcoming the disadvantage of long training time of traditional low-rank fitting expert models, and reducing the memory overhead required for training low-rank fitting expert models.

[0067] Furthermore, since LoRA experts are aligned with the data distribution they will process from the beginning of training, the initial downprojection matrix can guide the direction of parameter optimization during model training. This allows the multi-expert parameter fine-tuning method based on expert-aware differential initialization to improve the accuracy of the model on downstream tasks, thereby enhancing model performance.

[0068] In some embodiments, extracting the variation patterns of the data sample set to obtain feature vectors corresponding to LoRA experts, and obtaining the downprojection matrix of the corresponding LoRA expert based on the feature vectors, includes: extracting the variation patterns of the data sample set to obtain the principal component matrix of the data sample set, the principal component matrix including multiple feature vectors corresponding to the LoRA expert; extracting a specified number of target feature vectors in the principal component matrix according to the eigenvalues ​​sorted from largest to smallest; the specified number is the same as the rank parameter of the LoRA expert; and constructing the downprojection matrix of the corresponding LoRA expert based on the specified number of target feature vectors.

[0069] Wherein, the principal component matrix encodes the variation patterns of the data sample set, that is, the principal component matrix is a set of basis vectors that mathematically represent these variation patterns. Each basis vector (i.e., eigenvector) in the principal component matrix corresponds to a variation direction. Optionally, principal component analysis (PCA) is performed separately on the data sample set of each LoRA expert to obtain the principal component matrix corresponding to each LoRA expert. Principal component analysis includes the steps of standardization, covariance calculation and eigen decomposition, and the specific implementation of principal component analysis can be found in the description of the related art, which will not be repeated herein.

[0070] The rank parameter of a LoRA expert is a preset hyperparameter, which is related to the expert's learning ability and parameter scale during fine-tuning, wherein r << the original matrix dimension of the LoRA expert. To balance the parameter efficiency and representation capability of experts, the rank parameter is typically set to r=4 or 8; specifically, the number of expert parameters is proportional to r, so r can be increased for complex tasks and decreased for simple tasks or scenarios with limited resources. Optionally, in the principal component matrix, the eigenvectors are sorted in descending order of eigenvalues; the first r eigenvectors of each LoRA expert are transposed to obtain the lower projection matrix A of the LoRA expert through initialization, and the upper projection matrix B of the LoRA expert remains zero-initialized.

[0071] In this embodiment, by extracting the matrix formed by the first r principal components from the data sample set through principal component analysis, eigenvectors that can reflect high-importance variation patterns in the sample data set can be obtained. The r eigenvectors used to characterize the data variation patterns are encoded into the lower projection matrix that can be directly used by the LoRA expert, which ensures that the adaptation capability of the LoRA expert is aligned with the most significant variation direction in its input data from the beginning, and realizes the precise alignment between the expert capability and the data distribution.

[0072] In order to further balance the data-driven capability and random exploration capability of each LoRA expert, in some embodiments, Figure 3 a flowchart of a lower projection matrix mixing strategy is provided, as Figure 3 shown, obtaining the lower projection matrix of the corresponding LoRA expert according to the eigenvectors further includes the following steps:

[0073] step S231, acquiring a first initialization matrix formed by eigenvectors and a randomly generated second initialization matrix.

[0074] Wherein, the dimension of the initialization matrix is determined according to the architecture of the pre-trained model and the preset LoRA rank r. For the lower projection matrix A of LoRA, its dimension is fixed as [r, d in , which is used as the initialization matrix acquired in this step.

[0075] Understandably, when training a pre-trained model, it is also necessary to initialize the projection matrix. For example, a matrix with dimension [d] can be selected. out Let the matrix B of [,r] be directly set as the initial state of the upward projection matrix. in It is the input dimension of the pre-trained model; d out It is the output dimension of the pre-trained model.

[0076] Step S232: Based on the mixing intensity parameter, fuse the first initialization matrix and the second initialization matrix to obtain the fusion matrix.

[0077] The mixing intensity parameter (denoted as α) is a hyperparameter used to control the fusion ratio of the initialization matrix and the matrix composed of eigenvectors. To balance the specialization and generalization ability of the projection matrix, systematic experiments such as grid search can be conducted to find the optimal α value for the validation set performance on a specific task, a specific backbone model, and a specific dataset, thereby achieving the fusion of the initialization matrix and the matrix composed of eigenvectors.

[0078] For example, the fusion matrix = fusion intensity parameter × first initialization matrix + (1 - fusion intensity parameter) × second initialization matrix.

[0079] Step S233: Add noise to the fusion matrix to obtain the lower projection matrix of the corresponding LoRA expert.

[0080] Optionally, a parameter-adjustable random perturbation, i.e., noise, can be introduced into the fusion matrix. For example, this noise can be a Gaussian distribution with a mean of 0 and a variance of β². It is understood that other random noise distributions can also be used as needed. By adding controllable noise to the fusion matrix, the robustness of the model trained based on the lower projection matrix can be improved.

[0081] In this embodiment, the first initialization matrix is ​​used to provide prior knowledge for data-driven approaches, while the second initialization matrix and noise are used to provide the model with the ability to explore new feature directions. By fusing the first initialization matrix with a random second initialization matrix and adding controllable noise, this hybrid initialization strategy can balance the prior knowledge of data-driven approaches with the model's ability to explore new feature directions, providing an adaptive optimization starting point for LoRA experts with different capabilities and downstream text maps.

[0082] After obtaining the downprojection matrix in the previous embodiment, the initialized downprojection matrix can be standardized. Specifically, after obtaining the downprojection matrix of the corresponding LoRA expert based on the feature vector, the method further includes: performing norm normalization on the downprojection matrix; determining the variance scaling factor based on the dimension of the neural network layer in the pre-trained model; and scaling the downprojection matrix based on the variance scaling factor.

[0083] Norm normalization aims to control the weight magnitude of the lower projection matrix by normalizing it. Optionally, the Frobenius norm of the lower projection matrix is ​​calculated, and the lower projection matrix is ​​divided by its corresponding norm to obtain the normalized output.

[0084] Variance scaling is used to dynamically adjust the variance of the projection matrix based on the input and output dimensions of the neural network layers in the pre-trained model, in order to match the forward and backward propagation requirements of the neural network. The variance scaling factor is used to scale the projection matrix. This factor can be set according to the dimensions of the neural network layers: a larger dimension requires a smaller variance scaling factor, and vice versa.

[0085] Optionally, the variance scaling method initialized by sampling Kaiming is set to √(2 / d). in Multiply by √(2 / d) based on the lower projection matrix. in This ensures the stability of the variance of activation values ​​during forward propagation. It prevents data from disappearing or exploding during neural network propagation and allows the downprojection matrix to match the neural network initialization principles.

[0086] In this embodiment, the initialized matrix is ​​standardized by norm normalization and variance scaling to ensure the stability of LoRA experts during the training process.

[0087] In some embodiments, the pre-trained model further includes a backbone network and a gating network. Training data is input into the pre-trained model to obtain hidden state vectors generated by the pre-trained model based on the training data, and LoRA experts for the routes corresponding to each training data point. A data sample set corresponding to the LoRA experts is constructed based on the hidden state vectors, including: dividing the training data into token sequences and inputting the token sequences into the backbone network to obtain the input hidden state vectors of each token output by the backbone network; obtaining the LoRA experts for the routes corresponding to each token based on the routing decisions of the gating network in the pre-trained model, and the output hidden state vectors output by the LoRA experts based on the input hidden states; and integrating the input and output hidden state vectors of each token according to their corresponding LoRA experts to obtain a data sample set for the LoRA experts.

[0088] Here, a token is the basic processing unit obtained by the pre-trained model after splitting the training data, and each token has its corresponding hidden state vector. For example, the parameters of all LoRA experts are frozen, but the gating network is activated; the training data is split into token sequences, and these sequences are input into the pre-trained model. The model performs forward propagation on the data, and the backbone network in the pre-trained model calculates the corresponding input hidden state based on the training data; the gating network routes different tokens to different LoRA experts based on the hidden state vectors. Further, for each LoRA expert, the input hidden state vectors and output hidden state vectors corresponding to the tokens routed to that expert are collected and cached, resulting in a data sample set corresponding to each LoRA expert.

[0089] In this embodiment, an intelligent sampling strategy based on gated routing can adaptively select the most suitable expert for the training data, thereby improving routing efficiency and routing accuracy.

[0090] Furthermore, in some embodiments, the input hidden state vectors and output hidden state vectors of each token are integrated according to the corresponding LoRA experts to obtain a data sample set of LoRA experts. This includes: obtaining the gate value corresponding to the LoRA expert calculated by the gating network in the pre-trained model based on the input hidden state vector of the token; obtaining the target token with a gate value greater than or equal to a preset threshold from multiple tokens; and integrating the input hidden state vectors and output hidden state vectors of the target token according to the corresponding LoRA experts to obtain a data sample set of LoRA experts.

[0091] The gating value is a quantitative indicator representing the weight by which a token is assigned to a specific LoRA expert for processing. If the gating value is greater than or equal to a preset threshold, the token is considered to have a high correlation with the LoRA expert; conversely, if the gating value is less than or equal to the preset threshold, the token is considered to have a low correlation with the LoRA expert, and the LoRA expert is not activated.

[0092] The preset threshold can be a fixed value set in advance; or, the preset threshold can also be generated based on the gate value corresponding to each token. For example, using the Top-K strategy, the gate values ​​of multiple tokens are sorted from largest to smallest, and the gate value of the Kth position is selected as the preset threshold, where K is the preset value; or, the Top-P strategy can be used to accumulate the gate values ​​of multiple tokens from largest to smallest until the cumulative probability exceeds the preset value P.

[0093] In this embodiment, the hidden state vectors of target tokens whose gate values ​​exceed the preset threshold are collected by comparing whether the gate value exceeds the preset threshold, so that the data sample set constructed based on the hidden state vectors can accurately reflect the data distribution of actual LoRA experts.

[0094] Furthermore, the efficient fine-tuning method for multi-expert parameters based on expert-aware differential initialization also includes: obtaining the frequency at which LoRA experts are activated by the target token; and adjusting the routing strategy of LoRA experts in the pre-trained model when the frequency is lower than a specified frequency.

[0095] Specifically, the LoRA expert is considered activated if the token's gating value is greater than or equal to a preset threshold, and if the token's gating value is less than the preset threshold. The activation frequency is determined based on the number of times an expert is activated by a target token. Optionally, the activation count can be directly used as the activation frequency; alternatively, the token-level weight distribution of each target token can be obtained, and the activation frequency can be calculated by weighting the activation count and the token-level weight distribution together.

[0096] If the frequency is lower than a preset frequency, the LoRA expert is deemed insufficiently activated. Routing strategies are used to increase the activation frequency of LoRA experts, thereby enriching their data sample set. Routing strategies can include adjusting the sampling retention ratio of hidden state vectors, optimizing gating network initialization parameters, and reassessing the number of experts. For example, an exemption sampling and full retention strategy can be used: experts with frequencies lower than the preset frequency are marked; for the marked experts, the hidden vectors of all tokens routed to that expert are integrated to obtain the LoRA expert's data sample set. Another example is a strategy of dynamically increasing the sampling budget: temporarily increasing the expert's budget. Yet another example is a weighted mixed sampling strategy: prioritizing the retention of a preset number of samples (hidden state vectors) with the highest gating value, and then supplementing with random samples to obtain a data sample set, thereby balancing sample representativeness and diversity.

[0097] For ease of understanding, an EAD-LoRA model incorporating multiple LoRA experts is provided. Figure 4 A schematic diagram of preheating the EAD-LoRA model, as shown below. Figure 4 As shown, the EAD-LoRA model includes an embedding layer, a multi-head attention module, residual connections and a layer normalization model, and a feedforward neural network layer. Training data is obtained to produce training question-and-answer data, which is then input into the pre-trained model, which outputs the answers to the test questions.

[0098] Specifically, in the MoE-LoRA layer (feedforward neural network layer) of the pre-trained model, the gating network routes each token in the training question-and-answer data to a different LoRA expert. Each LoRA expert is configured with its corresponding upper projection matrix B and lower projection matrix A. Specifically, expert 1 corresponds to upper projection matrix B1 and lower projection matrix A1; expert 2 corresponds to upper projection matrix B1 and lower projection matrix A1, and so on, with expert n corresponding to upper projection matrix Bn and lower projection matrix An.

[0099] During the warm-up phase, the hidden state vectors of each expert are collected. Specifically, each LoRA expert receives the input hidden state vector generated by the preceding network based on the tokens in the training data, and generates an output hidden state vector based on its preset training weights and the LoRA matrix. A gating value G and a preset threshold τ are obtained for each token. If G > τ, it is determined that the token activates expert i. The hidden state of each expert is collected: if the token activates expert i, a sample dataset is constructed based on the input and output hidden state vectors corresponding to that token, resulting in expert 1 samples, expert 2 samples, ..., expert n samples.

[0100] PCA analysis is performed on the samples of each expert, and the top r target feature vectors are extracted from the principal component matrix. The LoRA matrix of the corresponding expert is initialized based on the target feature vectors, thus obtaining the EAD-LoRA model.

[0101] Furthermore, a hybrid variant, the "EAD-LoRA+" model, can be constructed. Specifically, after initializing the LoRA matrix of the corresponding expert based on the target feature vector to obtain the first initialization matrix, a second initialization matrix is ​​generated through random initialization. The first and second initialization matrices are then fused according to a specified ratio, thus obtaining the "EAD-LoRA+" model.

[0102] Experiments have shown that EAD-LoRA and its hybrid variant “EAD-LoRA+” achieve an average accuracy improvement of 1.6% to 7.7% over the uniformly initialized MoE-LoRA on multiple question-answering benchmark tests, demonstrating their ability to optimize the upper limit of model performance.

[0103] In this embodiment, by identifying experts with extremely low activation frequency or insufficient data collection, their routing behavior is improved, ensuring the number of samples in the data sample set and improving the accuracy of the change patterns extracted based on the data sample set.

[0104] When implementing an efficient fine-tuning method for expert parameters based on expert-aware differential initialization, a sample buffer for LoRA experts can be set up in memory. After obtaining the data sample set of LoRA experts, the data sample set is stored in the corresponding sample buffer. However, when the pre-trained model is a large-scale model, there may be a problem of excessive storage consumption of the data sample set. Based on this, in some embodiments, the efficient fine-tuning method for expert parameters based on expert-aware differential initialization further includes: when the storage consumption of the LoRA expert data sample set exceeds a specified threshold, executing at least one of the following storage strategies: storing the data sample set of LoRA experts in the central processing unit; converting the data format of the LoRA expert data sample set to a specified format; and downsampling the LoRA expert data sample set.

[0105] Moving the hidden state from the GPU to the CPU can alleviate memory pressure. The specified format is one with low storage cost. For example, the data sample set of LoRA experts is converted to a float32 NumPy array to balance data storage precision and memory usage. For example, downsampling includes setting a sample limit for each expert; if the number of samples exceeds the corresponding limit, then random downsampling is performed according to a preset ratio.

[0106] Optionally, an expert hidden state collector is also included in the pre-trained model. The storage strategy described above is executed by the expert hidden state collector.

[0107] In this embodiment, a memory optimization strategy is adopted to efficiently manage the storage volume of each expert's data sample set, ensuring the feasibility of storing each expert's data sample set independently under a large-scale model.

[0108] This embodiment also provides a multi-expert parameter high-efficiency fine-tuning device based on expert-perceptual differential initialization. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0109] Figure 5 This is a structural block diagram of a multi-expert parameter high-efficiency fine-tuning device based on expert-perception differential initialization according to an embodiment of this application, such as... Figure 5 As shown, the multi-expert parameter high-efficiency fine-tuning device 500 based on expert-perception differential initialization includes:

[0110] The acquisition module 501 is used to acquire the pre-trained model including the LoRA expert and the training data, and freeze the parameters of the LoRA expert in the pre-trained model.

[0111] The sample construction module 502 is used to input training data into the pre-trained model, obtain the corresponding hidden state vector generated by the pre-trained model based on the training data, as well as the LoRA expert corresponding to each training data, and construct a data sample set corresponding to the LoRA expert based on the hidden state vector.

[0112] The matrix construction module 503 is used to extract the change patterns of the data sample set, obtain the feature vectors corresponding to the LoRA experts, and obtain the down projection matrix of the corresponding LoRA experts based on the feature vectors.

[0113] Training module 504 is used to unfreeze the parameters of the LoRA expert and train the pre-trained model.

[0114] In some embodiments, the sample construction module 502 extracts the variation patterns of the data sample set to obtain feature vectors corresponding to the LoRA expert, and obtains the downprojection matrix of the corresponding LoRA expert based on the feature vectors. This includes: extracting the variation patterns of the data sample set to obtain the principal component matrix of the data sample set, the principal component matrix including multiple feature vectors corresponding to the LoRA expert; extracting a specified number of target feature vectors in the principal component matrix according to the eigenvalues ​​sorted from largest to smallest, the specified number being the same as the rank parameter of the LoRA expert; and constructing the downprojection matrix of the corresponding LoRA expert based on the specified number of target feature vectors.

[0115] In some embodiments, the matrix construction module 503 obtains the downprojection matrix of the corresponding LoRA expert based on the feature vectors, including: obtaining a first initialization matrix composed of feature vectors and a randomly generated second initialization matrix; fusing the first initialization matrix and the second initialization matrix according to the mixing intensity parameter to obtain a fusion matrix; adding noise to the fusion matrix to obtain the downprojection matrix of the corresponding LoRA expert.

[0116] In some embodiments, the pre-trained model further includes a backbone network and a gating network. The sample construction module 502 inputs training data into the pre-trained model to obtain the hidden state vectors generated by the pre-trained model based on the training data, as well as the LoRA experts for the routes corresponding to each training data point. It then constructs a data sample set corresponding to the LoRA experts based on the hidden state vectors, including: dividing the training data into token sequences and inputting the token sequences into the backbone network to obtain the input hidden state vectors of each token output by the backbone network; obtaining the LoRA experts for the routes corresponding to each token and the output hidden state vectors output by the LoRA experts based on the input hidden state vectors, according to the routing decisions of the gating network in the pre-trained model; and integrating the input hidden state vectors and output hidden state vectors of each token according to the corresponding LoRA experts to obtain the data sample set of the LoRA experts.

[0117] Optionally, the sample construction module 502 integrates the input hidden state vectors and output hidden state vectors of each token according to the corresponding LoRA expert to obtain a data sample set of LoRA experts, including: obtaining the gate value corresponding to the LoRA expert calculated by the gating network in the pre-trained model based on the input hidden state vector of the token; obtaining the target token with a gate value greater than or equal to a preset threshold from multiple tokens; and integrating the input hidden state vectors and output hidden state vectors of the target token according to the corresponding LoRA expert to obtain a data sample set of LoRA experts.

[0118] Optionally, the sample construction module 502 is also used to obtain the frequency at which LoRA experts are activated by the target token; if the frequency is lower than a preset frequency, the routing strategy of the LoRA experts in the pre-trained model is adjusted.

[0119] In some embodiments, after obtaining the lower projection matrix of the corresponding LoRA expert based on the eigenvector, the matrix construction module 503 is further used to perform norm normalization on the lower projection matrix; obtain the variance scaling factor, and perform variance scaling on the lower projection matrix based on the variance scaling factor.

[0120] In some embodiments, the multi-expert parameter high-efficiency fine-tuning device based on expert-perception differential initialization further includes a sample management module. The sample management module is used to execute at least one of the following storage strategies when the storage occupied by the data sample set of LoRA experts is greater than a specified threshold: storing the data sample set of LoRA experts to the central processing unit; converting the data format of the data sample set of LoRA experts to a specified format; and downsampling the data sample set of LoRA experts.

[0121] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0122] This embodiment also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0123] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0124] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0125] S1, obtain the pre-trained model including the LoRA expert and the training data, and freeze the parameters of the LoRA expert in the pre-trained model.

[0126] S2, input the training data into the pre-trained model to obtain the corresponding hidden state vector generated by the pre-trained model based on the training data, as well as the LoRA expert corresponding to each training data, and construct the data sample set corresponding to the LoRA expert based on the hidden state vector.

[0127] S3, extract the variation pattern of the data sample set, obtain the feature vector corresponding to the LoRA expert, and obtain the down projection matrix of the corresponding LoRA expert based on the feature vector.

[0128] S4, unfreeze the parameters of the LoRA expert and train the pre-trained model.

[0129] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0130] Furthermore, in conjunction with the efficient multi-expert parameter fine-tuning method based on expert-perceptive differential initialization in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the efficient multi-expert parameter fine-tuning methods based on expert-perceptive differential initialization in the above embodiments.

[0131] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0132] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0133] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for efficient fine-tuning of multi-expert parameters based on expert-perceptual differential initialization, characterized in that, The method includes: Obtain a pre-trained model including LoRA experts and training data, and freeze the parameters of the LoRA experts in the pre-trained model; The training data is input into the pre-trained model to obtain the corresponding hidden state vector generated by the pre-trained model based on the training data, as well as the LoRA expert corresponding to each training data, and a data sample set corresponding to the LoRA expert is constructed based on the hidden state vector. Extract the variation patterns of the data sample set to obtain the feature vectors corresponding to the LoRA expert, and obtain the down projection matrix of the corresponding LoRA expert based on the feature vectors; Unfreeze the parameters of the LoRA expert and train the pre-trained model.

2. The method according to claim 1, characterized in that, The step of extracting the change patterns of the data sample set to obtain the feature vector corresponding to the LoRA expert, and obtaining the downprojection matrix of the corresponding LoRA expert based on the feature vector, includes: Extract the variation patterns of the data sample set to obtain the principal component matrix of the data sample set, wherein the principal component matrix includes multiple feature vectors corresponding to LoRA experts; In the principal component matrix, a specified number of target feature vectors are extracted according to the eigenvalues ​​sorted from largest to smallest, and the specified number is the same as the rank parameter of the LoRA expert; Construct the lower projection matrix of the corresponding LoRA expert based on a specified number of target feature vectors.

3. The method according to claim 1 or 2, characterized in that, The step of obtaining the lower projection matrix of the corresponding LoRA expert based on the feature vector includes: Obtain a first initialization matrix composed of the feature vectors, and a randomly generated second initialization matrix; Based on the mixing intensity parameter, the first initialization matrix and the second initialization matrix are fused to obtain the fusion matrix; Noise is added to the fusion matrix to obtain the lower projection matrix of the corresponding LoRA expert.

4. The method according to claim 1, characterized in that, The pre-trained model further includes a backbone network and a gating network. The training data is input into the pre-trained model to obtain the corresponding hidden state vectors generated by the pre-trained model based on the training data, and the LoRA experts for each route corresponding to the training data. A data sample set corresponding to the LoRA experts is constructed based on the hidden state vectors, including: The training data is divided into token sequences, and the token sequences are input into the backbone network to obtain the input hidden state vector of each token output by the backbone network. Based on the routing decisions of the gated network for each token in the pre-trained model, the LoRA expert for the route corresponding to the token and the output hidden state vector output by the LoRA expert based on the input hidden state vector are obtained. The input hidden state vectors and output hidden state vectors of each token are integrated according to the corresponding LoRA expert to obtain the data sample set of the LoRA expert.

5. The method according to claim 4, characterized in that, The step involves integrating the input hidden state vectors and output hidden state vectors of each token according to their corresponding LoRA experts to obtain the data sample set of the LoRA experts, including: Obtain the gating value corresponding to the LoRA expert, calculated by the gating network in the pre-trained model based on the input hidden state vector of the token; Among the multiple tokens, obtain the target token whose gating value is greater than or equal to a preset threshold; The input hidden state vector and output hidden state vector of the target token are integrated according to the corresponding LoRA expert to obtain the data sample set of the LoRA expert.

6. The method according to claim 5, characterized in that, The method further includes: Obtain the frequency at which the LoRA expert is activated by the target token; If the frequency is lower than the preset frequency, adjust the routing strategy of the LoRA expert in the pre-trained model.

7. The method according to claim 1, characterized in that, After obtaining the lower projection matrix of the corresponding LoRA expert based on the eigenvectors, the method further includes: The lower projection matrix is ​​normalized using norm normalization. Obtain the variance scaling factor, and perform variance scaling processing on the lower projection matrix according to the variance scaling factor.

8. The method according to claim 1, characterized in that, The method further includes: If the storage usage of the LoRA expert's data sample set exceeds a specified threshold, execute at least one of the following storage strategies: store the LoRA expert's data sample set in the central processing unit; convert the data format of the LoRA expert's data sample set to a specified format; and downsample the LoRA expert's data sample set.

9. A multi-expert parameter high-efficiency fine-tuning device based on expert-perception differential initialization, characterized in that, include: The acquisition module is used to acquire a pre-trained model including LoRA experts and training data, and freeze the parameters of the LoRA experts in the pre-trained model. The sample construction module is used to input the training data into the pre-trained model, obtain the corresponding hidden state vector generated by the pre-trained model based on the training data, and the LoRA expert corresponding to each training data route, and construct a data sample set corresponding to the LoRA expert based on the hidden state vector; The matrix construction module is used to extract the change patterns of the data sample set, obtain the feature vectors corresponding to the LoRA expert, and obtain the down projection matrix of the corresponding LoRA expert based on the feature vectors. The training module is used to unfreeze the parameters of the LoRA expert and train the pre-trained model.

10. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute, at runtime, the efficient fine-tuning method for multi-expert parameters based on expert-perceptual differential initialization as described in any one of claims 1 to 8.