Asymmetric low-rank fine tuning model training method, device, equipment, medium and product

By constructing a multi-task dataset and utilizing an asymmetric low-rank adapter architecture, the problem of the difficulty in determining the number of B matrices in the HydraLoRA architecture in continuous modal multi-task classification scenarios is solved, the model is effectively applied in multimodal scenarios such as images and point clouds, and the performance of multi-task classification is improved.

CN120671765APending Publication Date: 2025-09-19CHINA MOBILE GROUP ZHEJIANG +3
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510624548.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The HydraLoRA architecture is difficult to be effectively applied in continuous modality multi-task classification scenarios. In particular, how to reasonably determine the number of B matrices required is an urgent problem to be solved.

Method used

Construct a multi-task dataset and use it to pre-train the initial model. Adopt an asymmetric low-rank adapter architecture and perform fine-tuning using a first matrix and multiple second matrices. The number of second matrices is determined based on the statistical vector of each sample data. Self-distillation loss and inter-task balance loss are used to improve training stability.

Benefits of technology

On the basis of reasonably determining the number of the second matrix, the application scenarios of the asymmetric low-rank adapter architecture are expanded, and the performance of the model in continuous modality multi-task classification scenarios is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671765A_ABST
    Figure CN120671765A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an asymmetric low-rank fine tuning model training method and device, equipment, a medium and a product. The method comprises the following steps: constructing a multi-task data set; the multi-task data set comprises a plurality of task classification data sets, and each task classification data set comprises sample data of a plurality of continuous modes; based on the multi-task data set, pre-training the initial model to obtain an asymmetric low-rank fine tuning model; the asymmetric low-rank fine tuning model comprises an asymmetric low-rank adapter, the asymmetric low-rank adapter is fine-tuned based on a first matrix and a plurality of second matrixes, and the number of the second matrixes is determined based on a statistical vector of each piece of sample data. By means of the mode, the problem that the number of the second matrixes needing to be used by the model is difficult to determine is solved, the model adopting the asymmetric low-rank adapter architecture is applied to a multi-task classification scene of a continuous mode, and the application scene of the asymmetric low-rank adapter architecture is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an asymmetric low-rank fine-tuning model training method, apparatus, equipment, medium and product. Background Art

[0002] LoRA (Low-Rank Adaptation) technology is an advanced model fine-tuning technology that optimizes the performance of the model on specific tasks by introducing a small number of adjustable parameters while maintaining the original structure and knowledge of the model and reducing computing resource consumption.

[0003] However, models using a single LoRA architecture still have problems in multi-task classification scenarios. They are difficult to meet the needs of multi-task classification scenarios because there may be crosstalk between data from different tasks. That is, the optimal gradient direction of one task may have a negative impact on the convergence of another task, resulting in the model's multi-task classification training failing to reach the global optimal solution. In multi-task classification scenarios, it is also possible to configure a separate LoRA for each task, but if a separate LoRA is configured for each task, too many additional parameters will be introduced into the model, and this design usually requires the introduction of more downstream data, which is not conducive to fine-tuning the model.

[0004] Based on this, some related technologies have proposed the HydraLoRA (Asymmetric Low Rank Adapter) architecture. HydraLoRA uses only one A matrix and multiple B matrices for fine-tuning. This design allows the model performance to exceed that of directly using multiple LoRAs for fine-tuning. This is because the principle of LoRA is to use one A matrix and one B matrix for fine-tuning. Directly using multiple LoRAs means that multiple A matrices and multiple B matrices are required, which introduces too many additional parameters into the model and is not conducive to model fine-tuning training. Among them, the A matrix focuses on extracting shared information of the task, while the B matrix focuses on extracting task-specific features.

[0005] At present, the effectiveness of the HydraLoRA architecture has been verified in the training of natural language tasks, but the HydraLoRA architecture is difficult to be effectively applied in continuous modality multi-task classification scenarios. For example, in multi-task classification scenarios, how to reasonably determine the number of B matrices required is also an urgent problem to be solved. Summary of the Invention

[0006] The embodiments of the present application provide an asymmetric low-rank fine-tuning model training method, device, equipment, medium and product to solve the technical problem that the existing HydraLoRA architecture is difficult to be effectively applied in continuous modal multi-task classification scenarios.

[0007] In the first aspect, an embodiment of the present application provides an asymmetric low-rank fine-tuning model training method, including: constructing a multi-task dataset; the multi-task dataset includes multiple task classification datasets, and each task classification dataset includes sample data of multiple continuous modes; based on the multi-task dataset, the initial model is pre-trained to obtain an asymmetric low-rank fine-tuning model; wherein, the asymmetric low-rank fine-tuning model includes an asymmetric low-rank adapter, and the asymmetric low-rank adapter is fine-tuned based on a first matrix and multiple second matrices, the first matrix is ​​used to extract shared features of multiple task classification datasets, and the second matrix is ​​used to extract specific task features of multiple task classification datasets, and the number of second matrices is determined based on the statistical vector of each sample data.

[0008] In one embodiment, the initial model is pre-trained based on a multi-task data set, and before obtaining an asymmetric low-rank fine-tuning model, it also includes: discretizing each sample data to obtain multiple sub-vectors corresponding to each sample data; determining the statistical vector of each sample data based on each sub-vector; clustering all statistical vectors to obtain multiple candidate quantity values; determining the variance ratio index of each candidate quantity value, and selecting a target quantity value from the multiple candidate quantity values; wherein the target quantity value is the minimum value among the candidate quantity values ​​whose variance ratio index is greater than a preset variance ratio threshold, and the target quantity value is the quantity of the second matrix.

[0009] In one embodiment, each sample data is discretized to obtain multiple sub-vectors corresponding to each sample data, including: segmenting each sample data to obtain multiple sub-blocks corresponding to each sample data; projecting each sub-block into a vector to obtain multiple sub-vectors corresponding to each sample data; wherein a sub-vector represents the probability of a sub-block being assigned to different visual words.

[0010] In one embodiment, the initial model includes a teacher model and a student model, and the teacher model and the student model have the same model weights when not pre-trained; based on a multi-task dataset, the initial model is pre-trained to obtain an asymmetric low-rank fine-tuning model, including: selecting the same sample data from the multi-task dataset as fine-tuning samples, inputting them into the teacher model and the student model respectively, and performing pre-training to obtain an asymmetric low-rank fine-tuning model; wherein the asymmetric low-rank fine-tuning model is the student model that has completed pre-training; during the pre-training process, the teacher model does not perform gradient updates, the student model performs gradient updates, and the model weights of the teacher model are updated through an exponential sliding average strategy, and gradually approach the model weights of the student model.

[0011] In one embodiment, the student model is pre-trained using a preset loss function, which includes a self-distillation loss and an inter-task balance loss; wherein the self-distillation loss is used to constrain the modules in the student model that are updated by the exponential sliding average strategy; the inter-task balance loss is used to constrain all second matrices in the student model, and the inter-task balance loss is determined based on the routing probability of each second matrix in the student model.

[0012] In one embodiment, constructing a multi-task dataset includes: obtaining multiple initial task classification datasets; each initial task classification dataset includes sample data of multiple continuous modes, and one sample data includes a sample classification image and a sample classification label corresponding to the sample classification image; adding a corresponding task classification dataset ID to each sample data to obtain multiple task classification datasets; arranging each sample data in each task classification dataset in sequence, and mixing all the sample data through randomly generated indexes to generate a multi-task dataset.

[0013] In the second aspect, an embodiment of the present application provides an asymmetric low-rank fine-tuning model training device, including: a construction module for constructing a multi-task dataset; the multi-task dataset includes multiple task classification datasets, and each task classification dataset includes sample data of multiple continuous modes; a training module for pre-training the initial model based on the multi-task dataset to obtain an asymmetric low-rank fine-tuning model; wherein the asymmetric low-rank fine-tuning model includes an asymmetric low-rank adapter, and the asymmetric low-rank adapter is fine-tuned based on a first matrix and multiple second matrices, the first matrix is ​​used to extract shared features of multiple task classification datasets, and the second matrix is ​​used to extract specific task features of multiple task classification datasets, and the number of second matrices is determined based on the statistical vector of each sample data.

[0014] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements any of the above-mentioned asymmetric low-rank fine-tuning model training methods.

[0015] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements any of the above-mentioned asymmetric low-rank fine-tuning model training methods.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned asymmetric low-rank fine-tuning model training methods.

[0017] The asymmetric low-rank fine-tuning model training method, apparatus, equipment, medium and product provided in the embodiments of the present application first construct a multi-task dataset, which includes sample data of multiple continuous modes of multiple task classification datasets, and uses the multi-task dataset to pre-train the initial model to obtain an asymmetric low-rank fine-tuning model, wherein the asymmetric low-rank fine-tuning model adopts an asymmetric low-rank adapter architecture, and the asymmetric low-rank adapter uses a first matrix and multiple second matrices for fine-tuning. The first matrix is ​​used to extract shared features of multiple task classification datasets, and the second matrix is ​​used to extract specific task features of multiple task classification datasets, and the number of second matrices can be determined according to the statistical vector of each sample data. In the above manner, the number of second matrices is determined by using the statistical vector of each sample data in the multi-task dataset, which can solve the problem that the number of second matrices required for the model is difficult to determine. At the same time, on the basis of reasonably determining the number of second matrices, the model using the asymmetric low-rank adapter architecture is applied to multi-task classification scenarios of continuous modes other than natural language tasks, expanding the application scenarios of the asymmetric low-rank adapter architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 It is a flow chart of the asymmetric low-rank fine-tuning model training method provided in an embodiment of the present application.

[0020] Figure 2 This is a schematic diagram of the principle of the LoRA architecture provided in the embodiment of the present application.

[0021] Figure 3 It is a schematic diagram of the principle of the HydraLoRA architecture provided in an embodiment of the present application.

[0022] Figure 4 This is a model architecture diagram of the asymmetric low-rank fine-tuning model training method provided in an embodiment of the present application.

[0023] Figure 5 It is a structural diagram of the asymmetric low-rank fine-tuning model training device provided in an embodiment of the present application.

[0024] Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0026] See also Figure 1 , Figure 1 This is a flow chart of the asymmetric low-rank fine-tuning model training method provided in the embodiment of the present application. Figure 1 As shown, in the embodiment of the present application, the asymmetric low-rank fine-tuning model training method includes steps S110 to S120, and each step is specifically as follows: S110: Build a multi-task dataset.

[0027] The multi-task dataset includes multiple task classification datasets, and each task classification dataset includes sample data of multiple continuous modes.

[0028] S120: Based on the multi-task dataset, pre-train the initial model to obtain an asymmetric low-rank fine-tuning model.

[0029] Among them, the asymmetric low-rank fine-tuning model includes an asymmetric low-rank adapter, which is fine-tuned based on a first matrix and multiple second matrices. The first matrix is ​​used to extract shared features of multiple task classification data sets, and the second matrix is ​​used to extract specific task features of multiple task classification data sets. The number of second matrices is determined based on the statistical vector of each sample data.

[0030] In order to better understand the technical solutions of the embodiments of this application, Figure 2 and Figure 3 See Figure 2 and Figure 3 , Figure 2 This is a schematic diagram of the principle of the LoRA architecture provided in the embodiment of the present application. Figure 3 It is a schematic diagram of the principle of the HydraLoRA architecture provided in an embodiment of the present application.

[0031] like Figure 2 As shown in the figure, the principle of LoRA is to use an A matrix and a B matrix for fine-tuning. That is, in the LoRA (Low-Rank Adaptation) architecture, the weights of the model's incremental update are expressed as the product of two low-rank matrices (i.e., the A matrix and the B matrix), which can significantly reduce the number of parameters that need to be trained in the model. The fine-tuning process of LoRA can be expressed by the following formula: ; in, Represents the weight of the model pre-training; Represents the weight updated during the model fine-tuning process; Represents the final trained weight of the model (obtained by adding the pre-trained weights and the weights updated during fine-tuning), that is, the weight used for final inference.

[0032] In order to improve the efficiency of model parameter training, LoRA does not directly train the weights updated in the fine-tuning process, but uses two low-rank matrices (i.e., A matrix and B matrix) to approximate their values, then: ; in, In essence, it is a size The matrix can be decomposed into The A matrix and size are This approach maintains model performance while reducing the model's requirements for computing and storage resources, making the model more lightweight and suitable for resource-constrained environments. This technical advantage also expands LoRA's application range, including but not limited to natural language processing (NLP), where LoRA is often used for lightweight fine-tuning of large models to adapt to specific tasks. Furthermore, LoRA has been expanded to include multiple variants to further optimize its performance, making it more adaptable and efficient in different application scenarios.

[0033] However, models using a single LoRA architecture still have problems in multi-task classification scenarios. They are difficult to meet the needs of multi-task classification scenarios because there may be crosstalk between data from different tasks. That is, the optimal gradient direction of one task may have a negative impact on the convergence of another task, resulting in the model's multi-task classification training failing to reach the global optimal solution. In multi-task classification scenarios, it is also possible to configure a separate LoRA for each task. However, if a separate LoRA is configured for each task, too many additional parameters will be introduced into the model, and this design usually requires the introduction of more downstream data, which is not conducive to fine-tuning the model. Therefore, the existing technology proposes an asymmetric low-rank adapter (HydraLoRA) architecture.

[0034] like Figure 3 As shown, HydraLoRA uses only one A matrix and multiple B matrices ( Figure 3 Use B1 to B K This design enables the model performance to exceed the performance of directly using multiple LoRA for fine-tuning.

[0035] It should be noted that the HydraLoRA architecture also refers to the advantages of the MoE architecture (Mixture of Experts) and makes weighted selections for different B matrices. This design is adopted because research has found that the A matrix in LoRA focuses on extracting shared information of the task, while the B matrix focuses on extracting specific task features. Therefore, the processing flow of the HydraLoRA architecture can be expressed as follows: ; in, Represents input data, Represents output data, is the initial weight; each B matrix is ​​equivalent to an expert in the MoE architecture, and the The routing weight of the expert (i.e. B matrix Routing probability ) can be expressed by the following formula: ; in, It is a simple linear layer used to complete the projection of features; softmax represents the activation function.

[0036] HydraLoRA has achieved excellent performance in multi-task training of natural language, but it is difficult to be effectively applied in multi-task classification scenarios of continuous modalities. For example, in multi-task classification scenarios, how to reasonably determine the number of B matrices required is also an urgent problem to be solved.

[0037] Based on this, the embodiments of the present application improve the existing model training method to enhance model performance, and apply the model using the HydraLoRA architecture to multi-task classification scenarios of multi-modal or continuous modalities such as images and point clouds.

[0038] Specifically, a multi-task dataset is first constructed. The multi-task dataset needs to include data from multiple task classification datasets, and each task classification dataset includes sample data from multiple continuous modes.

[0039] Furthermore, the initial model is pre-trained based on the multi-task dataset to obtain an asymmetric low-rank fine-tuning model.

[0040] Among them, the asymmetric low-rank fine-tuning model adopts the asymmetric low-rank adapter (HydraLoRA) architecture. The asymmetric low-rank adapter uses a first matrix (i.e., A matrix) and multiple second matrices (i.e., B matrices) for fine-tuning training. The number of second matrices (i.e., B matrices) can be determined based on the statistical vector of each sample data.

[0041] The asymmetric low-rank fine-tuning model training method provided in the embodiment of the present application first constructs a multi-task dataset, which includes sample data of multiple continuous modes of multiple task classification datasets, and uses the multi-task dataset to pre-train the initial model to obtain an asymmetric low-rank fine-tuning model, wherein the asymmetric low-rank fine-tuning model adopts an asymmetric low-rank adapter architecture, and the asymmetric low-rank adapter uses a first matrix and multiple second matrices for fine-tuning. The first matrix is ​​used to extract shared features of multiple task classification datasets, and the second matrix is ​​used to extract specific task features of multiple task classification datasets, and the number of second matrices can be determined according to the statistical vector of each sample data. In the above manner, the number of second matrices is determined by using the statistical vector of each sample data in the multi-task dataset, which can solve the problem that the number of second matrices required for the model is difficult to determine. At the same time, on the basis of reasonably determining the number of second matrices, the model using the asymmetric low-rank adapter architecture is applied to multi-task classification scenarios of continuous modes other than natural language tasks, expanding the application scenarios of the asymmetric low-rank adapter architecture.

[0042] In some embodiments, the initial model is pre-trained based on a multi-task data set, and before obtaining an asymmetric low-rank fine-tuning model, it also includes: discretizing each sample data to obtain multiple sub-vectors corresponding to each sample data; determining the statistical vector of each sample data based on each sub-vector; clustering all statistical vectors to obtain multiple candidate quantity values; determining the variance ratio index of each candidate quantity value, and selecting a target quantity value from the multiple candidate quantity values; wherein the target quantity value is the minimum value among the candidate quantity values ​​whose variance ratio index is greater than a preset variance ratio threshold, and the target quantity value is the quantity of the second matrix.

[0043] In the embodiment of the present application, the asymmetric structure and weight initialization strategy of HydraLoRA are selected.

[0044] Specifically, HydraLoRA is used in the LoRA layer of the initial model, and the HydraLoRA in each LoRA layer uses a size of The first matrix (i.e. A matrix) and K dimensions are The second matrix (i.e., B matrix) is fine-tuned for training, where K is a positive integer. Indicates the input dimension of the current LoRA layer, Indicates the output dimension of the current LoRA layer, Indicates the rank of the current LoRA layer. In addition, the initial model has a size of Matrix Routing for the second matrix (i.e., B matrix).

[0045] Optionally, for the weight initialization strategy, the first matrix (i.e., A matrix) and are randomly initialized from the standard normal distribution, and the K second matrices (i.e., B matrices) are initialized with zero.

[0046] It should be noted that since HydraLoRA is proposed for natural language tasks, TF-IDF (Term Frequency–Inverse Document Frequency) can be used to measure the similarity between sample data of discrete modalities, and the K-means clustering algorithm can be further used to determine the optimal K value selection. In multi-task classification scenarios, sample data is usually continuous modal, for example, sample data of image modality and point cloud modality are both continuous modal data. Based on this, the embodiment of the present application proposes an idea for selecting K values ​​on continuous modal data to determine the number of the second matrix (i.e., the B matrix).

[0047] First, the encoder of the DALL-E model can be used to discretize each sample data to obtain multiple sub-vectors corresponding to each sample data.

[0048] Specifically, assuming that the total number of sample data in the multi-task dataset is M; for each sample data, the sample data is segmented, and the sample data can be segmented into corresponding sub-blocks. At this time, each sub-block can be projected into a vector through the encoder of the DALL-E model to obtain multiple sub-vectors corresponding to the sample data.

[0049] Among them, a sub-vector represents the probability of a sub-block being assigned to different visual words.

[0050] For example, The sample data is divided into corresponding sub-blocks, the encoder of the DALL-E model can The sub-blocks are projected into a sub-vector with a dimension of 8192 , assuming that there are 8192 visual words in the visual codebook, the sub-vector represents the probability of the corresponding sub-block being assigned to each visual word; after projecting all sub-blocks into vectors, multiple sub-vectors corresponding to the sample data can be obtained.

[0051] Furthermore, based on each sub-vector, a statistical vector of each sample data is determined.

[0052] Specifically, for each sample data, a statistical vector of the sample data can be determined according to its corresponding multiple sub-vectors.

[0053] Among them, The statistical vector of sample data The calculation formula is as follows: ; in, Represents a high-dimensional vector (e.g., 8192 dimensions), which can represent the distribution of the entire sample data on visual words; ” means element-wise multiplication.

[0054] Furthermore, all statistical vectors are clustered to obtain multiple candidate quantity values.

[0055] Specifically, assuming that the total number of sample data in the multi-task dataset is M, after the above steps, M statistical vectors can be obtained. At this time, K-means clustering is performed on the M statistical vectors to obtain multiple candidate quantity values, that is, multiple candidate K values.

[0056] Further, the variance ratio index of each candidate quantity value is determined, and a target quantity value is selected from multiple candidate quantity values; wherein the target quantity value is the minimum value among the candidate quantity values ​​whose variance ratio index is greater than a preset variance ratio threshold, and the target quantity value is the quantity of the second matrix.

[0057] In this embodiment, the Calinski-Harabasz Index (CHI, also known as the variance ratio index or the Calinski-Harabasz index) is used to quantify the clustering effect. A larger variance ratio index indicates a better clustering effect.

[0058] Among them, the calculation formula of the variance ratio index of the candidate K value is as follows: ; in, Indicates the The number of statistical vectors in a cluster; Indicates the The cluster centers of the clusters; Represents the mean of M statistical vectors.

[0059] Substituting each candidate K value into the above formula, the variance ratio index of each candidate K value (i.e. value).

[0060] At this time, first select from all candidate K values The candidate K value whose value is greater than the preset variance ratio threshold is selected from The candidate K values ​​whose values ​​are greater than the preset variance ratio threshold are selected The candidate K value with the smallest value is used as the target quantity value (ie, the target K value), and the target quantity value is the quantity of the second matrix (ie, the B matrix).

[0061] It can be understood that the above-mentioned processing ideas for sample data of continuous modality are not limited to sample data of image modality. For example, the point cloud sample data can be discretized with the help of a dVAE model trained on sample data of point cloud modality, and then the above-mentioned process can be repeated.

[0062] In some embodiments, each sample data is discretized to obtain multiple sub-vectors corresponding to each sample data, including: segmenting each sample data to obtain multiple sub-blocks corresponding to each sample data; projecting each sub-block into a vector to obtain multiple sub-vectors corresponding to each sample data; wherein a sub-vector represents the probability of a sub-block being assigned to different visual words.

[0063] Specifically, the encoder of the DALL-E model may be used to discretize each sample data to obtain multiple sub-vectors corresponding to each sample data.

[0064] Taking the sample data of image modality as an example, assuming that the total number of sample classification images of the multi-task dataset is M; for each sample classification image, the sample classification image is segmented, and the sample classification image can be segmented into corresponding image sub-blocks. At this time, each image sub-block can be projected into a vector through the encoder of the DALL-E model to obtain multiple sub-vectors corresponding to the sample classification image.

[0065] Among them, a sub-vector represents the probability of an image sub-block being assigned to different visual words.

[0066] For example, The sample classification image is segmented into corresponding image sub-blocks, the encoder of the DALL-E model can convert the The image sub-block is projected into a sub-vector with a dimension of 8192 , assuming that there are 8192 visual words in the visual codebook, the sub-vector represents the probability of the corresponding image sub-block being assigned to each visual word; after projecting all image sub-blocks into vectors, multiple sub-vectors corresponding to the sample classification image can be obtained.

[0067] In some embodiments, the initial model includes a teacher model and a student model, and the teacher model and the student model have the same model weights when not pre-trained; based on a multi-task dataset, the initial model is pre-trained to obtain an asymmetric low-rank fine-tuning model, including: selecting the same sample data from the multi-task dataset as fine-tuning samples, inputting them into the teacher model and the student model respectively, and performing pre-training to obtain an asymmetric low-rank fine-tuning model; wherein the asymmetric low-rank fine-tuning model is a student model that has completed pre-training; during the pre-training process, the teacher model does not perform gradient updates, the student model performs gradient updates, and the model weights of the teacher model are updated through an exponential sliding average strategy, and gradually approach the model weights of the student model.

[0068] Because the first matrix (i.e., the A matrix) is used to extract shared features from different task classification datasets, and during pre-training, sample data from different task classification datasets is interleaved and fed into the initial model for training, the training gradient of the initial model will experience significant fluctuations during pre-training, hindering the initial model's stable and rapid convergence to achieve effective shared feature extraction. Based on this, the present embodiment introduces a self-distillation strategy for the training of the first matrix (i.e., the A matrix).

[0069] See also Figure 4 , Figure 4 This is a model architecture diagram of the asymmetric low-rank fine-tuning model training method provided in an embodiment of the present application.

[0070] Specifically, the initial model includes a teacher model and a student model. The teacher model and the student model have the same model weights when not pre-trained, that is, the initial values ​​of the teacher model and the student model are exactly the same.

[0071] Further, if Figure 4 As shown in the figure, the same sample data is continuously selected from the multi-task dataset as the fine-tuning sample, and is input into the teacher model at the T-1th time step and the student model at the Tth time step for pre-training to obtain an asymmetric low-rank fine-tuning model. The asymmetric low-rank fine-tuning model is the student model that completes the pre-training.

[0072] It should be noted that during the pre-training process, the teacher model does not perform gradient updates, while the student model performs normal pre-training. The model weights of the teacher model are updated using the exponential moving average (EMA) strategy in the following way, gradually approaching the model weights of the student model: ; in, Represents the model weight of the teacher model at the Tth time step (including all trainable modules such as the A matrix and the B matrix, and can also flexibly define modules that need to use the EMA strategy), Represents the model weight of the teacher model at the T-1th time step (including all trainable modules such as the A matrix and the B matrix, and can also flexibly define modules that need to use the EMA strategy); represents the model weight of the student model at the Tth time step; It is a momentum parameter. By changing its value, the dependence of the teacher model update process on the past time can be controlled. Optionally, a scheduling strategy can be adopted to control The value of changes dynamically during the fine-tuning process.

[0073] In some embodiments, the student model is pre-trained using a preset loss function, which includes a self-distillation loss and an inter-task balance loss; wherein the self-distillation loss is used to constrain the modules in the student model that are updated by the exponential sliding average strategy; the inter-task balance loss is used to constrain all second matrices in the student model, and the inter-task balance loss is determined based on the routing probability of each second matrix in the student model.

[0074] In this embodiment, the student model is pre-trained using a preset loss function. In order to constrain the training of the student model, the embodiment of the present application introduces self-distillation loss and inter-task balance loss into the preset loss function.

[0075] Self-distillation loss is used to constrain the module in the student model that is updated by the exponential sliding average strategy. The calculation formula is as follows: ; in, represents the output of the student model during training; represents the output of the teacher model during training; Represents the smoothed L1 loss function, which is a regression task loss function that combines the advantages of L1 loss (absolute error) and L2 loss (square error).

[0076] Optionally, any module in the student model that uses the EMA strategy can be constrained with self-distillation loss, such as the A matrix, B matrix, or other modules in the student model.

[0077] Furthermore, in order to fully learn the specific task features of different tasks, the embodiment of the present application introduces an inter-task balance loss. The inter-task balance loss is used to constrain the learning process of all second matrices (i.e., B matrices) in the student model. The inter-task balance loss is determined based on the routing probability of each second matrix (i.e., B matrix) in the student model. The calculation formula is as follows: ; in, and are predefined hyperparameters; is the number of B matrices in the student model; is a natural constant; , Indicates the The routing probability of a B matrix, The calculation method is The number of times a B matrix is ​​selected is divided by the total number of routing times; express The maximum value in ; express The second-to-last value in .

[0078] In some embodiments, constructing a multi-task dataset includes: obtaining multiple initial task classification datasets; each initial task classification dataset includes sample data of multiple continuous modalities, and one sample data includes a sample classification image and a sample classification label corresponding to the sample classification image; adding a corresponding task classification dataset ID to each sample data to obtain multiple task classification datasets; arranging each sample data in each task classification dataset in sequence, and mixing all sample data through randomly generated indexes to generate a multi-task dataset.

[0079] A multi-task dataset is a dataset that contains sample data from multiple tasks, typically merging multiple task classification datasets. To construct a multi-task dataset, in addition to the necessary sample classification images and their corresponding sample classification labels, task information is also required to indicate the source of the data for subsequent training.

[0080] Specifically, first obtain T (T is a positive integer) initial task classification data sets, which are recorded as , unify the format of classification datasets for each initial task.

[0081] For the Initial task classification dataset , which includes multiple continuous modal sample data, one sample data includes a sample classification image and a sample classification label corresponding to the sample classification image, the sample data Can be recorded as , represents the sample classification image, Indicates the sample classification label corresponding to the sample classification image.

[0082] Furthermore, a corresponding task classification data set ID is added to each sample data to obtain multiple task classification data sets.

[0083] Specifically, a corresponding task classification dataset ID is added to each sample data of each initial task classification dataset to identify the data source, that is, the initial task classification dataset The Sample data Add the corresponding task classification dataset ID, and the new data is represented as ,in It is The ID of a task classification dataset. The ID of a task classification dataset is usually the name of the task classification dataset.

[0084] Furthermore, each sample data in each task classification dataset is arranged in sequence, and all sample data are fully mixed through randomly generated indexes to generate a unified multi-task dataset.

[0085] It should be noted that although this application is explained using a multi-task classification scenario as an example, this application is actually also applicable to other multi-task scenarios, including but not limited to multi-task segmentation, natural language understanding, and even multimodal multi-task scenarios.

[0086] The asymmetric low-rank fine-tuning model training method provided in the embodiment of the present application improves the existing HydraLoRA and model training methods, realizes the determination of the number of B matrices on the continuous modality dataset, introduces self-distillation loss to improve the stability of multi-task training, and introduces inter-task balance loss to constrain the training of B matrices, which can effectively improve model performance.

[0087] The present application also provides an asymmetric low-rank fine-tuning model training device. Figure 5 , Figure 5 5 is a schematic diagram of the structure of the asymmetric low-rank fine-tuning model training device provided in the embodiment of the present application. In the embodiment of the present application, the asymmetric low-rank fine-tuning model training device includes a construction module 510 and a training module 520.

[0088] The construction module 510 is used to construct a multi-task dataset.

[0089] The multi-task dataset includes multiple task classification datasets, and each task classification dataset includes sample data of multiple continuous modes.

[0090] The training module 520 is used to pre-train the initial model based on the multi-task dataset to obtain an asymmetric low-rank fine-tuning model.

[0091] Among them, the asymmetric low-rank fine-tuning model includes an asymmetric low-rank adapter, which is fine-tuned based on a first matrix and multiple second matrices. The first matrix is ​​used to extract shared features of multiple task classification data sets, and the second matrix is ​​used to extract specific task features of multiple task classification data sets. The number of second matrices is determined based on the statistical vector of each sample data.

[0092] In some embodiments, the training module 520 is used to discretize each sample data to obtain multiple sub-vectors corresponding to each sample data; determine the statistical vector of each sample data based on each sub-vector; cluster all statistical vectors to obtain multiple candidate quantity values; determine the variance ratio index of each candidate quantity value, and select a target quantity value from the multiple candidate quantity values; wherein the target quantity value is the minimum value among the candidate quantity values ​​whose variance ratio index is greater than a preset variance ratio threshold, and the target quantity value is the quantity of the second matrix.

[0093] In some embodiments, the training module 520 is used to segment each sample data separately to obtain multiple sub-blocks corresponding to each sample data; project each sub-block into a vector to obtain multiple sub-vectors corresponding to each sample data; wherein a sub-vector represents the probability of a sub-block being assigned to different visual words.

[0094] In some embodiments, the initial model includes a teacher model and a student model, and the teacher model and the student model have the same model weights when not pre-trained.

[0095] The training module 520 is used to select the same sample data from the multi-task dataset as fine-tuning samples, input them into the teacher model and the student model respectively, perform pre-training, and obtain an asymmetric low-rank fine-tuning model; wherein, the asymmetric low-rank fine-tuning model is the student model that has completed pre-training; during the pre-training process, the teacher model does not perform gradient updates, the student model performs gradient updates, and the model weights of the teacher model are updated through an exponential sliding average strategy, and gradually approach the model weights of the student model.

[0096] In some embodiments, the student model is pre-trained using a preset loss function, which includes a self-distillation loss and an inter-task balance loss; wherein the self-distillation loss is used to constrain the modules in the student model that are updated by the exponential sliding average strategy; the inter-task balance loss is used to constrain all second matrices in the student model, and the inter-task balance loss is determined based on the routing probability of each second matrix in the student model.

[0097] In some embodiments, a construction module 510 is used to obtain multiple initial task classification data sets; each initial task classification data set includes sample data of multiple continuous modalities, and one sample data includes a sample classification image and a sample classification label corresponding to the sample classification image; a corresponding task classification data set ID is added to each sample data to obtain multiple task classification data sets; each sample data in each task classification data set is arranged in sequence, and all sample data are mixed through randomly generated indexes to generate a multi-task data set.

[0098] The embodiment of the present application also provides an electronic device, Figure 6 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the asymmetric low-rank fine-tuning model training method.

[0099] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0100] An embodiment of the present application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the asymmetric low-rank fine-tuning model training method provided by the above methods.

[0101] An embodiment of the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the asymmetric low-rank fine-tuning model training method provided by the above methods.

[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0103] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for training an asymmetric low-rank fine-tuning model, characterized in that: include: Build multi-task datasets; The multi-task dataset includes a plurality of task classification datasets, each of which includes sample data of a plurality of continuous modalities; Pre-training the initial model based on the multi-task dataset to obtain an asymmetric low-rank fine-tuning model; In which, the asymmetric low-rank fine-tuning model includes an asymmetric low-rank adapter, which is fine-tuned based on a first matrix and multiple second matrices. The first matrix is ​​used to extract shared features of multiple task classification data sets, and the second matrix is ​​used to extract specific task features of multiple task classification data sets. The number of the second matrices is determined based on the statistical vector of each sample data.

2. The asymmetric low-rank fine-tuning model training method according to claim 1, characterized in that Before pre-training the initial model based on the multi-task dataset to obtain the asymmetric low-rank fine-tuning model, the method further includes: Performing discretization processing on each of the sample data to obtain a plurality of sub-vectors corresponding to each of the sample data; Determine a statistical vector of each of the sample data based on each of the sub-vectors; Clustering all the statistical vectors to obtain multiple candidate quantity values; determining a variance ratio index for each candidate quantity value, and selecting a target quantity value from the plurality of candidate quantity values; The target quantity value is the minimum value among the candidate quantity values ​​whose variance ratio index is greater than a preset variance ratio threshold, and the target quantity value is the quantity of the second matrix.

3. The asymmetric low-rank fine-tuning model training method according to claim 2, characterized in that The discretization processing is performed on each sample data to obtain a plurality of sub-vectors corresponding to each sample data, including: Performing segmentation processing on each sample data respectively to obtain a plurality of sub-blocks corresponding to each sample data; Projecting each of the sub-blocks into a vector to obtain a plurality of sub-vectors corresponding to each of the sample data; Wherein, one of the sub-vectors represents the probability of assigning one of the sub-blocks to different visual words.

4. The asymmetric low-rank fine-tuning model training method according to claim 1, characterized in that The initial model includes a teacher model and a student model, and the teacher model and the student model have the same model weight when not pre-trained; The pre-training of the initial model based on the multi-task dataset to obtain an asymmetric low-rank fine-tuning model includes: Selecting the same sample data from the multi-task dataset as fine-tuning samples, inputting them into the teacher model and the student model respectively for pre-training to obtain the asymmetric low-rank fine-tuning model; Among them, the asymmetric low-rank fine-tuning model is a student model that has completed pre-training; during the pre-training process, the teacher model does not perform gradient updates, the student model performs gradient updates, and the model weights of the teacher model are updated through an exponential sliding average strategy, and gradually approach the model weights of the student model.

5. The asymmetric low-rank fine-tuning model training method according to claim 4, characterized in that The student model is pre-trained using a preset loss function, wherein the preset loss function includes a self-distillation loss and an inter-task balance loss; The self-distillation loss is used to constrain the module in the student model that is updated by the exponential moving average strategy; The inter-task balance loss is used to constrain all the second matrices in the student model, and the inter-task balance loss is determined based on the routing probability of each second matrix in the student model.

6. The asymmetric low-rank fine-tuning model training method according to claim 1, characterized in that The multi-task dataset is constructed, including: Acquire multiple initial task classification data sets; each of the initial task classification data sets includes the sample data of multiple continuous modalities, and one of the sample data includes a sample classification image and a sample classification label corresponding to the sample classification image; Adding a corresponding task classification data set ID to each of the sample data to obtain a plurality of the task classification data sets; Each of the sample data in each of the task classification data sets is arranged in sequence, and all of the sample data are mixed using randomly generated indexes to generate the multi-task data set.

7. An asymmetric low-rank fine-tuning model training device, characterized in that: include: Building module for constructing multi-task datasets; The multi-task dataset includes a plurality of task classification datasets, each of which includes sample data of a plurality of continuous modalities; A training module, configured to pre-train the initial model based on the multi-task dataset to obtain an asymmetric low-rank fine-tuning model; In which, the asymmetric low-rank fine-tuning model includes an asymmetric low-rank adapter, which is fine-tuned based on a first matrix and multiple second matrices. The first matrix is ​​used to extract shared features of multiple task classification data sets, and the second matrix is ​​used to extract specific task features of multiple task classification data sets. The number of the second matrices is determined based on the statistical vector of each sample data.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the asymmetric low-rank fine-tuning model training method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the asymmetric low-rank fine-tuning model training method as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the asymmetric low-rank fine-tuning model training method as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Task model generation method, electronic equipment, storage medium and program product

    CN121052286A

  • Task model generation method, electronic device, storage medium and program product

    CN121052286B