Characteristic distillation method, device, equipment and medium of multi-modal large model

By measuring and aligning the characteristic data of the teacher model and student model, the problem of decreasing generalization ability of students' models is solved, and data processing efficiency and generalization ability are improved.

CN120338040APending Publication Date: 2025-07-18GUANGZHOU YUNCONG INFORMATION TECH CO LTD +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510398271.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the generalization ability of the student model is reduced after the feature replacement of the artificial intelligence model, resulting in the problem of low data processing efficiency.

Method used

By evaluating the similarity measurement of the characteristic data of the teacher model and the student model, calculating the characteristic similarity value and the gap value, obtaining the distillation loss value, and aligning it when it is less than the preset threshold, updating the student model, and optimizing it with the training data.

Benefits of technology

It improves the data processing efficiency of the student model, enhances its generalization ability, reduces feature redundancy and useless feature transmission, and improves the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338040A_ABST
    Figure CN120338040A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, particularly provides a feature distillation method, device and equipment for a multi-modal large model and a medium, and aims to solve the problem of low recognition accuracy of a student model after feature replacement. In order to achieve the purpose, the method comprises the steps that similarity measurement is conducted on feature data of a teacher model and feature data of a student model, and a feature similarity value is obtained; obtaining a feature difference value according to the feature data of the teacher model and the feature data of the student model; obtaining a distillation loss value according to the feature similarity value and the feature difference value; determining that the distillation loss value is smaller than a preset threshold value, and performing alignment processing on the feature data of the student model and the feature data of the teacher model to obtain an updated student model; obtaining student model training loss according to the distillation loss value; performing training processing on the updated student model according to the student model training loss and the training data to obtain an optimized student model; and processing the to-be-processed data according to the optimized student model to obtain a processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a method, device, equipment and medium for feature distillation of a multimodal large model. Background Art

[0002] Currently, when deploying an artificial intelligence large model, due to the artificial intelligence large model's capabilities of processing large-scale data, processing complex problems, and having higher accuracy, the artificial intelligence large model is usually deployed in a cloud server to process the data to be processed. When processing specific data to be processed, the artificial intelligence large model needs to be adjusted. However, adjusting the artificial intelligence large model takes a lot of time and effort, and there will also be problems of high model deployment cost and low data processing efficiency during subsequent use. Therefore, the model compression technology based on feature distillation is widely used in artificial intelligence large models to compress the artificial intelligence large model. In feature distillation, the artificial intelligence large model is also called the teacher model, and the artificial intelligence small model is also called the student model.

[0003] In the prior art, the feature data in the student model is replaced with the feature data of the teacher model to improve the data processing ability of the student model. However, the feature data of the teacher model may be low-level features, and affected by the law of scale, after the teacher model undergoes feature distillation, the generalization ability of the student model decreases, resulting in the technical problem of low data processing efficiency of the student model. Summary of the Invention

[0004] In order to overcome the above defects, this application is proposed to provide a solution to solve or at least partially solve the technical problem that the generalization ability of the student model after feature replacement decreases, resulting in low data processing efficiency of the student model.

[0005] In a first aspect, this application provides a method for feature distillation of a multimodal large model, including:

[0006] Obtain the feature data of the teacher model, the feature data of the student model, the training data, and the data to be processed;

[0007] Perform similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value;

[0008] Obtain a feature gap value according to the feature data of the teacher model and the feature data of the student model;

[0009] Obtain a distillation loss value according to the feature similarity value and the feature gap value;

[0010] Determine that the distillation loss value is less than a preset threshold, align the feature data of the student model and the feature data of the teacher model to obtain an updated student model;

[0011] Obtain the training loss of the student model according to the distillation loss value;

[0012] Train the updated student model according to the training loss of the student model and the training data to obtain an optimized student model;

[0013] Process the data to be processed according to the optimized student model to obtain a processing result.

[0014] In a technical solution of the above feature distillation method for a multi-modal large model, the obtaining of the feature similarity value by performing similarity measurement on the feature data of the teacher model and the feature data of the student model includes:

[0015] Perform dimensionality reduction processing on the feature data of the teacher model and the feature data of the student model respectively to obtain the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model;

[0016] Perform linear transformation processing on the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model respectively to obtain a teacher model transformation matrix and a student model transformation matrix;

[0017] Perform activation processing on the teacher model transformation matrix and the student model transformation matrix respectively to obtain teacher model activation data and student model activation data;

[0018] Perform similarity measurement on the teacher model activation data and the student model activation data to obtain the feature similarity value.

[0019] In a technical solution of the above feature distillation method for a multi-modal large model, the obtaining of the feature gap value according to the feature data of the teacher model and the feature data of the student model includes:

[0020] Perform dimensionality reduction processing on the feature data of the teacher model and the feature data of the student model respectively to obtain teacher model dimensionality-reduced data and student model dimensionality-reduced data;

[0021] Obtain a dimensionality-reduced data difference according to the teacher model dimensionality-reduced data and the student model dimensionality-reduced data;

[0022] Obtain the feature gap value according to the dimensionality-reduced data difference.

[0023] In a technical solution of the above feature distillation method for a multi-modal large model, the obtaining of the distillation loss value according to the feature similarity value and the feature gap value includes:

[0024] Obtain the feature similarity value and the feature difference value of each layer of features in the feature data of the teacher model and the feature data of the student model;

[0025] Sum up the feature similarity value of each layer of features and the feature difference value of each layer of features to obtain the distillation loss value.

[0026] In a technical solution of the above feature distillation method for a multi-modal large model, the obtaining the training loss of the student model according to the distillation loss value includes:

[0027] Obtain the prediction error;

[0028] According to the distillation loss value and a preset weight, obtain an adjusted distillation loss value;

[0029] According to the adjusted distillation loss value and the prediction error, obtain the training loss of the student model.

[0030] In a technical solution of the above feature distillation method for a multi-modal large model, the determining that the distillation loss value is less than a preset threshold, and aligning the feature data of the student model and the feature data of the teacher model to obtain an updated student model includes:

[0031] Obtain the dimension of the feature data of the student model and the dimension of the feature data of the teacher model;

[0032] Determine that the dimension of the feature data of the student model is the same as the dimension of the feature data of the teacher model, and replace the feature data of the student model with the feature data of the teacher model to obtain the updated student model.

[0033] In a technical solution of the above feature distillation method for a multi-modal large model, the method further includes:

[0034] Determine that the dimension of the feature data of the student model is less than the dimension of the feature data of the teacher model;

[0035] Perform dimensionality reduction processing on the feature data of the teacher model to obtain teacher model feature dimensionality reduction data;

[0036] Replace the feature data of the student model with the teacher model feature dimensionality reduction data to obtain the updated student model.

[0037] In a technical solution of the above feature distillation method for a multi-modal large model, the method further includes:

[0038] Determine that the dimension of the feature data of the student model is greater than the dimension of the feature data of the teacher model;

[0039] Perform dimensionality increase processing on the feature data of the teacher model to obtain the dimensionality-increased feature data of the teacher model;

[0040] Replace the feature data of the student model with the dimensionality-increased feature data of the teacher model to obtain the updated student model.

[0041] In a second aspect, the present application provides a feature distillation device for a multimodal large model, including:

[0042] An acquisition module that acquires the feature data of the teacher model, the feature data of the student model, training data, and data to be processed;

[0043] An analysis module for performing similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value;

[0044] The analysis module is further configured to obtain a feature gap value according to the feature data of the teacher model and the feature data of the student model;

[0045] The analysis module is further configured to obtain a distillation loss value according to the feature similarity value and the feature gap value;

[0046] An alignment module for determining that the distillation loss value is less than a preset threshold, and performing alignment processing on the feature data of the student model and the feature data of the teacher model to obtain an updated student model;

[0047] The analysis module is further configured to obtain the training loss of the student model according to the distillation loss value;

[0048] A training module for performing training processing on the updated student model according to the training loss of the student model and the training data to obtain an optimized student model;

[0049] An identification module for processing the data to be processed according to the optimized student model to obtain a processing result.

[0050] In a third aspect, the present application provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the method according to any one of the first aspects through the computer program.

[0051] In a fourth aspect, the present application provides a computer-readable storage medium, in which multiple program codes are stored, and the program codes are adapted to be loaded and run by a processor to execute the method according to any one of the first aspects.

[0052] The present application provides a feature distillation method, device, equipment and medium for a multi-modal large model. Specifically, the method is as follows: obtain the feature data of the teacher model, the feature data of the student model, the training data and the data to be processed; perform a similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value; obtain a feature gap value according to the feature data of the teacher model and the feature data of the student model; obtain a distillation loss value according to the feature similarity value and the feature gap value; determine that the distillation loss value is less than a preset threshold, align the feature data of the student model and the feature data of the teacher model to obtain an updated student model; obtain the training loss of the student model according to the distillation loss value; perform a training process on the updated student model according to the training loss of the student model and the training data to obtain an optimized student model; process the data to be processed according to the optimized student model to obtain a processing result, thereby improving the processing efficiency of the student model. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Referring to the accompanying drawings, the disclosure of the present application will become more understandable. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present application.

[0054] In addition, similar numbers in the drawings are used to represent similar components, where:

[0055] Figure 1 is a schematic flowchart of Embodiment 1 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application;

[0056] Figure 2 is a schematic flowchart of Embodiment 2 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application;

[0057] Figure 3 is a schematic flowchart of Embodiment 3 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application;

[0058] Figure 4 is a schematic flowchart of Embodiment 4 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application;

[0059] Figure 5 is a schematic flowchart of Embodiment 5 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application;

[0060] Figure 6 is a schematic flowchart of Embodiment 6 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application;

[0061] Figure 7 is a schematic flowchart of Embodiment 7 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application;

[0062] Figure 8 Schematic flowchart of Embodiment 8 of a feature distillation method for a multimodal large model provided by an embodiment of the present application;

[0063] Fig. 9 Schematic structural diagram of Embodiment 1 of a feature distillation device for a multimodal large model provided by an embodiment of the present application;

[0064] Fig.10 Schematic structural diagram of Embodiment 1 of an electronic device provided by an embodiment of the present application;

[0065] Reference numerals list :

[0066] 11: Acquisition module; 12: Analysis module; 13: Alignment module; 14: Training module; 15: Recognition module; 21: Processor; 22: Memory. Detailed implementation manners

[0067] The following describes some implementation manners of the present application with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principle of the present application and are not intended to limit the protection scope of the present application.

[0068] In the description of the present application, "module" and "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various suitable sensors, communication ports, a memory, and may also include a software part, such as program code, or may be a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. A non-transitory computer-readable storage medium includes any suitable medium that can store program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and so on. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one A or B" or "at least one of A and B" has a meaning similar to "A and / or B" and may include only A, only B, or A and B. The singular terms "a" and "this" may also include the plural form.

[0069] In the prior art, the features of the teacher model and the features of the student model are combined, and the features of the teacher model are passed to the student model through predefined weights. However, passing the features of the teacher model to the student model will result in redundancy of feature data and passing a large amount of useless features to the student model, thereby causing the technical problem of low data processing efficiency of the student model.

[0070] Here, some terms involved in this application are explained first:

[0071] Average pooling: Average pooling is an operation in machine learning used to reduce the feature dimension, reduce the computational amount, and retain important information at the same time.

[0072] Linear transformation matrix: A linear transformation matrix is a special matrix that can be applied to a vector through mathematical operations (such as matrix multiplication) to transform word vectors or sentence vectors.

[0073] Non-linear activation function: It is used to transform or increase the complexity of the input data, enabling the network to learn and perform more complex tasks. Common non-linear activation functions include sigmoid, ReLU, tanh, etc.

[0074] Bilinear matrix: It is a function that receives two input vectors and generates an output. It is called "bilinear" because the function operates linearly on each input independently.

[0075] Attention mechanism: The attention mechanism is an artificial intelligence technology. This mechanism allows the system to "focus" on the key parts when processing information, improving the efficiency and accuracy of recognition and understanding.

[0076] Based on this, to solve the above technical problems, this application provides a new feature distillation method for a multi-modal large model to improve the processing efficiency of the student model.

[0077] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the drawings.

[0078] Figure 1 It is a schematic flowchart of Embodiment 1 of a feature distillation method for a multi-modal large model provided by an embodiment of this application. As Figure 1 shown, specifically, the method includes:

[0079] Step S101: Obtain the feature data of the teacher model, the feature data of the student model, the training data, and the data to be processed.

[0080] In this embodiment, both the teacher model and the student model contain multiple pieces of feature data, and the feature data of the teacher model and the feature data of the student model are feature maps.

[0081] Step S102: Perform similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value.

[0082] In this embodiment, the feature data of the teacher model and the feature data of the student model are both matrix data, and the feature similarity value is obtained through a similarity calculation formula.

[0083] In this embodiment, for example, the Euclidean distance between the feature data of the teacher model and the feature data of the student model can be calculated to obtain the feature similarity value.

[0084] Step S103: Obtain a feature gap value according to the feature data of the teacher model and the feature data of the student model.

[0085] In this embodiment, when the student model learns the feature data of the teacher model, it should learn the difference between the feature data of the teacher model and the feature data of the student model. Therefore, the feature gap value can be obtained according to the difference between the feature data of the teacher model and the feature data of the student model.

[0086] Step S104: Obtain a distillation loss value according to the feature similarity value and the feature gap value.

[0087] In this embodiment, the feature similarity value and the feature gap value are summed, and the sum value obtained is the distillation loss value.

[0088] Step S105: Determine that the distillation loss value is less than a preset threshold, and align the feature data of the student model and the feature data of the teacher model to obtain an updated student model.

[0089] In this embodiment, it is determined that the distillation loss value is less than a preset threshold. After processing the dimension of the feature data of the teacher model, the dimension of the processed feature data of the teacher model is made consistent with the dimension of the feature data of the student model, and then the feature data of the student model is replaced with the processed feature data of the teacher model to obtain an updated student model.

[0090] In this embodiment, for example, the preset threshold is 0.01.

[0091] Step S106: Obtain the training loss of the student model according to the distillation loss value.

[0092] In this embodiment, the training loss of the student model consists of two parts: the distillation loss value and the prediction loss after each training. The sum value of the distillation loss value and the prediction loss is the training loss of the student model.

[0093] Step S107: Perform training processing on the updated student model according to the training loss of the student model and the training data to obtain an optimized student model.

[0094] In this embodiment, according to the above method, feature distillation can be performed on large models based on the Transformer architecture, such as large language models, vision language models, and multi-modal large language models.

[0095] In this embodiment, during each training, the training data is input into the updated student model to obtain a trained model. Then, the test data is input into the trained model to obtain a test result. According to the test result and the true result of the test data, a prediction loss value is obtained. According to the prediction loss value and the distillation loss value, the training loss of the student model is obtained. Then, the model parameters of the updated student model are updated using the training loss of the student model. When it is determined that the training loss of the student model is less than the preset loss value, the current student model is the optimized student model.

[0096] Step S108: Process the data to be processed according to the optimized student model to obtain a processing result.

[0097] In this embodiment, the data to be processed is input into the optimized student model, and the optimized student model processes the data to be processed to obtain a processing result of the data to be processed.

[0098] In this embodiment, for example, according to the optimized student model, image data can be recognized to obtain the category of the object in the image, the category probability, and the position of the object in the image.

[0099] In this embodiment, the feature data of the teacher model, the feature data of the student model, the training data, and the data to be processed are obtained; the similarity between the feature data of the teacher model and the feature data of the student model is measured to obtain a feature similarity value; based on the feature data of the teacher model and the feature data of the student model, a feature gap value is obtained; based on the feature similarity value and the feature gap value, a distillation loss value is obtained; it is determined that the distillation loss value is less than a preset threshold, and the feature data of the student model and the feature data of the teacher model are aligned to obtain an updated student model; based on the distillation loss value, the training loss of the student model is obtained; based on the training loss of the student model and the training data, the updated student model is trained to obtain an optimized student model; based on the optimized student model, the data to be processed is processed to obtain a processing result. Compared with the prior art, in the low-level artificial intelligence student model, the feature data is replaced with the feature data of the artificial intelligence teacher model, which reduces the generalization ability of the student model and further leads to low processing efficiency of the artificial intelligence student model. In this application, the similarity between the feature data of the teacher model and the feature data of the student model is measured to obtain a feature similarity value, then based on the feature data of the teacher model and the feature data of the student model, a feature gap value is obtained, and based on the feature similarity value and the feature gap value, a distillation loss value is obtained. When it is determined that the distillation loss value is less than a preset threshold, the feature data of the student model and the feature data of the teacher model are aligned to obtain an updated student model, then based on the distillation loss value, the training loss of the student model is obtained, and then based on the training loss of the student model and the training data, the updated student model is trained to obtain an optimized student model. Finally, based on the optimized student model, the data to be processed is processed to obtain a processing result, thereby improving the processing efficiency of the student model.

[0100] Figure 2 FIG. is a schematic flowchart of Embodiment 2 of a feature distillation method for a multimodal large model provided by an embodiment of the present application. On the basis of the above embodiment, as Figure 2 shown, specifically, a specific implementation manner of step S102 is:

[0101] Step S201: The feature data of the teacher model and the feature data of the student model are respectively dimension-reduced to obtain the dimension-reduced feature data of the teacher model and the dimension-reduced feature data of the student model.

[0102] In this embodiment, the feature data of the teacher model and the feature data of the student model are respectively dimension-reduced by convolution or pooling to obtain the dimension-reduced feature data of the teacher model and the dimension-reduced feature data of the student model.

[0103] In this embodiment, for example, through average pooling, the dimensionality reduction processing is respectively performed on the feature data of the teacher model and the feature data of the student model to obtain the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model.

[0104] Step S202: Respectively perform linear transformation processing on the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model to obtain the teacher model transformation matrix and the student model transformation matrix.

[0105] In this embodiment, the product of the dimensionality-reduced feature data of the teacher model and the linear transformation matrix of the teacher model gives the teacher model transformation matrix, and the product of the dimensionality-reduced feature data of the student model and the linear transformation matrix of the student model gives the student model transformation matrix.

[0106] Step S203: Respectively perform activation processing on the teacher model transformation matrix and the student model transformation matrix to obtain the teacher model activation data and the student model activation data.

[0107] In this embodiment, using a non-linear activation function, respectively perform activation processing on the teacher model transformation matrix and the student model transformation matrix to obtain the teacher model activation data and the student model activation data.

[0108] Step S204: Perform similarity measurement on the teacher model activation data and the student model activation data to obtain a feature similarity value.

[0109] In this embodiment, according to Formula 1:

[0110]

[0111] The feature similarity value W distill is obtained. Wherein, q T is the teacher model activation data, k S is the student model activation data, W Q -K is a bilinear matrix, P T and P S are position encoding information, d is a preset number of channels, and softmax is an activation function.

[0112] In this embodiment, the teacher model activation data and the student model activation data in Formula 1 are transformed from features of different dimensions. Therefore, using a bilinear matrix can improve the generalization of the student model.

[0113] In this embodiment, dimensionality reduction processing is respectively performed on the feature data of the teacher model and the feature data of the student model to obtain the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model; linear transformation processing is respectively performed on the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model to obtain the teacher model transformation matrix and the student model transformation matrix; activation processing is respectively performed on the teacher model transformation matrix and the student model transformation matrix to obtain the teacher model activation data and the student model activation data; similarity measurement is performed on the teacher model activation data and the student model activation data to obtain the feature similarity value.

[0114] Figure 3 FIG. 3 is a schematic flowchart of Embodiment 3 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application. On the basis of the above embodiment, as Figure 3 shown, specifically, one implementation manner of step S103 includes:

[0115] Step S301: Respectively perform dimensionality reduction processing on the feature data of the teacher model and the feature data of the student model to obtain the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model.

[0116] In this embodiment, according to the full-channel average pooling algorithm, dimensionality reduction processing is respectively performed on the feature data of the teacher model and the feature data of the student model to obtain the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model.

[0117] Step S302: Obtain the dimensionality-reduced data difference according to the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model.

[0118] In this embodiment, the difference between the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model is the dimensionality-reduced data difference.

[0119] Step S303: Obtain the feature gap value according to the dimensionality-reduced data difference.

[0120] In this embodiment, the L2 norm of the dimensionality-reduced data difference is the feature gap value.

[0121] In this embodiment, dimensionality reduction processing is respectively performed on the feature data of the teacher model and the feature data of the student model to obtain the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model; the dimensionality-reduced data difference is obtained according to the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model; the feature gap value is obtained according to the dimensionality-reduced data difference.

[0122] Figure 4 FIG. 4 is a schematic flowchart of Embodiment 4 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application. On the basis of the above embodiment, as Figure 4 shown, specifically, one implementation manner of step S104 includes:

[0123] Step S401: Obtain the feature similarity value and the feature difference value of each layer of features in the feature data of the teacher model and the student model.

[0124] Step S402: Sum up the feature similarity value of each layer of features and the feature difference value of each layer of features to obtain the distillation loss value.

[0125] In this embodiment, according to Formula 2:

[0126] L FD =∑ layers ∑ B W distill Distil n (2)

[0127] Obtain the distillation loss value L FD . Wherein, layers is the number of layers of features, B is the batch size of the input data, W distill is the feature similarity value, and Distil n is the feature difference value.

[0128] In this embodiment, obtain the feature similarity value of each layer of features and the feature difference value of each layer of features in the feature data of the teacher model and the student model; sum up the feature similarity value of each layer of features and the feature difference value of each layer of features to obtain the distillation loss value.

[0129] Figure 5 This is a schematic flowchart of Embodiment 5 of the feature distillation method for a multi-modal large model provided by an embodiment of the present application. On the basis of the above embodiments, as Figure 5 shown, specifically, one implementation manner of step S106 includes:

[0130] Step S501: Obtain the prediction error.

[0131] Step S502: Obtain the adjusted distillation loss value according to the distillation loss value and the preset weight.

[0132] In this embodiment, obtain the adjusted distillation loss value according to the product of the distillation loss value and the preset weight.

[0133] In this embodiment, the preset weight can adjust the loss of feature distillation and is set according to needs, and its value is greater than 0 and less than 1.

[0134] Step S503: Obtain the training loss of the student model according to the adjusted distillation loss value and the prediction error.

[0135] In this embodiment, the sum value of the adjusted distillation loss value and the prediction error is the training loss of the student model.

[0136] In this embodiment, taking the feature similarity value as part of the training loss of the student model can suppress the gradient fluctuation caused by large differences in the training data during the training process.

[0137] In this embodiment, a prediction error is obtained; an adjusted distillation loss value is obtained according to the distillation loss value and a preset weight; and a training loss of the student model is obtained according to the adjusted distillation loss value and the prediction error.

[0138] Figure 6 FIG. is a schematic flowchart of Embodiment 6 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application. On the basis of the above embodiment, as Figure 6 shown, specifically, one implementation manner of step S105 includes:

[0139] Step S601: Obtain the dimension of the feature data of the student model and the dimension of the feature data of the teacher model.

[0140] In this embodiment, the feature data of the student model and the feature data of the teacher model are both matrices. Therefore, the dimension of the feature data of the student model and the dimension of the feature data of the teacher model are the size of the matrix.

[0141] Step S602: Determine that the dimension of the feature data of the student model is the same as the dimension of the feature data of the teacher model, and replace the feature data of the student model with the feature data of the teacher model to obtain an updated student model.

[0142] In this embodiment, if it is determined that the dimension of the feature data of the student model is the same as the dimension of the feature data of the teacher model, the feature data of the student model is deleted, and the feature data of the teacher model is added to the student model.

[0143] In this embodiment, the dimension of the feature data of the student model and the dimension of the feature data of the teacher model are obtained; it is determined that the dimension of the feature data of the student model is the same as the dimension of the feature data of the teacher model, and the feature data of the student model is replaced with the feature data of the teacher model to obtain an updated student model.

[0144] Figure 7 FIG. is a schematic flowchart of Embodiment 7 of a feature distillation method for a multi-modal large model provided by an embodiment of the present application. As Figure 7 shown, after step S601, the method includes:

[0145] Step S701: Determine that the dimension of the feature data of the student model is smaller than the dimension of the feature data of the teacher model.

[0146] Step S702: Perform dimensionality reduction processing on the feature data of the teacher model to obtain teacher model feature dimensionality reduction data.

[0147] In this embodiment, through convolution and pooling, the feature data of the teacher model is dimensionally reduced to obtain the dimensionally reduced feature data of the teacher model.

[0148] Step S703: Replace the feature data of the student model with the dimensionally reduced feature data of the teacher model to obtain an updated student model.

[0149] In this embodiment, it is determined that the dimension of the feature data of the student model is smaller than the dimension of the feature data of the teacher model; the feature data of the teacher model is dimensionally reduced to obtain the dimensionally reduced feature data of the teacher model; the feature data of the student model is replaced with the dimensionally reduced feature data of the teacher model to obtain an updated student model.

[0150] Figure 8 It is a schematic flowchart of Embodiment 8 of a feature distillation method for a multimodal large model provided by an embodiment of the present application. As Figure 8 shown, after step S601, the method further includes:

[0151] Step S801: Determine that the dimension of the feature data of the student model is larger than the dimension of the feature data of the teacher model.

[0152] Step S802: Perform dimension elevation processing on the feature data of the teacher model to obtain the dimension-elevated feature data of the teacher model.

[0153] In this embodiment, through transposed convolution, the feature data of the teacher model is dimensionally elevated to obtain the dimension-elevated feature data of the teacher model.

[0154] Step S803: Replace the feature data of the student model with the dimension-elevated feature data of the teacher model to obtain an updated student model.

[0155] In this embodiment, it is determined that the dimension of the feature data of the student model is larger than the dimension of the feature data of the teacher model; the feature data of the teacher model is dimensionally elevated to obtain the dimension-elevated feature data of the teacher model; the feature data of the student model is replaced with the dimension-elevated feature data of the teacher model to obtain an updated student model.

[0156] Furthermore, the present application also provides a feature distillation device for a multimodal large model.

[0157] Fig. 9 It is a schematic structural diagram of Embodiment 1 of a feature distillation device for a multimodal large model provided by an embodiment of the present application. As Fig. 9As shown in the figure, the feature distillation device of the multi-modal large model in the embodiment of the present application mainly includes an acquisition module 11, an analysis module 12, an alignment module 13, a training module 14, and an identification module 15. One or more of the above modules can be combined into one module. In some embodiments, the acquisition module 11 can be configured to acquire the feature data of the teacher model, the feature data of the student model, the training data, and the data to be processed. The analysis module 12 can be configured to perform a similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value. The analysis module 12 can also be configured to obtain a feature gap value according to the feature data of the teacher model and the feature data of the student model. The analysis module 12 can also be configured to obtain a distillation loss value according to the feature similarity value and the feature gap value. The alignment module 13 can be configured to determine that the distillation loss value is less than a preset threshold, and align the feature data of the student model and the feature data of the teacher model to obtain an updated student model. The analysis module 12 can also be configured to obtain the training loss of the student model according to the distillation loss value. The training module 14 can be configured to perform a training process on the updated student model according to the training loss of the student model and the training data to obtain an optimized student model; the identification module 15 can be configured to process the data to be processed according to the optimized student model to obtain a processing result.

[0158] The above-mentioned feature distillation device of the multi-modal large model is used to execute Figures 1 to 8 the embodiment of the feature distillation method of the multi-modal large model shown in the figure. The technical principles, the technical problems solved, and the technical effects produced by the two are similar. Those skilled in the art of this technology can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the feature distillation device of the multi-modal large model can refer to the content described in the embodiment of the feature distillation method of the multi-modal large model whose execution subject is the feature distillation device of the multi-modal large model, and will not be elaborated here.

[0159] Those skilled in the art can understand that all or part of the processes in the method of the above-mentioned embodiment of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate forms, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0160] Furthermore, the present application also provides an electronic device.

[0161] Fig.10 FIG. 1 is a schematic structural diagram of Embodiment 1 of an electronic device provided by an embodiment of the present application. As Fig.10 shown, the electronic device includes at least one processor 21 and a memory 22. The memory 22 can be configured to store a program for executing the feature distillation method of the multi-modal large model in the above-mentioned Figures 1 to 8 shown embodiment. The processor 21 can be configured to execute the program in the memory 22, and the program includes, but is not limited to, a program for executing a feature distillation method of a multi-modal large model in the above-mentioned method embodiment. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The electronic device can be a control device including various electronic devices.

[0162] In the present application, the above-mentioned device and equipment can be an artificial intelligence-oriented training and inference all-in-one machine, intelligent terminal device, intelligent wearable and other intelligent devices. The above-mentioned device and equipment can run offline and have low-power and low-latency algorithm inference capabilities.

[0163] Furthermore, the present application also provides a computer-readable storage medium. In an embodiment of a computer-readable storage medium according to the present application, the computer-readable storage medium can be configured to store and execute the above-mentioned method Figures 1 to 8A program for the feature distillation method of the multi-modal large model of the illustrated embodiment. This program can be loaded and run by a processor to implement the above-mentioned feature distillation method of the multi-modal large model. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. This computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present application is a non-transitory computer-readable storage medium.

[0164] Furthermore, it should be understood that since the setting of each module is only to illustrate the functional units of the device of the present application, the corresponding physical devices of these modules can be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, the number of each module in the figure is only illustrative.

[0165] Those skilled in the art can understand that the various modules in the device can be adaptively split or combined. Such splitting or combination of specific modules will not cause the technical solution to deviate from the principle of the present application. Therefore, the technical solutions after splitting or combination will all fall within the protection scope of the present application.

[0166] So far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principle of the present application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present application.

Claims

1. A feature distillation method for multi-modal large models, characterized in that, Including: Obtain the feature data of the teacher model, the feature data of the student model, the training data, and the data to be processed; Perform similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value; Obtain a feature gap value according to the feature data of the teacher model and the feature data of the student model; Obtain a distillation loss value according to the feature similarity value and the feature gap value; Determine that the distillation loss value is less than a preset threshold, align the feature data of the student model and the feature data of the teacher model to obtain an updated student model; Obtain the training loss of the student model according to the distillation loss value; Train the updated student model according to the training loss of the student model and the training data to obtain an optimized student model; Process the data to be processed according to the optimized student model to obtain a processing result.

2. The method according to claim 1, characterized in that, The performing similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value includes: Perform dimensionality reduction processing on the feature data of the teacher model and the feature data of the student model respectively to obtain the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model; Perform linear transformation processing on the dimensionality-reduced feature data of the teacher model and the dimensionality-reduced feature data of the student model respectively to obtain a teacher model transformation matrix and a student model transformation matrix; Perform activation processing on the teacher model transformation matrix and the student model transformation matrix respectively to obtain teacher model activation data and student model activation data; Perform similarity measurement on the teacher model activation data and the student model activation data to obtain the feature similarity value.

3. The method according to claim 1, wherein The obtaining a feature gap value according to the feature data of the teacher model and the feature data of the student model includes: Perform dimensionality reduction processing on the feature data of the teacher model and the feature data of the student model respectively to obtain the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model; Obtain a dimensionality-reduced data difference according to the dimensionality-reduced data of the teacher model and the dimensionality-reduced data of the student model; Obtain the feature gap value according to the dimensionality-reduced data difference.

4. The method according to claim 1, wherein The obtaining a distillation loss value according to the feature similarity value and the feature gap value includes: Obtain the feature similarity value of each layer of features and the feature gap value of each layer of features in the feature data of the teacher model and the feature data of the student model; Perform a summation process on the feature similarity value of each layer of features and the feature gap value of each layer of features to obtain a distillation loss value.

5. The method according to claim 1, wherein The obtaining the training loss of the student model according to the distillation loss value includes: Obtain a prediction error; Obtain an adjusted distillation loss value according to the distillation loss value and a preset weight; Obtain the training loss of the student model according to the adjusted distillation loss value and the prediction error.

6. The method according to claim 1, characterized in that The determining that the distillation loss value is less than a preset threshold, aligning the feature data of the student model and the feature data of the teacher model to obtain an updated student model includes: Obtain the dimension of the feature data of the student model and the dimension of the feature data of the teacher model; Determine that the dimension of the feature data of the student model is the same as that of the feature data of the teacher model, and replace the feature data of the student model with the feature data of the teacher model to obtain the updated student model.

7. The method according to claim 6, wherein The method further includes: Determine that the dimension of the feature data of the student model is smaller than that of the feature data of the teacher model; Perform dimensionality reduction processing on the feature data of the teacher model to obtain the dimension-reduced feature data of the teacher model; Replace the feature data of the student model with the dimension-reduced feature data of the teacher model to obtain the updated student model.

8. The method according to claim 6, characterized in that, The method further includes: Determine that the dimension of the feature data of the student model is larger than that of the feature data of the teacher model; Perform dimensionality increase processing on the feature data of the teacher model to obtain the dimension-increased feature data of the teacher model; Replace the feature data of the student model with the dimension-increased feature data of the teacher model to obtain the updated student model.

9. A feature distillation device for a multi-modal large model, characterized in that, It includes: An acquisition module that acquires the feature data of the teacher model, the feature data of the student model, the training data, and the data to be processed; An analysis module for performing similarity measurement on the feature data of the teacher model and the feature data of the student model to obtain a feature similarity value; The analysis module is further configured to obtain a feature gap value according to the feature data of the teacher model and the feature data of the student model; The analysis module is further configured to obtain a distillation loss value according to the feature similarity value and the feature gap value; An alignment module for determining that the distillation loss value is less than a preset threshold, and performing alignment processing on the feature data of the student model and the feature data of the teacher model to obtain an updated student model; The analysis module is further configured to obtain a training loss of the student model according to the distillation loss value; A training module for training the updated student model according to the training loss of the student model and the training data to obtain an optimized student model; An identification module for processing the data to be processed according to the optimized student model to obtain a processing result.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.

11. A computer-readable storage medium storing multiple program codes, characterized in that, The program code is suitable for being loaded and run by the processor to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Financial risk judgment method and device and feature analysis model training method and device

    CN120931396A

  • Training method and device of palm recognition model and access control recognition method and device

    CN121564768A

  • Palm identification model training method and device, and access control identification method and device

    CN121564768B