Student Model Optimization Method, Device, and Readable Storage Medium Based on Distillation Learning
By calculating vector relationship differences on the attention layer of the student model and the teacher model, determining the target loss value of the student model and optimizing the processing, the problem of traditional methods ignoring the internal impact of the model is solved, and the optimization performance and prediction accuracy of the student model are improved.
Patent Information
- Application Number
- CN202510132774.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The traditional method of calculating loss value ignores the internal influence of the model during the distillation of the model, resulting in the student model being unable to fully learn the deep characteristics of the teacher model, affecting the optimization performance of the student model.
By calculating vector relationship differences on the attention layer of the student model and the teacher model, the target loss value of the student model is determined and the student model is optimized based on this. The specific steps include: calculating the attention layer vector relationship in each coding module of the student model and the teacher model, calculating the vector differences between each coding module based on these relationships, and finally determining the optimization goals of the student model based on these differences.
By learning the attention layer difference between the teacher model and the student model, the performance of the student model is optimized, and the prediction accuracy and computational efficiency of the student model are improved.
Smart Images

Figure CN119558353B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distillation learning processing, and in particular to a method for optimizing a student model based on distillation learning, an electronic device, and a computer-readable storage medium. Background Art
[0002] Model distillation is a technique for transferring the knowledge of a large and complex teacher model to a small student model, enabling the student model to inherit the performance advantages of the teacher model while maintaining a low computational cost and a faster inference speed.
[0003] During the model distillation process, in order to make the output of the student model as close as possible to the value that is really desired to be predicted, the difference between the value output by the teacher model and the value output by the student model, and the difference between the value output by the student model and the true value can be compared. The two are combined to obtain the loss value of the student model, and the student model is optimized with the loss value.
[0004] It can be seen that the loss value is an important parameter for measuring the quality of the student model, and the accuracy of the loss value determines the optimization ability of the model. The traditional loss value calculation method mainly calculates based on the predicted value finally output by the model, but ignores the internal influence of the model, resulting in the student model being unable to fully learn the deep features of the teacher model and affecting the optimization performance of the student model. Summary of the Invention
[0005] The main technical problem to be solved by this application is to provide a method for optimizing a student model based on distillation learning, an electronic device, and a computer-readable storage medium to improve the optimization effect of the student model.
[0006] To solve the above technical problems, a technical solution adopted in this application is: to provide an optimization method for a student model based on distillation learning. The optimization method for the student model based on distillation learning is applied to a student model optimization device. The student model optimization device includes a student model and a teacher model. The student model includes at least one first encoding module, and each first encoding module includes at least one first attention layer. The teacher model includes at least one second encoding module, and each second encoding module includes at least one second attention layer. The method includes: inputting the data to be processed into the student model to obtain query vectors, key vectors, and value vectors of each first attention layer in each first encoding module; determining a first vector relationship value of each first encoding module and / or a second vector relationship value between any two first encoding modules according to at least one of the query vectors, key vectors, and value vectors of each first attention layer; inputting the data to be processed into the teacher model to obtain query vectors, key vectors, and value vectors of each second attention layer in each second encoding module; determining a third vector relationship value of each second encoding module and / or a fourth vector relationship value between any two second encoding modules according to at least one of the query vectors, key vectors, and value vectors of each second attention layer; determining a first relationship difference between the first vector relationship value of each first encoding module in the student model and the third vector relationship value of the corresponding second encoding module in the teacher model according to the attribute information of the encoding module, and / or, determining a second relationship difference between the second vector relationship value in the student model and the corresponding fourth vector relationship value in the teacher model according to the attribute information of the encoding module; determining the target loss value of the student model according to the first relationship difference and / or the second relationship difference; and optimizing the student model according to the target loss value to obtain an optimized student model.
[0007] To solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, including a memory and a processor. The memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the above-mentioned optimization method for the student model based on distillation learning.
[0008] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium including stored program data. When the program data is executed by a processor, it is used to implement the above-mentioned optimization method for the student model based on distillation learning.
[0009] In the above solution, the first vector relationship value of each first encoding module and / or the second vector relationship value between any two first encoding modules are determined based on at least one of the query vector, key vector, and value vector of each attention layer in each first encoding module of the student model; the third vector relationship value of each second encoding module and / or the fourth vector relationship value between any two second encoding modules are determined based on at least one of the query vector, key vector, and value vector of each attention layer in each second encoding module of the teacher model. Thus, not only can the vector relationships of the attention layers inside the model be fully learned, but also the vector relationships between multiple encoding modules in the same model can be learned; then, the first relationship difference is determined based on the first vector relationship value and the third vector relationship value and / or the second relationship difference is determined based on the second vector relationship value and the fourth vector relationship value; the target loss value of the student model is determined based on the first relationship difference and / or the second relationship difference, and the student model is optimized based on the target loss value. Thus, the student model is optimized through the differences between the student model and the teacher model in the attention layer, improving the performance of the student model. Description of the Drawings
[0010] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:
[0011] Figure 1 is a schematic flowchart of an exemplary embodiment of the student model optimization method shown in the present application;
[0012] Figure 2 is Figure 1 a schematic flowchart of an exemplary embodiment of step S120 in the student model optimization method shown;
[0013] Figure 3 is a schematic framework diagram of an exemplary embodiment of the student model optimization method shown in the present application;
[0014] Figure 4 is a schematic structural diagram of an exemplary embodiment of the student model optimization device shown in the present application;
[0015] Figure 5 is a schematic structural diagram of an embodiment of an electronic device provided by the present application;
[0016] Figure 6 is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. Detailed Embodiments
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. In addition, it should be noted that for the convenience of description, only the parts related to the present application rather than all the structures are shown in the accompanying drawings. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0018] First of all, it should be noted that distillation learning usually transfers the knowledge of a large and complex teacher model to a smaller and more efficient student model. The teacher model is generally a trained model used to provide high-quality prediction results to guide the learning of the student model. The student model is generally a model with a simpler structure and fewer parameters, which can improve its own performance by imitating the teacher model and is more efficient than the teacher model.
[0019] In order to better enable the student model to learn deeper structures in the teacher model, the present application provides a method for optimizing the student model based on distillation learning, hereinafter referred to as the student model optimization method for short. It deeply considers the influence of the attention layer on distillation learning, and optimizes the student model by using the difference between the attention layer of the teacher model and the attention layer of the student model, improving the optimization performance of the student model. For details, please refer to Figure 1 , Figure 1 is a schematic flowchart of an exemplary embodiment of the student model optimization method shown in the present application.
[0020] The execution subject of the student model optimization method can be a terminal device, a server or other processing devices. Among them, the terminal device can be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The execution subject of the student model optimization method can also be a student model optimization device. In some possible implementation manners, the student model optimization method can be implemented by a processor calling computer-readable instructions stored in a memory.
[0021] Specifically, the student model optimization method of this embodiment includes the following steps:
[0022] S110: Input the data to be processed into the student model to obtain the query vector, key vector, and value vector of each first attention layer in each first coding module.
[0023] The data to be processed can be a data to be processed selected from a data set. Exemplarily, it can be any data in the data set, or the data with the best quality in the data set. The embodiments of the present application do not limit this. It should be noted that a data to be processed can include, but is not limited to, image data, text data, etc.
[0024] The student model refers to the model that needs to be optimized. In the embodiments of the present application, the process of model optimization is mainly based on distillation learning using the teacher model. The teacher model refers to the model pre-trained through a data set. Based on the teacher model, a student model with fewer parameters is initialized, and then the student model is continuously optimized through the teacher model.
[0025] The first encoding module is the encoding module in the student model, and there is at least one first encoding module in the student model. Exemplarily, the first encoding module can be an image encoding module for extracting image features of image data; it can also be a text encoding module for extracting text features of text data; or it can be a fusion encoding module for fusing image features and / or text features to obtain fusion features.
[0026] The first encoding module includes at least one first attention layer. When processing each element of the data to be processed, the attention layer can dynamically focus on other relevant elements to obtain the context information of the data to be processed. The architecture of the first encoding module can be based on the Transformer architecture, which includes an embedding layer, an attention layer, a fully connected layer, a linear transformation layer, etc. After the data to be processed is input to the first attention layer of the student model after being processed, linear transformation processing is performed on the data to be processed to obtain the query vector, key vector, and value vector of the data to be processed.
[0027] S120: Determine the first vector relationship value of each first encoding module and / or the second vector relationship value between any two first encoding modules according to at least one of the query vector, key vector, and value vector of each first attention layer.
[0028] The first vector relationship value includes the relationship between the vectors in the first encoding module. Exemplarily, the relationship value between the vectors of each first attention layer in the first encoding module can be obtained first, and then the first vector relationship value of the first encoding module is determined based on the relationship values of each first attention layer. In each first attention layer, the relationship value between the vectors is determined by the query vector, key vector, or value vector in the first attention layer. For example, the first vector relationship value can be determined by the relationship between any two of the query vector, key vector, and value vector, or the first vector relationship value can be determined by the relationship among the query vector, key vector, and value vector.
[0029] The second vector relationship value includes the relationship between vectors of any two first encoding modules in the student model. In the student model, the first attention layers of any two first encoding modules correspond to each other. Therefore, the cross-relationship value of the vectors between the two first encoding modules corresponding to each first attention layer can be obtained first, and then the second vector relationship value between the two first encoding modules can be determined based on the cross-relationship values of each first attention layer.
[0030] S130: Input the data to be processed into the teacher model to obtain the query vector, key vector, and value vector of each second attention layer in each second encoding module.
[0031] The teacher model can be a pre-trained model for optimizing the student model. The model complexity of the teacher model is greater than that of the student model. In some embodiments, the teacher model can optimize only one first encoding module in the student model, or can optimize two or more first encoding modules in the student model. For example, when the student model includes three first encoding modules, which can be an image encoding module, a text encoding module, and a fusion encoding module respectively, the teacher model can optimize only the image encoding module; or only the text encoding module or the fusion encoding module, which is not specifically limited herein, or can optimize two or three of them.
[0032] The second encoding module is an encoding module in the teacher model, and the teacher model includes at least one second encoding module. Exemplarily, the second encoding module can be an image encoding module for extracting image features of image data; or a text encoding module for extracting text features of text data; or a fusion encoding module for fusing image features and / or text features to obtain fusion features.
[0033] The second encoding module includes at least one second attention layer. The architecture of the second encoding module can also be based on the Transformer architecture, which includes an embedding layer, an attention layer, a fully connected layer, a linear transformation layer, etc. After the data to be processed is input into the second attention layer of the teacher model after being processed, linear transformation processing is performed on the data to be processed to obtain the query vector, key vector, and value vector of the data to be processed. It should be noted that the student model is a model with a smaller number of parameters, and there are differences between the first attention layer of the student model and the second attention layer of the teacher model. Therefore, in the process of model optimization of the student model, it is necessary to learn the performance of the second attention layer of the teacher model and improve the first attention layer of the student model, so as to ensure that on the basis of not being weaker than the teacher model in performance, it can also be more efficient and simple than the teacher model.
[0034] S140: Determine the third vector relationship value of each second encoding module and / or the fourth vector relationship value between any two second encoding modules according to at least one of the query vector, key vector, and value vector of each second attention layer.
[0035] The third vector relationship value includes the relationship between vectors in the second encoding module. Exemplarily, the relationship value between vectors of each second attention layer in the second encoding module can be obtained first, and then the third vector relationship value of the second encoding module can be determined based on the relationship values of each second attention layer. In each second attention layer, the relationship value between vectors is determined by the query vector, key vector, or value vector in the second attention layer. For example, the third vector relationship value can be determined by the relationship between any two of the query vector, key vector, and value vector, or the third vector relationship value can be determined by the relationship among the query vector, key vector, and value vector.
[0036] The fourth vector relationship value includes the relationship between vectors between any two second encoding modules in the teacher model. In the teacher model, the second attention layers of any two second encoding modules correspond to each other. Therefore, the cross-relationship value of vectors between the two second encoding modules corresponding to each second attention layer can be obtained first, and then the fourth vector relationship value between the two second encoding modules can be determined based on the cross-relationship values of each second attention layer.
[0037] It should be noted that the time when the data to be processed is input into the teacher model and the student model respectively can be that the student model is first, or the teacher model is first, or the teacher model and the student model can be input simultaneously, which is not specifically limited here.
[0038] S150: Determine the first relationship difference between the first vector relationship value of each first encoding module in the student model and the third vector relationship value of the corresponding second encoding module in the teacher model according to the attribute information of the encoding module, and / or determine the second relationship difference between the second vector relationship value in the student model and the fourth vector relationship value of the corresponding encoding module in the teacher model according to the attribute information of the encoding module.
[0039] The attribute information may include the functional attributes of the encoding module, such as the encoding module for processing images and the encoding module for processing text. Before determining the first relationship difference, first determine the corresponding relationship between the first encoding module in the student model and the second encoding module in the teacher model according to the attribute information of the encoding module. For example, the first encoding module for processing images in the student model is determined to be in a corresponding relationship with the second encoding module for processing images in the teacher model, and the first encoding module for processing text in the student model is determined to be in a corresponding relationship with the second encoding module for processing text in the teacher model; then, the first relationship difference between the two is determined according to the first vector relationship value of the corresponding first encoding module and the third vector relationship value of the second encoding module.
[0040] The first relational difference is used to measure the amount of information loss of the first vector relational value compared to the corresponding third vector relational value, that is, the amount of information loss of the first encoding module of the student model compared to the corresponding second encoding module of the teacher model. The first relational difference is the difference between a single first encoding module in the student model and the corresponding second encoding module in the teacher model, and can learn the excellent performance of each second encoding module of the teacher model.
[0041] The second relational difference is usually used when there are multiple encoding modules in both the student model and the teacher model, and is used to measure the amount of information loss of the second vector relational value compared to the corresponding fourth vector relational value. The second relational difference is the difference between two first encoding modules in the student model and the corresponding two second encoding modules in the teacher model, from which the implicit relationship between different second encoding modules in the teacher model can be learned. When the two encoding modules belong to encoding modules with different functional attributes, for example, one is an image encoding module and the other is a text encoding module, the implicit knowledge between the image encoding module and the text encoding module in the teacher model can be fully learned, improving the cross-modal processing ability of the student model.
[0042] S160: Determine the target loss value of the student model according to the first relational difference and / or the second relational difference.
[0043] The target loss value is used to measure whether the student model meets the preset requirements. Exemplarily, in addition to determining the target loss value of the student model according to the first relational difference and / or the second relational difference, it can also be determined according to the first loss value of the loss function between the prediction result of the student model and the prediction result of the teacher model, and the second loss value of the loss function between the prediction result of the student model and the actual result. As an example, it can be determined as the target loss value based on the first relational difference, the first loss value, and the second loss value. As another example, it can also be determined as the target loss value based on the second relational difference, the first loss value, and the second loss value. As another example, it can also be determined as the target loss value based on the first relational difference, the second relational difference, the first loss value, and the second loss value. Among them, weights can be set to adjust each parameter to obtain the target loss value.
[0044] S170: Optimize the student model according to the target loss value to obtain an optimized student model.
[0045] After obtaining the target loss value, the student model optimization device adjusts the student model with the aim of reducing the target loss value to obtain an optimized student model; then processes the next data to be processed with the optimized student model until the student model meets the preset requirements, that is, the prediction accuracy of the student model reaches the requirements.
[0046] It can be seen that the student model optimization method based on distillation learning according to the embodiments of the present application determines the first vector relationship value of each first coding module and / or the second vector relationship value between any two first coding modules according to at least one of the query vector, key vector, and value vector of each attention layer in each first coding module of the student model; determines the third vector relationship value of each second coding module and / or the fourth vector relationship value between any two second coding modules according to at least one of the query vector, key vector, and value vector of each attention layer in each second coding module of the teacher model, thereby not only being able to fully learn the vector relationship of the attention layer inside the model, but also being able to learn the vector relationship between multiple coding modules in the same model; then determines the first relationship difference according to the first vector relationship value and the third vector relationship value and / or determines the second relationship difference according to the second vector relationship value and the fourth vector relationship value; determines the target loss value of the student model according to the first relationship difference and / or the second relationship difference, and optimizes the student model according to the target loss value, thereby optimizing the student model through the difference between the student model and the teacher model at the attention layer, and improving the performance of the student model.
[0047] To clearly illustrate the student model optimization method according to the embodiments of the present application, various existing situations will be described:
[0048] As an example, the data to be processed is input into the student model to obtain the query vector, key vector, and value vector of each first attention layer in each first coding module; the first vector relationship value of each first coding module is determined according to at least one of the query vector, key vector, and value vector of each first attention layer; the data to be processed is input into the teacher model to obtain the query vector, key vector, and value vector of each second attention layer in each second coding module; the third vector relationship value of each second coding module is determined according to at least one of the query vector, key vector, and value vector of each second attention layer; the first relationship difference between the first vector relationship value of each first coding module in the student model and the third vector relationship value of the corresponding second coding module in the teacher model is determined according to the attribute information of the coding module; the target loss value of the student model is determined according to the first relationship difference; and the student model is optimized according to the target loss value to obtain the optimized student model.
[0049] Among them, the first vector relationship value can be determined according to the query vectors of each first attention layer, determined according to the key vectors of each first attention layer, determined according to the key vector value vectors of each first attention layer, determined according to the query vectors and key vectors of each first attention layer, determined according to the query vectors and value vectors of each first attention layer, determined according to the key vectors and value vectors of each first attention layer, or determined according to the query vectors, key vectors and value vectors of each first attention layer; similarly, the second vector relationship value can also be determined according to the query vectors of each second attention layer, determined according to the key vectors of each second attention layer, determined according to the key vector value vectors of each second attention layer, determined according to the query vectors and key vectors of each second attention layer, determined according to the query vectors and value vectors of each second attention layer, determined according to the key vectors and value vectors of each second attention layer, or determined according to the query vectors, key vectors and value vectors of each second attention layer.
[0050] As another example, the data to be processed is input into the student model to obtain the query vectors, key vectors and value vectors of each first attention layer in each first encoding module; the second vector relationship value between any two first encoding modules is determined according to at least one of the query vectors, key vectors and value vectors of each first attention layer; the data to be processed is input into the teacher model to obtain the query vectors, key vectors and value vectors of each second attention layer in each second encoding module; the fourth vector relationship value between any two second encoding modules is determined according to at least one of the query vectors, key vectors and value vectors of each second attention layer; the second relationship difference between the second vector relationship value in the student model and the corresponding fourth vector relationship value in the teacher model is determined according to the attribute information of the encoding module; the target loss value of the student model is determined according to the second relationship difference; the student model is optimized according to the target loss value to obtain the optimized student model.
[0051] As another example, the data to be processed is input into the student model to obtain the query vectors, key vectors and value vectors of each first attention layer in each first encoding module; the first vector relationship value of each first encoding module is determined according to at least one of the query vectors, key vectors and value vectors of each first attention layer, and the second vector relationship value between any two first encoding modules is determined according to at least one of the query vectors, key vectors and value vectors of each first attention layer; the data to be processed is input into the teacher model to obtain the query vectors, key vectors and value vectors of each second attention layer in each second encoding module; the third vector relationship value of each second encoding module is determined according to at least one of the query vectors, key vectors and value vectors of each second attention layer, and the fourth vector relationship value between any two second encoding modules is determined according to at least one of the query vectors, key vectors and value vectors of each second attention layer;
[0052] Among them, the first relationship difference between the first vector relationship value of each first coding module in the student model and the third vector relationship value of the corresponding second coding module in the teacher model can be determined according to the attribute information of the coding module; the target loss value of the student model can be determined according to the first relationship difference; the student model can be optimized according to the target loss value to obtain an optimized student model. Alternatively, the second relationship difference between the second vector relationship value in the student model and the fourth vector relationship value of the corresponding one in the teacher model can be determined according to the attribute information of the coding module; the target loss value of the student model can be determined according to the second relationship difference; the student model can be optimized according to the target loss value to obtain an optimized student model. Further alternatively, the first relationship difference between the first vector relationship value of each first coding module in the student model and the third vector relationship value of the corresponding second coding module in the teacher model can be determined according to the attribute information of the coding module, and the second relationship difference between the second vector relationship value in the student model and the fourth vector relationship value of the corresponding one in the teacher model can be determined according to the attribute information of the coding module; the target loss value of the student model can be determined according to the first relationship difference and the second relationship difference; the student model can be optimized according to the target loss value to obtain an optimized student model.
[0053] Based on the above embodiments, the embodiments of the present application use Figure 2 The flowchart details how to obtain the first vector relationship value of the first coding module and the second vector relationship value between any two first coding modules. Please refer to Figure 2 , Figure 2 which Figure 1 is a schematic flowchart of an exemplary embodiment of step S120 in the student model optimization method shown. Specifically, the process of step S120 for determining the first vector relationship value of each first coding module and / or the second vector relationship value between any two first coding modules according to at least one of the query vector, key vector, and value vector of each first attention layer includes the following steps:
[0054] S210: Determine the first target vector similarity of each first attention layer according to the vector similarity between the query vector, key vector, and value vector of each first attention layer in the first coding module.
[0055] After the student model optimization device inputs the data to be processed into the student model, it extracts the query vector, key vector, and value vector of each first attention layer in each first coding module; then obtains the vector similarity between the query vector, key vector, and value vector, and the vector similarity characterizes the vector relationship between the query vector, key vector, and value vector; and determines the first target vector similarity of each first attention layer according to the vector similarity.
[0056] In some embodiments, the first target vector similarity may be determined by the vector similarity between the query vector and the key vector, or by the vector similarity between the query vector and the value vector, or by the vector similarity between the key vector and the value vector.
[0057] In some other embodiments, the first target vector similarity can be determined by the vector similarity between the query vector, the key vector and the value vector. Specifically, at each first attention layer, the sum of the vector similarities between the query vector and the query vector, the query vector and the key vector, and the query vector and the value vector is calculated to obtain the vector similarity of the query vector; at each first attention layer, the sum of the vector similarities between the key vector and the query vector, the key vector and the key vector, and the key vector and the value vector is calculated to obtain the vector similarity of the key vector; at each first attention layer, the sum of the vector similarities between the value vector and the query vector, the value vector and the key vector, and the value vector and the value vector is calculated to obtain the vector similarity of the value vector; the sum of the vector similarities between the vector similarity of the query vector, the vector similarity of the key vector, and the vector similarity of the value vector in the same first attention layer is determined as the first target vector similarity of the corresponding first attention layer.
[0058] For each second attention layer of each second encoding module in the teacher model, the target vector similarity of each second attention layer can also be calculated according to the above method.
[0059] For ease of understanding, the query vector can be represented as Q, the key vector can be represented as K, and the value vector can be represented as V. QKV is extracted from each first attention layer. Then, at each first attention layer, the sum of the vector similarities between Q and QKV is calculated to obtain the vector similarity of the query vector, which can be expressed as ; Calculate the sum of the vector similarities between K and QKV to obtain the vector similarity of the key vector, which can be expressed as ; Calculate the sum of the vector similarities between V and QKV to obtain the vector similarity of the value vector, which can be expressed as ;Will , and The sum between them is determined as the first target vector similarity corresponding to the first attention layer.
[0060] S220: Sum the similarities of the first target vectors of each first attention layer to obtain the first vector relationship value of the first encoding module.
[0061] After obtaining the first target vector similarity of each first attention layer in the first encoding module, the student model optimization device sums up the first target vector similarities of all the first attention layers in the first encoding module to obtain the first vector relationship value of the first encoding module. In some other embodiments, the student model optimization device may also first sum up the query vectors in each first attention layer of the first encoding module to obtain a target query vector; sum up the key vectors in each first attention layer of the first encoding module to obtain a target key vector; sum up the value vectors in each first attention layer of the first encoding module to obtain a target value vector; and then calculate the vector similarity based on the target query vector, the target key vector, and the target value vector to obtain the first vector relationship value of the first encoding module.
[0062] Exemplarily, on each first attention layer, the process of calculating the sum of the vector similarities between the query vector and the query vector, the query vector and the key vector, and the query vector and the value vector to obtain the vector similarity of the query vector includes: respectively calculating the vector products between the transpose of the query vector and the query vector, the key vector, and the value vector; respectively performing activation processing on each vector product to obtain the vector similarity between the query vector and the query vector, the vector similarity between the query vector and the key vector, and the vector similarity between the query vector and the value vector; and summing up the vector similarity between the query vector and the query vector, the vector similarity between the query vector and the key vector, and the vector similarity between the query vector and the value vector to obtain the vector similarity of the query vector. Thus, the relationships between vectors can be learned through simple vector similarity.
[0063] On each first attention layer, the process of calculating the sum of the vector similarities between the key vector and the query vector, the key vector and the key vector, and the key vector and the value vector to obtain the vector similarity of the key vector; and the process of calculating the sum of the vector similarities between the value vector and the query vector, the value vector and the key vector, and the value vector and the value vector to obtain the vector similarity of the value vector can refer to the above calculation process of the query vector.
[0064] In some other embodiments, the methods for calculating the vector similarities between the query vector and the query vector, the query vector and the key vector, and the query vector and the value vector may also include but are not limited to cosine similarity, Euclidean distance, or Hamming distance, etc.
[0065] Exemplarily, the calculation of the vector similarity between any two vectors satisfies the following formula:
[0066]
[0067] Where represents the layer of the attention layer, the vector and the vector The vector similarity between; denotes the softmax activation function; denotes the transpose of a certain vector of the th layer attention layer, belongs to [1, 3], when ; denotes the transpose of the query vector in the th layer, when ; denotes the transpose of the key vector of the th layer, when ; denotes the transpose of the value vector of the th layer; denotes a certain vector of the th layer attention layer; belongs to [1, 3], when ; denotes the query vector in the th layer, when ; denotes the key vector of the th layer, when ; denotes the value vector of the th layer; denotes the th layer attention layer, denotes the attention layer dimension in the encoding module.
[0068] S230: Determine the second target vector similarity of each first attention layer according to the vector similarity between the query vector, key vector, and value vector of each first attention layer of one of the first encoding modules in any two first encoding modules and the query vector, key vector, and value vector of the corresponding first attention layer of the other first encoding module.
[0069] When there are multiple first encoding modules in the student model, the relationship between any two of them can also be learned. For the convenience of description, any two first encoding modules in the student model can be the first student encoding module and the second student encoding module. Specifically, the student model optimization device inputs the data to be processed into the first student encoding module and the second student encoding module of the student model, and obtains the query vector, key vector, and value vector of each first attention layer in the first student encoding module and the second student encoding module; then obtains the vector similarity between the query vector, key vector, and value vector of each first attention layer in the first student encoding module and the query vector, key vector, and value vector of the corresponding attention layer in the second student encoding module; and determines the second target vector similarity between the first student encoding module and the second student encoding module on each first attention layer according to the vector similarity.
[0070] In some embodiments, the second target vector similarity can be determined by the vector similarity between the query vectors in each first attention layer of the first student encoding module and the query vectors in the corresponding first attention layer of the second student encoding module; it can also be determined by the vector similarity between the key vectors in each first attention layer of the first student encoding module and the key vectors in the corresponding first attention layer of the second student encoding module; it can also be determined by the vector similarity between the value vectors in each first attention layer of the first student encoding module and the value vectors in the corresponding first attention layer of the second student encoding module; it can also be determined by the vector similarity between the query vectors, key vectors, and value vectors in each first attention layer of the first student encoding module and the query vectors, key vectors, and value vectors in the corresponding attention layer of the second student encoding module respectively.
[0071] In some other embodiments, the process of determining the second target vector similarity may further include the following steps: on each first attention layer, obtain the sum of the vector similarities between the query vector in one first encoding module and the query vectors, key vectors, and value vectors in another first encoding module to obtain the first cross vector similarity; on each first attention layer, obtain the vector similarity between the key vector in one first encoding module and the query vectors, key vectors, and value vectors in another first encoding module to obtain the second cross vector similarity; on each first attention layer, obtain the vector similarity between the value vector in one first encoding module and the query vectors, key vectors, and value vectors in another first encoding module to obtain the third cross vector similarity; determine the sum of the first cross vector similarity, the second cross vector similarity, and the third cross vector similarity as the second target vector similarity of the corresponding first attention layer in the corresponding two first encoding modules. One of the first encoding modules can be referred to as the first student encoding module, and the other first encoding module can be referred to as the second student encoding module.
[0072] For any two second attention layers of each second encoding module in the teacher model, the target vector similarity between any two second encoding modules on each second attention layer can also be calculated by referring to the above method.
[0073] As an example, the query vector can be represented as Q, the key vector can be represented as K, and the value vector can be represented as V. The QKV is extracted from the first attention layer of each first encoding module. Then, on each first attention layer, the sum of the vector similarities between Q in the first student encoding module and the QKV in the second student encoding module is calculated to obtain the cross-vector similarity of the query vector in the first student encoding module, and the sum of the vector similarities between Q in the second student encoding module and the QKV in the first student encoding module is calculated to obtain the cross-vector similarity of the query vector in the second student encoding module. The sum of the two is the first cross-vector similarity between the first student encoding module and the second student encoding module. The sum of the vector similarities between K in the first student encoding module and the QKV in the second student encoding module is calculated to obtain the cross-vector similarity of the key vector in the first student encoding module, and the sum of the vector similarities between K in the second student encoding module and the QKV in the first student encoding module is calculated to obtain the cross-vector similarity of the key vector in the second student encoding module. The sum of the two is the second cross-vector similarity between the first student encoding module and the second student encoding module. The sum of the vector similarities between V in the first student encoding module and the QKV in the second student encoding module is calculated to obtain the cross-vector similarity of the value vector in the first student encoding module, and the sum of the vector similarities between V in the second student encoding module and the QKV in the first student encoding module is calculated to obtain the cross-vector similarity of the value vector in the second student encoding module. The sum of the two is the third cross-vector similarity between the first student encoding module and the second student encoding module. The sum of the first cross-vector similarity, the second cross-vector similarity, and the third cross-vector similarity is determined as the second target vector similarity.
[0074] If the student model includes an image encoding module, a text encoding module, and a fusion encoding module, in some embodiments, the second target vector similarity between the image encoding module and the text encoding module can be obtained; the second target vector similarity between the image encoding module and the fusion encoding module can also be obtained; and the second target vector similarity between the text encoding module and the fusion encoding module can also be obtained.
[0075] S240: Sum up the second target vector similarities of each first attention layer to obtain the second vector relationship value corresponding to the two first encoding modules.
[0076] After obtaining the second target vector similarity of each first attention layer in any two first encoding modules, the student model optimization device sums up the second target vector similarities of all first attention layers to obtain the second vector relationship value between the corresponding two first encoding modules. In some other embodiments, the student model optimization device can also first sum up the query vectors in each first attention layer of all first encoding modules to obtain a target query vector; sum up the key vectors in each first attention layer of the first encoding module to obtain a target key vector; sum up the value vectors in each first attention layer of the first encoding module to obtain a target value vector; and then calculate the second vector relationship value between any two encoding modules based on the target query vector, the target key vector, and the target value vector.
[0077] Exemplarily, on each first attention layer, the calculation of the vector similarity between the vectors in the first student encoding module and the second student encoding module satisfies the following formula:
[0078]
[0079] Wherein, represents the vector similarity between two vectors, represents the softmax activation function, represents the transpose of a certain vector in the -th attention layer of the first student encoding module, belongs to [1, 3], when is the case, represents the query vector in the -th layer of the first student encoding module, when is the case, represents the key vector in the -th layer of the first student encoding module, when is the case, represents the value vector in the -th layer of the first student encoding module; represents a certain vector in the -th attention layer of the second student encoding module, belongs to [1, 3], when is the case, represents the query vector in the -th layer of the second student encoding module, when is the case, represents the key vector in the -th layer of the second student encoding module, when is the case, represents the value vector in the -th layer of the second student encoding module; represents the Layer attention layer, indicating the dimension of the attention layer in the encoding module; since the dimensions of each first encoding module in the student model may be different, learnable weights and are required to align the dimensions of the first encoding module.
[0080] Among them, the calculation process of the third vector relationship value of the teacher model can refer to the calculation process of the first vector relationship value of the student model; the calculation process of the fourth vector relationship value of the teacher model can refer to the calculation process of the second vector relationship value of the student model, which will not be elaborated here.
[0081] After obtaining the first vector relationship value of the student model and the third vector relationship value of the teacher model, the student model optimization device can use a preset loss function to calculate the loss between the first vector relationship value and the third vector relationship value, obtaining the initial relationship difference between the first vector relationship value and the third vector relationship value; since the teacher model and the student model belong to different models, after obtaining the initial relationship difference, it is also necessary to perform weighted processing on the initial relationship difference according to the alignment weight between the student model and the teacher model to obtain the first relationship difference. Among them, the preset loss function can be kldivergence (Kullback-Leibler Divergence, KL divergence), root mean squared error (Root Mean Squared Error, RMSE), etc.
[0082] In some other embodiments, after obtaining the second vector relationship value of the student model and the fourth vector relationship value of the teacher model, the student model optimization device can also use a preset loss function to calculate the loss between the second vector relationship value and the fourth vector relationship value, obtaining the initial relationship difference between the second vector relationship value and the fourth vector relationship value; since the teacher model and the student model belong to different models, after obtaining the initial relationship difference, it is also necessary to perform weighted processing on the initial relationship difference according to the alignment weight between the student model and the teacher model to obtain the second relationship difference.
[0083] The student model optimization device can determine the target loss value of the student model according to the first relationship difference. Specifically, calculate the sum of the first relationship differences between each first vector relationship value and the corresponding third vector relationship value, and determine the sum of the first relationship differences as the target loss value of the student model; it can also determine the target loss value of the student model according to the second relationship difference. Specifically, calculate the sum of the second relationship differences between each second vector relationship value and the corresponding fourth vector relationship value, and determine the sum of the second relationship differences as the target loss value of the student model; it can also determine the target loss value of the student model according to the first relationship difference and the second relationship difference. Specifically, the sum of the first relationship difference and the second relationship difference can be determined as the target loss value of the student model. In some other embodiments, the target loss value may also include the difference between the prediction result of the student model and the prediction result of the teacher model, as well as the direct difference between the prediction result of the student model and the actual result.
[0084] Finally, as an example, the student model can be a multimodal model, such as GLIP (Grounded Language-Image Pre-training). The structure of distillation learning can refer to Figure 3 , both the student model and the teacher model include an image encoding module, a text encoding module, and a fusion encoding module. At least one first encoding module of the student model includes a student image encoding module, a student text encoding module, and a student fusion encoding module. At least one second encoding module of the teacher model includes a teacher image encoding module, a teacher text encoding module, and a teacher fusion encoding module. The image encoding module can be a swin backbone, the text encoding module can be BERT (Bidirectional Encoder Representations from Transformers), and the fusion encoding module can be DyHead (Dynamic Head). The fusion encoding module is used to output fused image features and fused text features. For the fused image features, ATSS (Adaptive Training Sample Selection) can be used for the training of object detection.
[0085] The data to be processed includes the image data to be processed and the text data to be processed. The image data to be processed is input into the student image encoding module for encoding processing to obtain the query vector, key vector, and value vector of each first attention layer in the student image encoding module, as well as the image features output by the student image encoding module. The text data to be processed is input into the student text encoding module for encoding processing to obtain the query vector, key vector, and value vector of each first attention layer in the student text encoding module, as well as the text features output by the student text encoding module. The image features output by the student image encoding module and the text features output by the student text encoding module are input into the student fusion encoding module to obtain the query vector, key vector, and value vector of each first attention layer in the student fusion encoding module. Before inputting the image features into the fusion encoding module, they can be processed by a feature pyramid network to obtain image features of different scales.
[0086] The image data to be processed is input into the teacher image encoding module for encoding processing to obtain the query vector, key vector, and value vector of each second attention layer in the teacher image encoding module, as well as the image features output by the teacher image encoding module. The text data to be processed is input into the teacher text encoding module for encoding processing to obtain the query vector, key vector, and value vector of each second attention layer in the teacher text encoding module, as well as the text features output by the teacher text encoding module. The image features output by the teacher image encoding module and the text features output by the teacher text encoding module are input into the teacher fusion encoding module to obtain the query vector, key vector, and value vector of each second attention layer in the teacher fusion encoding module.
[0087] After that, the first vector relationship value of the student image encoding module and the first vector relationship value of the student text encoding module are calculated respectively, and the third vector relationship value of the teacher image encoding module and the third vector relationship value of the teacher text encoding module are calculated respectively. Exemplarily, the calculation of the first vector relationship value and the third vector relationship value satisfies the following formula:
[0088]
[0089] where, represents the first vector relationship value or the third vector relationship value; represents the number of layers of the attention layer, belongs to [1, L]; belongs to [1, 3], belongs to [1, 3], representing the query vector, key vector, and value vector respectively; represents the layer of the attention layer, the vector and the vector the vector similarity between them; is a hyperparameter, representing the vector and the vector The weight between
[0090] Calculate the second vector relationship value between the student image encoding module and the student text encoding module, and the fourth vector relationship value between the teacher image encoding module and the teacher text encoding module. Exemplarily, the calculation of the second vector relationship value or the fourth vector relationship value satisfies the following formula:
[0091]
[0092] Where represents the second vector relationship value or the fourth vector relationship value; represents the number of layers of the attention layer, belongs to [1, L]; belongs to [1, 3], belongs to [1, 3], and respectively represents the query vector, the key vector, and the value vector; represents the layer of the attention layer in the image encoding module, the vector and the layer of the attention layer in the text encoding module, the vector the vector similarity between; is a hyperparameter, representing the vector and the vector the weight between.
[0093] For the first vector relationship value of the student fusion encoding module and the third vector relationship value of the teacher fusion encoding module, since there are not only image features but also text features in the fusion encoding module, the value vector is different from that of the image encoding module and the text encoding module. There will be both the image value vector of the image data to be processed and the text value vector of the text data to be processed. Exemplarily, the calculation of the fusion encoding module satisfies the following formula:
[0094]
[0095] Where represents the first vector relationship value of the student fusion encoding module or the third vector relationship value of the teacher fusion encoding module, represents the number of layers of the attention layer of the fusion encoding module, belongs to [1, ; belongs to [1, 4], belongs to [1, 4], and respectively represents the query vector, the key vector, the image value vector, and the text value vector; represents the layer of the attention layer, the vector and the vector the vector similarity between; is a hyperparameter representing the vector and the vector weight between them.
[0096] Furthermore, the calculation of satisfies the following formula:
[0097]
[0098] where, represents the vector similarity between the vector and the vector in the layer attention layer; represents the softmax activation function, represents the transpose of a certain vector in the layer attention layer, belongs to [1, 4], when then represents the transpose of the query vector in the layer, when then represents the transpose of the key vector in the layer, when then represents the transpose of the image value vector in the layer; when then represents the transpose of the text value vector in the layer; represents a certain vector in the layer attention layer; belongs to [1, 4], when then represents the query vector in the layer, when then represents the key vector in the layer, when then represents the image value vector in the layer, when then represents the text value vector in the layer; represents the layer attention layer, represents the attention layer dimension in the fusion coding module.
[0099] After obtaining the first vector relationship value of the student image encoding module, the first vector relationship value of the student text encoding module, the first vector relationship value of the student fusion encoding module, the third vector relationship value of the teacher image encoding module, the third vector relationship value of the teacher text encoding module, and the third vector relationship value of the teacher fusion encoding module, as well as the second vector relationship value between the student image encoding module and the student text encoding module, and the fourth vector relationship value between the teacher image encoding module and the teacher text encoding module, calculate the first relationship difference between the first vector relationship value of the student image encoding module and the third vector relationship value of the teacher image encoding module; calculate the first relationship difference between the first vector relationship value of the student text encoding module and the third vector relationship value of the teacher text encoding module; calculate the first relationship difference between the first vector relationship value of the student fusion encoding module and the third vector relationship value of the teacher fusion encoding module; calculate the second relationship difference between the second vector relationship value and the fourth vector relationship value; combine these four relationship differences to obtain the target loss value of the student model. Exemplarily, its calculation formula satisfies the following formula:
[0100]
[0101] Among them, represents the target loss value, represents the KL divergence calculation, represents the first vector relationship value of the student image encoding module, represents the third vector relationship value of the teacher image encoding module, represents the alignment weight between the student image encoding module and the teacher image encoding module; represents the first vector relationship value of the student text encoding module, represents the third vector relationship value of the teacher text encoding module, represents the alignment weight between the student text encoding module and the teacher text encoding module; represents the second vector relationship value of the student model; represents the fourth vector relationship value of the teacher model, represents the alignment weight between the student model and the teacher model; represents the first vector relationship value of the student fusion encoding module, represents the third vector relationship value of the teacher fusion encoding module, represents the alignment weight between the student fusion encoding module and the teacher fusion encoding module; represents the weight obtained through multi-modal model training.
[0102] Please refer to Figure 4 , Figure 4It is a schematic structural diagram of an exemplary embodiment of the student model optimization device shown in the present application. The student model optimization device 400 includes an input module 410, a relationship value determination module 420, a relationship difference determination module 430, a loss value determination module 440, and an optimization module 450. The input module 410 is configured to input the data to be processed into the student model to obtain the query vector, key vector, and value vector of each first attention layer in each first encoding module, and input the data to be processed into the teacher model to obtain the query vector, key vector, and value vector of each second attention layer in each second encoding module. The relationship value determination module 420 is configured to determine the first vector relationship value of each first encoding module and / or the second vector relationship value between any two first encoding modules according to at least one of the query vector, key vector, and value vector of each first attention layer, and determine the third vector relationship value of each second encoding module and / or the fourth vector relationship value between any two second encoding modules according to at least one of the query vector, key vector, and value vector of each second attention layer. The relationship difference determination module 430 is configured to determine the first relationship difference between the first vector relationship value of each first encoding module in the student model and the third vector relationship value of the corresponding second encoding module in the teacher model according to the attribute information of the encoding module, and / or determine the second relationship difference between the second vector relationship value in the student model and the fourth vector relationship value of the corresponding second encoding module in the teacher model according to the attribute information of the encoding module. The loss value determination module 440 is configured to determine the target loss value of the student model according to the first relationship difference and / or the second relationship difference. The optimization module 450 is configured to optimize the student model according to the target loss value to obtain an optimized student model.
[0103] In the above solution, the student model optimization device determines the first vector relationship value of each first encoding module and / or the second vector relationship value between any two first encoding modules according to at least one of the query vector, key vector, and value vector of each attention layer in each first encoding module of the student model; determines the third vector relationship value of each second encoding module and / or the fourth vector relationship value between any two second encoding modules according to at least one of the query vector, key vector, and value vector of each attention layer in each second encoding module of the teacher model. Thus, it can not only fully learn the vector relationship of the attention layer inside the model, but also learn the vector relationship between multiple encoding modules in the same model. Then, it determines the first relationship difference according to the first vector relationship value and the third vector relationship value and / or determines the second relationship difference according to the second vector relationship value and the fourth vector relationship value; determines the target loss value of the student model according to the first relationship difference and / or the second relationship difference, and optimizes the student model according to the target loss value. Thus, the student model is optimized through the difference between the student model and the teacher model at the attention layer, improving the performance of the student model.
[0104] Among them, the functions of each module can be referred to in the embodiments of the student model optimization method, which will not be elaborated here.
[0105] To implement the student model optimization method of the above embodiments, the present application proposes another electronic device. For details, please refer to Figure 5 , Figure 5 is a schematic structural diagram of an embodiment of the electronic device provided by the present application.
[0106] The electronic device 500 includes a memory 510 and a processor 520. Among them, the memory 510 and the processor 520 are coupled.
[0107] The memory 510 is used to store program data, and the processor 520 is used to execute the program data to implement the student model optimization method of the above embodiments.
[0108] In this embodiment, the processor 520 can also be referred to as a CPU (Central Processing Unit). The processor 520 may be an integrated circuit chip with signal processing capabilities. The processor 520 may also be a general-purpose processor, a digital signal processor (Digital Signal Processing, DSP), an application-specific integrated circuit (Application-Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 520 may also be any conventional processor, etc.
[0109] The present application also provides a computer-readable storage medium, such as Figure 6 shown, the computer-readable storage medium 600 is used to store program data 610. When the program data 610 is executed by the processor, it is used to implement the student model optimization method in the method embodiments of the present application.
[0110] In the embodiments of the method for optimizing the student model of the present application, when the method involved exists in the form of a software functional unit and is sold or used as an independent product, it can be stored in a device, such as a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0111] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A student model optimization method based on distillation learning, characterized in that: The student model optimization method based on distillation learning is applied to a student model optimization device, the student model optimization device includes a student model and a teacher model, the student model includes at least one first encoding module, each of the first encoding modules includes at least one first attention layer, the teacher model includes at least one second encoding module, each of the second encoding modules includes at least one second attention layer, the method includes: Inputting the data to be processed into the student model to obtain a query vector, a key vector and a value vector of each first attention layer in each first encoding module, wherein the data to be processed includes image data to be processed and text data to be processed; Determine the first vector relationship value of each of the first coding modules and / or the second vector relationship value between any two first coding modules according to at least one of the query vector, key vector and value vector of each first attention layer, wherein the second target vector similarity of each first attention layer is determined according to the vector similarity between the query vector, key vector and value vector of each first attention layer of one first coding module in any two first coding modules and the query vector, key vector and value vector of the corresponding first attention layer of another first coding module; sum the second target vector similarities of each first attention layer to obtain the second vector relationship value of the corresponding two first coding modules; Inputting the data to be processed into the teacher model to obtain a query vector, a key vector and a value vector of each second attention layer in each second encoding module; Determine a third vector relationship value of each of the second encoding modules and / or a fourth vector relationship value between any two second encoding modules according to at least one of the query vector, the key vector, and the value vector of each second attention layer; Determine, according to the attribute information of the encoding module, a first relationship difference between a first vector relationship value of each of the first encoding modules in the student model and a third vector relationship value of the corresponding second encoding module in the teacher model, and / or determine, according to the attribute information of the encoding module, a second relationship difference between a second vector relationship value in the student model and a fourth vector relationship value in the teacher model; Determine a target loss value of the student model according to the first relationship difference and / or the second relationship difference; The student model is optimized according to the target loss value to obtain an optimized student model.
2. The student model optimization method based on distillation learning according to claim 1, characterized in that: The step of determining the first vector relationship value of each first encoding module according to at least one of the query vector, the key vector and the value vector of each first attention layer comprises: Determining first target vector similarities of each first attention layer according to vector similarities between query vectors, key vectors, and value vectors of each first attention layer in the first encoding module; The first target vector similarities of each first attention layer are summed to obtain the first vector relationship value of the first encoding module.
3. The student model optimization method based on distillation learning according to claim 2, characterized in that: The step of determining the similarity of the first target vectors of each first attention layer according to the vector similarities between the query vector, the key vector and the value vector of each first attention layer in the first encoding module comprises: At each first attention layer, calculating the sum of vector similarities between the query vector and the query vector, between the query vector and the key vector, and between the query vector and the value vector to obtain the vector similarity of the query vector; At each first attention layer, calculate the sum of vector similarities between the key vector and the query vector, between the key vector and the key vector, and between the key vector and the value vector to obtain the vector similarity of the key vector; At each first attention layer, calculate the sum of vector similarities between the value vector and the query vector, between the value vector and the key vector, and between the value vector and the value vector to obtain the vector similarity of the value vector; The sum of the vector similarities of the query vector, the key vector and the value vector in the same first attention layer is determined as the first target vector similarity corresponding to the first attention layer.
4. The student model optimization method based on distillation learning according to claim 3, characterized in that: The step of calculating the sum of vector similarities between the query vector and the query vector, between the query vector and the key vector, and between the query vector and the value vector at each first attention layer to obtain the vector similarity of the query vector comprises: calculating the vector product between the transpose of the query vector and the query vector, the key vector and the value vector respectively; Perform activation processing on each vector product respectively to obtain the vector similarity between the query vector and the query vector, the vector similarity between the query vector and the key vector, and the vector similarity between the query vector and the value vector; The vector similarity between the query vector and the query vector, the vector similarity between the query vector and the key vector, and the vector similarity between the query vector and the value vector are summed to obtain the vector similarity of the query vector.
5. The student model optimization method based on distillation learning according to claim 1, characterized in that: The step of determining the similarity of the second target vectors of each first attention layer according to the vector similarity between the query vector, the key vector and the value vector of each first attention layer of one first encoding module in any two first encoding modules and the query vector, the key vector and the value vector of the corresponding first attention layer of the other first encoding module comprises: At each first attention layer, obtaining the sum of vector similarities between a query vector in a first encoding module and a query vector, a key vector, and a value vector in another first encoding module to obtain a first cross vector similarity; At each first attention layer, obtaining a vector similarity between a key vector in a first encoding module and a query vector, a key vector, and a value vector in another first encoding module to obtain a second cross vector similarity; At each first attention layer, obtaining a vector similarity between a value vector in a first encoding module and a query vector, a key vector, and a value vector in another first encoding module to obtain a third cross vector similarity; The sum of the first cross-vector similarity, the second cross-vector similarity and the third cross-vector similarity is determined as the second target vector similarity corresponding to the first attention layer in the two first encoding modules.
6. The student model optimization method based on distillation learning according to claim 1, characterized in that: At least one first encoding module of the student model includes a student image encoding module, a student text encoding module and a student fusion encoding module. The step of inputting the data to be processed into the student model to obtain the query vector, key vector and value vector of each first attention layer in each first encoding module includes: Input the image data to be processed into the student image encoding module for encoding processing, and obtain the query vector, key vector and value vector of each first attention layer in the student image encoding module, as well as the image features output by the student image encoding module; Input the text data to be processed into the student text encoding module for encoding processing, and obtain the query vector, key vector and value vector of each first attention layer in the student text encoding module, as well as the text features output by the student text encoding module; The image features output by the student image encoding module and the text features output by the student text encoding module are input into the student fusion encoding module to obtain the query vector, key vector and value vector of each first attention layer in the student fusion encoding module.
7. The student model optimization method based on distillation learning according to claim 1, characterized in that: The step of determining the first relationship difference between the first vector relationship value of each of the first encoding modules in the student model and the third vector relationship value of the corresponding second encoding module in the teacher model according to the attribute information of the encoding module comprises: Using a preset loss function to perform loss calculation on the first vector relationship value and the third vector relationship value to obtain an initial relationship difference between the first vector relationship value and the third vector relationship value; The initial relationship difference is weighted according to the alignment weight between the student model and the teacher model to obtain the first relationship difference.
8. The student model optimization method based on distillation learning according to claim 1, characterized in that: The step of determining the target loss value of the student model according to the first relationship difference and / or the second relationship difference comprises: Determine the target loss value of the student model according to the first relationship difference, or, The target loss value of the student model is determined according to the second relationship difference, or, A target loss value for the student model is determined according to the first relationship difference and the second relationship difference.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: include: Program data is stored, and when the program data is executed by a processor, it is used to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image text model processing method and image text retrieval system
CN116521833A
Zero sample image classification system and method
CN119295791A
Target detection model robustness test method based on cross-task learning
CN119357066A