Knowledge distillation method and electronic device

By unifying the dimensions of the splicing matrix of the Transformer network and combining it with multiple distillation loss, the problem of inconsistent head numbers in the multi-head attention mechanism is solved, enabling full transfer and absorption of knowledge and improving the training accuracy and resource utilization efficiency of the student network.

CN120932073BActive Publication Date: 2025-12-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511449445.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-12-12
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

In resource-constrained scenarios, the inconsistent number of heads in the multi-head attention mechanism of Transformer-based neural network models leads to knowledge loss, which affects the student network's full learning and absorption of the teacher network's knowledge and makes it difficult to effectively perform knowledge distillation.

Method used

By unifying the splicing matrix dimensions of the teacher network and the student network, and combining multiple distillation loss and task loss, the model parameters of the student network are adjusted, thus solving the problem of inconsistent attention map numbers and achieving full knowledge transfer and absorption.

Benefits of technology

It improves the training accuracy and knowledge distillation effect of student networks, making it suitable for application in edge devices and reducing computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932073B_ABST
    Figure CN120932073B_ABST
Patent Text Reader

Abstract

The application provides a knowledge distillation method and an electronic device, which can be applied to the field of artificial intelligence technology. The knowledge distillation method comprises the following steps: determining a plurality of first attention maps for a teacher network and a plurality of second attention maps for a student network respectively; performing dimension normalization on a first splicing matrix obtained by splicing the plurality of first attention maps and a second splicing matrix obtained by splicing the plurality of second attention maps, to obtain a first attention matrix and a second attention matrix with the same dimension; determining a multiple distillation loss for knowledge distillation from the first attention map to the second attention map according to the matrix features of the first attention matrix and the second attention matrix; and in the training process of the student network, adjusting the model parameters of the student network according to the multiple distillation loss and the task loss of the student network until all samples in a sample set used for training the student network are polled, to obtain a target student network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a knowledge distillation method and an electronic device. BACKGROUND

[0002] In the current era of rapid development of artificial intelligence, neural network models based on self-attention mechanism have been widely used in natural language processing, computer vision, speech recognition, image processing and many other fields due to their feature extraction and processing capabilities. However, the running process of the neural network model requires a large amount of storage and computing resources. The knowledge distillation method can solve these problems. Knowledge distillation aims to extract knowledge from a superior but complex teacher network and pass it to a relatively simple student network, thereby reducing the storage and computing overhead of the model while ensuring certain performance. SUMMARY

[0003] In view of the above problems, the present application provides a knowledge distillation method and an electronic device for improving the prediction accuracy of the student network and the knowledge distillation effect.

[0004] One aspect of the present application provides a knowledge distillation method, which comprises: determining a plurality of first attention maps for a teacher network and a plurality of second attention maps for a student network, respectively; performing dimension normalization on a first splicing matrix obtained by splicing the plurality of first attention maps and a second splicing matrix obtained by splicing the plurality of second attention maps, to obtain first attention matrices and second attention matrices with the same dimension; determining a multiple distillation loss for knowledge distillation from the first attention maps to the second attention maps according to the matrix features of the first attention matrices and the matrix features of the second attention matrices; and adjusting the model parameters of the student network according to the multiple distillation loss and the task loss of the student network during the training process of the student network until all samples in a sample set used for training the student network are polled, to obtain a target student network, the sample set comprising training samples and sample labels, and the task loss being used to describe the difference between the output result of the student network according to the training samples and the sample labels.

[0005] Another aspect of the present application also provides an electronic device comprising: one or more processors; a memory for storing one or more computer programs, the one or more processors executing the one or more computer programs to implement the steps of the above knowledge distillation method.

[0006] According to the embodiment of the present application, by respectively determining the first attention map for the teacher network and the second attention map for the student network, and performing dimension normalization on the first splicing matrix based on the plurality of first attention maps and the second splicing matrix based on the plurality of second attention maps, the first attention matrix and the second attention matrix with the same dimension are obtained, the multiple distillation loss for knowledge distillation from the first attention map to the second attention map is determined according to the matrix features of the first attention matrix and the matrix features of the second attention matrix, the model parameters of the student network are adjusted according to the multiple distillation loss and the task loss of the student network training, and the target student network is obtained. Since the dimensions of the splicing matrices of the teacher network and the student network are unified in the process of knowledge distillation, the attention matrices with the same dimension are obtained, and the multiple distillation loss is calculated according to the matrix features of the attention matrices after dimension unification; then the student network is trained by using the multiple distillation loss and the task loss, which effectively solves the knowledge loss problem caused by the inconsistency of the number of heads of the multi-head attention mechanism of the image retrieval model of the teacher network and the image retrieval model of the student network in the related technology, so that the student network can fully learn and absorb the knowledge of the teacher network, and thus the retrieval accuracy of the image retrieval model of the student network is improved. BRIEF DESCRIPTION OF DRAWINGS

[0007] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:

[0008] Figure 1 A scene diagram of application of the knowledge distillation method according to the embodiment of the present application is shown;

[0009] Figure 2 A flowchart of the knowledge distillation method according to the embodiment of the present application is shown.

[0010] Figure 3 An architecture diagram of the knowledge distillation according to the embodiment of the present application is shown.

[0011] Figure 4 An architecture diagram of the image retrieval model according to the embodiment of the present application is shown.

[0012] Figure 5 A structure block diagram of the knowledge distillation device according to the embodiment of the present application is shown.

[0013] Figure 6 A block diagram of the electronic device suitable for implementing the knowledge distillation method according to the embodiment of the present application is shown schematically. DETAILED DESCRIPTION

[0014] Embodiments of the present application will be described below with reference to the accompanying drawings. It should be understood, however, that the description which follows is illustrative only and is not intended to limit the scope of the present application. In the detailed description, procedures, apparatuses, and methods that are well known and commonly used in the art will not be described in detail in order to avoid obscuring the concept of the present application.

[0015] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "includes" and tautological equivalents thereof, means the inclusion of the stated features, steps, operations, and / or components but not to the exclusion of one or more other features, steps, operations, or components.

[0016] All terms used herein including technical and scientific terms have the same meaning as commonly understood by one of ordinary skill in the art unless otherwise defined. It should be noted that the terms used herein are defined as consistent with the context where used and should not be construed as ideally or over-abstractly.

[0017] In the case where expressions similar to "at least one of A, B, and C, etc." are used, it should generally be interpreted to include at least one of A, B, and C, etc. (e.g., "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).

[0018] Currently, Transformer-based neural network models are applied in many fields, such as natural language processing for text generation, machine translation, and other tasks; in the field of computer vision, image classification, object detection, and other functions. These applications have promoted the intelligent development of various industries and improved production efficiency and service quality. Given the problem that the operation of a Transformer-based neural network model will occupy a large amount of storage and computing resources, knowledge distillation can solve this problem. However, in the knowledge distillation process, the number of heads in the multi-head attention mechanism of the teacher network is often greater than the number of heads in the multi-head attention mechanism of the student network. The multi-head attention mechanism is the core component of the Transformer model, and different numbers of heads play different roles in capturing features and representing information. The difference in the number of heads makes it difficult to directly construct a distillation loss, because it is not possible to simply one-to-one match and compare the multi-head attention output of the teacher network with the corresponding output of the student network, and current technologies do not focus on the multi-head attention mechanism of the Transformer architecture core, do not consider the problem of inconsistent attention maps caused by the difference in the number of attention heads between the teacher and student networks, do not design a grouping strategy to establish a precise correspondence of attention maps, and cannot solve the problem of matching attention maps in the "many-to-few" or "few-to-many" scenario, which easily loses key knowledge in the attention layer during distillation, resulting in difficulty in effectively conducting knowledge distillation and hindering the full learning and absorption of knowledge from the teacher network by the student network.

[0019] Therefore, there is an urgent need for a technical solution to overcome the above technical problems. The knowledge distillation method provided by the embodiments of the present application can unify the dimensions of the splicing matrices of the teacher network and the student network, and combine multiple distillation losses and task losses to jointly train the student network, solve the problem of knowledge loss, improve the training accuracy and knowledge distillation effect of the student network, and provide a new solution for the efficient application of Transformer-based neural network models in resource-constrained scenarios.

[0020] Figure 1 An application scenario diagram of the knowledge distillation method according to an embodiment of the present application is shown.

[0021] As Figure 1 shown, the application scenario of this embodiment can include a teacher network and a student network. The teacher network includes a plurality of first attention maps 101-1, and the student network includes a plurality of second attention maps 102-1. The application scenario can also include a first splicing matrix 101-2, a second splicing matrix 102-2, a first attention matrix 101-3, a second attention matrix 102-3, a multiple distillation loss 103, a training sample 104, an output result 105, a sample label 106, a task loss 107, and a target loss 108 obtained from the multiple distillation loss 103 and the task loss 107 of the student network.

[0022] The teacher network comprises a plurality of first attention maps 101-1, and the student network comprises a plurality of second attention maps 102-1. Dimension normalization is performed on a first spliced matrix 101-2 obtained by splicing the plurality of first attention maps 101-1 and a second spliced matrix 102-2 obtained by splicing the plurality of second attention maps 102-1, to obtain first attention matrices 101-3 and second attention matrices 102-3 of the same dimension. According to the matrix characteristics of the first attention matrices 101-3 and the second attention matrices 102-3, a multiple distillation loss 103 for knowledge distillation from the first attention maps to the second attention maps can be determined. By inputting a training sample 104 to the attention maps of the student network, an output result 105 can be obtained. According to the output result 105 and a sample label 106, a task loss 107 of the student network can be calculated. The target loss 108 obtained according to the multiple distillation loss 103 and the task loss 107 can be used to feedback adjust the student network until the training sample 104 is polled, and a target student network is obtained.

[0023] It should be understood that Figure 1 The number of attention maps in the above formula is only illustrative. According to the needs of implementation, there can be any number of attention maps.

[0024] The following will describe the knowledge distillation method of the embodiments of the present application based on the scenario described above, by Figure 1 Figures 2-4 The knowledge distillation method of the embodiments of the present application will be described in detail.

[0025] Figure 2 A flowchart of the knowledge distillation method according to the embodiments of the present application is shown.

[0026] As shown in Figure 2 The knowledge distillation method of the embodiments includes operations S210-S240.

[0027] In operation S210, a plurality of first attention maps for the teacher network and a plurality of second attention maps for the student network are determined respectively.

[0028] In operation S220, dimension normalization is performed on a first spliced matrix obtained by splicing the plurality of first attention maps and a second spliced matrix obtained by splicing the plurality of second attention maps, to obtain first attention matrices and second attention matrices of the same dimension.

[0029] In operation S230, according to the matrix characteristics of the first attention matrices and the matrix characteristics of the second attention matrices, a multiple distillation loss for knowledge distillation from the first attention maps to the second attention maps is determined.

[0030] ​In operation S240, during the training of the student network, the model parameters of the student network are adjusted according to the multiple distillation loss and the task loss of the student network until all samples in the sample set used to train the student network are polled to obtain the target student network. The sample set includes training samples and sample labels. The task loss is used to describe the difference between the output of the student network based on the training samples and the sample labels.

[0031] In some embodiments, the teacher network may include (n can represent the number, L can be an abbreviation for layer, and t can be an abbreviation for teacher.) The student network can include a first attention layer and can include... (s can be an abbreviation for Student) There are two second attention layers. Multiple first attention maps can be determined from the first attention layers. Multiple second attention maps can be determined from the second attention layers. The number of first attention maps can be greater than the number of second attention maps.

[0032] In some embodiments, the same number of students may be selected from both the teacher network and the student network. () When performing "one-to-one" knowledge distillation using the first and second attention layers (which can represent a one-to-one relationship between the teacher network and the student network), the number of these first or second attention layers selected must not exceed the total number of layers in the student network. It also does not exceed the total number of layers in the teacher network. ,Right now , .

[0033] In the knowledge distillation process based on the Transformer architecture, due to the difference in the number of heads in the multi-head attention mechanism between the teacher network and the student network during the one-to-one knowledge distillation between the first and second attention layers, the number of first attention graphs in the first attention layer is inconsistent with the number of second attention graphs in the second attention layer. This leads to a discrepancy between the number of first attention graphs in the first attention layer and the number of second attention graphs in the second attention layer, thus affecting the knowledge distillation process based on the Transformer architecture. Figure 1 The distillation loss in a one-to-one comparison is difficult to apply directly. If forced to use it, information loss will occur due to the mismatch in the number of attention maps, severely impacting the effective learning and absorption of knowledge from the teacher network by the student network, making it difficult to achieve the goal of knowledge distillation guidance or improving student network performance. Therefore, this invention provides a method for achieving knowledge distillation even when the number of attention maps between the teacher network and the student network is inconsistent.

[0034] Next, in order to elaborate on the technical details of the present invention, we will refer to the foregoing selection. Analyze any one of the "one-to-one" correspondence layers. Define this correspondence layer as... The range of values ​​for j is . (That is, j can be 1 to 1) Any integer between these values ​​corresponds to a different "one-to-one" correspondence layer. It should be noted that since all "one-to-one" correspondence layers follow the same processing logic and principles during distillation, therefore... The derived conclusions, methods, and formulas are applicable to all other selected "one-to-one" corresponding layers, without the need for repeated derivation for each layer individually. By analyzing a single layer, the results can be generalized to all "one-to-one" layers participating in the distillation process.

[0035] For any of the "one-to-one" corresponding layers selected above The multi-head attention mechanism of the layers in the teacher network includes Size, correspondingly producing First attention map , , ..., In student networks, the number of heads in a multi-head attention mechanism is... Correspondingly generated The second attention map is .

[0036] Then, the aforementioned teacher network First attention map and student network The second attention map is divided into the same number of ( The group. In knowledge distillation, the teacher network needs to be included. First attention map and student network The second attention map is divided into the same number of... Grouping is used to establish a one-to-one correspondence between groups. Specifically, attention maps can be distributed evenly in order of their numbers, so that each group contains as many attention maps as possible, allowing the last group to have slightly fewer.

[0037] For example: when the teacher network has There are 1 first attention maps, numbered 1.1, 1.2, 1.3, 1.4, 1.5, and 1.6 respectively. The student network has... The two second attention maps are numbered 2.1, 2.2, 2.3, 2.4, 2.5, and 2.6, and are divided into... When grouping, the teacher network is grouped as [1.1,1.2,1.3] and [1.4,1.5,1.6], and the student network is grouped as [2.1,2.2,2.3] and [2.4,2.5,2.6]; when the teacher network has There are 1 first attention maps, numbered 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, and 1.7 respectively. The student network has... The two second attention maps are numbered 2.1, 2.2, 2.3, 2.4, and 2.5, and are divided into... When grouping, the teacher network is grouped as [1.1,1.2,1.3], [1.4,1.5], and [1.6,1.7], while the student network is grouped as [2.1,2.2], [2.3,2.4], and [2.5]. The core of this grouping method is to ensure the number of groups for both the teacher and student networks. Consistency is achieved through intra-group alignment (such as calculating distillation losses) to facilitate knowledge transfer. The specific implementation of group partitioning does not affect the process for determining multiple distillation losses in this embodiment of the invention, and this embodiment is also applicable to... Extreme situations.

[0038] Based on the grouping logic described above, the following section will use a corresponding group from the teacher network and the student network (denoted as...). The range of values ​​for i is 1. Let's take ) as an example for analysis. Since the processing methods for all corresponding groups are exactly the same, therefore, for The analytical conclusions can be directly generalized to other groups.

[0039] In the corresponding group In the diagram, the subgroup of the teacher network contains p first attention maps (i.e., from that layer of the teacher network). The first attention map is assigned to The number is p); the subgroups of the student network contain q second attention maps (i.e., from this layer of the student network). The second attention map is assigned to The quantity is q). The quantities of p and q can be the same or different.

[0040] It needs to be clarified that the dimensions of all attention maps in both the teacher and student networks can be uniform: each attention map can be dimensional. The matrix is ​​a square matrix (n represents the number of tokens input to the training samples of the current layer), because the core function of the attention map is to characterize the correlation strength between n tokens, and the rows and columns of the matrix correspond to different tokens.

[0041] To integrate attention information within a group and align features between the teacher and student networks, the first and second attention maps within the group can be stitched together.

[0042] For the p first attention graphs of the teacher network, they are concatenated sequentially along the column direction (i.e., all columns of the first first attention graph are immediately followed by all columns of the second first attention graph, and so on), forming a new matrix, namely the first concatenation matrix. Since each first attention map is The matrix, after concatenating p first attention maps column-wise, is the first concatenated matrix. The dimension is ,Right now Similarly, for the q second attention graphs of the student network, a second concatenation matrix is ​​formed by concatenating them in the same column-wise manner. Its dimensions are ,Right now This stitching operation integrates information from multiple attention maps within the group into a single matrix, laying the data foundation for subsequent solutions to dimensionality discrepancies and the construction of distillation loss.

[0043] In knowledge distillation, the construction of the loss function requires that the two feature vectors involved in the loss function calculation have the same dimension; otherwise, it is impossible to directly compare the similarity or measure the difference between the two feature vectors. In the intra-group alignment stage of this embodiment, this problem is specifically manifested in the first concatenation matrix of the teacher network subgroup. The second splicing matrix of the student network subgroup The dimensions do not match, the first concatenated matrix The dimension is (n is the number of input tokens, p is the number of first attention maps within the teacher network subgroup); second concatenation matrix The dimension is ( (The number of second attention maps within the student network subgroup). This addresses the issue of differing attention heads between teachers and students. and In the case of inequality, the first concatenation matrix Second splicing matrix number of columns and The differences between the two attention maps prevent them from being directly used to construct the distillation loss. Therefore, it is necessary to normalize the dimensions of the first concatenation matrix obtained by concatenating multiple first attention maps and the second concatenation matrix obtained by concatenating multiple second attention maps to obtain first and second attention matrices with the same dimensions.

[0044] For the first attention matrix and the second attention matrix unified in dimension, a multiple distillation loss can be determined from the first attention map to the second attention map according to the matrix characteristics of the first attention matrix and the second attention matrix, for example, a first multiple distillation loss for describing the difference between the core feature vectors of the first attention matrix and the second attention matrix; a second multiple distillation loss for quantifying the statistical quantity difference between the teacher network and the student network in the graph distribution characteristics; and a third multiple distillation loss for describing the attention distribution difference between the normalized first attention map and the normalized second attention map. The three multiple distillation losses complement each other, can capture the knowledge difference between the teacher and the student network from different aspects, and ensure the comprehensiveness of knowledge transfer.

[0045] In the training process of the student network, a training sample set can be used. The training data set used by the student network applied to different scenes will be different. For natural language related tasks, a text data set can be used, which can be a multi-language corpus, a labeled sentiment analysis data set or a question and answer pair data set, etc. In image processing tasks, an image data set can be used, for example, a detection image data set labeled with target position and category, or a continuous video frame sequence, or an image text pair or a video audio pair.

[0046] Taking the training sample set as an image text pair for example, the training sample set can include training samples and sample labels, for example, the training sample can be a piece of image query text, such as a cat, a beautiful landscape, a city night scene text, etc., and the sample label can be a cat related image, such as an image of a cat napping in the sun, an image of a black and white cat playing in the grass, etc., such as an image of magnificent mountains and blue sky, an image of beautiful sunset scenery, etc., or an image of flashing city lights and high-rise buildings, an image of a busy street at night, etc. The task loss can be determined based on the difference between the output result of the student network according to the training sample in the sample set (i.e. the output cat, beautiful landscape or city night scene image) and the sample label. According to the task loss and the multiple distillation loss, the model parameters of the student network can be adjusted until the training samples are polled, and the target student network is obtained.

[0047] In some embodiments, the teacher network can include a first text encoding model and a first image encoding model, and the student network can include a second text encoding model and a second image encoding model obtained by knowledge distillation of the first text encoding model and the first image encoding model of the teacher network. The difference between the first text encoding model and the second text encoding model, and the first image encoding model and the second image encoding model can be that the number of heads of the multi-head attention mechanism of the first text encoding model is greater than that of the second text encoding model, and the number of heads of the multi-head attention mechanism of the first image encoding model is greater than that of the second image encoding model. The text encoding model can be used to process text data and convert the text data into a vector representation containing semantic information. The text encoding model can understand the meaning of words, phrases, etc. in the text. For example, given a text describing a cat, the text encoding model can identify keywords such as "cat" and "orange" in the text. The image encoding model can convert input image data into a feature vector with semantic meaning. For example, when identifying a picture of a landscape, the image encoding model can extract features such as mountains, rivers, and trees.

[0048] In some embodiments, an image retrieval model can be constructed according to the image encoding model and the text encoding model, which can implement the following functions: by using a piece of text as a query, one or more images consistent with the content description of the text can be retrieved from an image library. For example, when the input query is "a cat", the image retrieval model will return images related to "cat", such as an image of an orange cat napping in the sun, an image of a black and white cat playing on the grass, etc. In addition, when the input query is "beautiful scenery", the image retrieval model can return images of magnificent mountains and blue sky, images of beautiful sunset scenery, etc. Similarly, when the input query is "city night scene", the image retrieval model can return images of flashing city lights and high-rise buildings, images of night scenes on busy streets, etc. The image retrieval model obtained by the knowledge distillation method of the embodiments of the present application can help users quickly and accurately find images consistent with their description through simple text queries, providing a convenient image search and browsing experience.

[0049] According to an embodiment of the present application, by respectively determining a first attention map for a teacher network and a second attention map for a student network, and performing dimension normalization on a first splicing matrix based on splicing of a plurality of first attention maps and a second splicing matrix based on splicing of a plurality of second attention maps, a first attention matrix and a second attention matrix of the same dimension are obtained, a multiple distillation loss for knowledge distillation from the first attention map to the second attention map is determined according to the matrix features of the first attention matrix and the second attention matrix, and the model parameters of the student network are adjusted according to the multiple distillation loss and the task loss of the student network training, to obtain a target student network. Since the dimensions of the splicing matrices of the teacher network and the student network are unified in the process of knowledge distillation, attention matrices of the same dimension are obtained, and the multiple distillation loss is calculated according to the matrix features of the attention matrices after dimension unification; and the student network is trained by using the multiple distillation loss and the task loss together, effectively solving the knowledge loss problem caused by the inconsistency of the number of heads of the multi-head attention mechanism of the image retrieval model of the teacher network and the image retrieval model of the student network in related technologies, so that the student network can fully learn and absorb the knowledge of the teacher network, thereby improving the retrieval accuracy of the image retrieval model of the student network.

[0050] In some embodiments, in the process of training the student network, a training sample set matched with the image retrieval model can be called from the database according to the model identifier of the student network, such as the image retrieval model, the training sample set is synchronously input to the teacher network and the student network, the difference between the attention maps of the student network and the teacher network is measured by the distillation loss, the learning ability of the student network to the original task is ensured by the task loss of the student network, and the student network trained based on the distillation loss and the task loss and the model parameters of the student network are stored in the database or other storage medium, so that when processing the image retrieval task, the image retrieval model of the student network can be directly called from the database or other storage medium.

[0051] In some embodiments, when processing the image retrieval task, the user can send an image retrieval request through a terminal device, such as a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., for example, inputting a piece of text "a cat picture" to the terminal device. The server calls the trained image retrieval model from the database or other storage medium in response to the image retrieval request, inputs the image retrieval request into the image retrieval model, and sends the picture about the cat output by the image retrieval model to the terminal device, so that the user can browse the picture about the cat through the terminal device.

[0052] In some embodiments, the student network can be applied in the field of computer vision, for example, taking a visual model pre-trained based on an image dataset as a teacher network, and the student network obtained by knowledge distillation on the visual model can be applied to image classification, object detection, image retrieval, etc. For example, a vehicle in automatic driving determines whether it is a tunnel or a parking lot according to the current road image; for example, in the field of product quality inspection, whether a product has scratches and stains is detected according to the product image; for example, in the field of product recommendation, similar images are retrieved according to the input product image.

[0053] The student network can also be applied in the field of natural language processing, for example, taking a natural language model pre-trained based on a text dataset as a teacher network, and the student network obtained by knowledge distillation on the natural language model can be used for semantic understanding, text classification, sentiment analysis or question and answer system. For example, in the field of intelligent customer service, the student network can perform semantic understanding and automatic reply according to the user's question. For example, in the field of article recommendation, the student network can classify and recommend articles.

[0054] The student network can also be applied in the field of multi-modal interaction, for example, taking a multi-modal pre-training model pre-trained based on an image-text pair dataset as a teacher network, and the student network obtained by knowledge distillation on the multi-modal pre-training model can be used for multi-modal interaction tasks. For example, in the field of image retrieval, according to the user's input of "a cat", the image of "a cat napping in the sun" is returned.

[0055] By distilling the attention mechanism of the teacher network, the student network can inherit the ability of the teacher network while reducing the consumption of computing resources, which is suitable for running on edge devices.

[0056] Figure 3 An architecture diagram of knowledge distillation according to an embodiment of the present application is shown.

[0057] As shown in Figure 3 , taking a set of corresponding groups in the teacher network and the student network as an example, the sub-group of the teacher network in the group can include p first attention maps 101-1 (i.e. p first attention maps from the first attention layer of the teacher network are allocated to the first attention map 101-1), and the number of the first attention map 101-1 is p), and the sub-group of the student network can include q second attention maps 102-1 (i.e. q second attention maps from the second attention layer corresponding to the first attention layer of the student network are allocated to the second attention map 102-1), and the number of the second attention map 102-1 is q). Figure 3 ​​​​​​The middle can include a first splicing matrix 101-2 based on the first attention map 101-1, a second splicing matrix 102-2 based on the second attention map 102-1, a first weight matrix 301, a second weight matrix 302, a first attention matrix 101-3, a second attention matrix 102-3, a first re-distillation loss 303, a second re-distillation loss 304, a probability normalization layer 305, a normalized first attention matrix 101-4, a normalized second attention matrix 102-4, and a third distillation loss 306.

[0058] As shown in Figure 3 , the first weight matrix 301 can be a matrix for dimension transformation of the first splicing matrix 101-2, and the second weight matrix 302 can be a matrix for dimension transformation of the second splicing matrix 102-2.

[0059] Continuing to refer to Figure 3 , the process of dimension normalization of the first splicing matrix based on the plurality of first attention maps and the second splicing matrix based on the plurality of second attention maps in the operation S220 can include the following operations: dimension reduction of the first splicing matrix by using the first weight matrix, dimension reduction of the second splicing matrix by using the second weight matrix, so that the first attention matrix obtained after dimension reduction of the first splicing matrix and the second attention matrix obtained after dimension reduction of the second splicing matrix are square matrices of a predetermined dimension, and the first weight matrix and the second weight matrix are learnable matrices updated by a back propagation algorithm with the training of the student network. The predetermined dimension can be n rows and n columns, and n represents the number of tokens of the training sample input to the current layer.

[0060] Embodiments of the present application introduce learnable first weight matrices and second weight matrices , the first weight matrix , the dimension of , the first splicing matrix is transformed into the first attention matrix of n rows and n columns by matrix multiplication operation . Similarly, the second weight matrix , the dimension of . The second splicing matrix is transformed into the second attention matrix of n rows and n columns by matrix multiplication operation . The first weight matrix and the second weight matrix The parameters of the first weight matrix and the second weight matrix are dynamically optimized in the student network training process through a back propagation algorithm, so as to realize adaptive feature mapping from the spliced matrix to the target dimension feature matrix.

[0061] According to the embodiment of the present application, by reducing and unifying the dimensions of the first spliced matrix and the second spliced matrix, the dimension alignment of the attention matrix can be realized, the calculation complexity and memory occupation are reduced, and the training efficiency and accuracy of the model are improved.

[0062] In some embodiments, according to the dimension-normalized first attention matrix and the second attention matrix, a multiple knowledge distillation loss of the first attention graph to the second attention graph in a knowledge distillation process can be determined. For example, the principal component analysis method is used to determine the eigenvectors of the first attention matrix and the second attention matrix respectively, and according to the eigenvectors, a first multiple distillation loss of the first attention graph and the second attention graph is determined, the first multiple distillation loss is used to describe the attention mode difference between the first attention graph and the second attention graph; according to the graph distribution characteristics of the first attention matrix and the second attention matrix, a second multiple distillation loss of the first attention graph and the second attention graph is determined, the second multiple distillation loss is used to describe the statistical difference between the first attention graph and the second attention graph; according to the attention distribution characteristics of the first attention matrix and the second attention matrix, a third multiple distillation loss of the first attention graph and the second attention graph is determined, the third multiple distillation loss is used to describe the attention distribution difference between the first attention graph and the second attention graph.

[0063] In some embodiments, although the first attention matrix and the second attention matrix may both be square matrices , directly comparing the similarity of the two square matrices still has problems such as high calculation complexity and noise sensitivity. Therefore, the embodiment of the present application introduces the principal component analysis (PCA) algorithm to extract the core features of the first attention matrix and the second attention matrix respectively, and determines the first multiple distillation loss according to the eigenvectors of the core features.

[0064] In some embodiments, the feature vectors include a first feature vector of a first attention matrix and a second feature vector of a second attention matrix, wherein the first feature vector is related to the maximum eigenvalue of the first attention map matrix and the second feature vector is related to the maximum eigenvalue of the second attention matrix; determining the first redistillation loss of the first attention map and the second attention map based on the feature vectors includes: processing the first feature vector and the second feature vector based on cosine similarity to obtain the first redistillation loss.

[0065] For example, the first attention matrix Perform PCA analysis to solve for the first attention matrix. The eigenvalues ​​are determined, and the first eigenvector corresponding to the largest eigenvalue is selected. (n-dimensional column vector); for the second attention matrix Perform the same operation to obtain the second feature vector. (n-dimensional column vector). The core function of PCA is to map high-dimensional data to a low-dimensional space through linear transformation, preserving the most discriminative information in the data (i.e., the direction of the largest eigenvalue). The first eigenvector extracted here. Second eigenvector It can be regarded as a condensed representation of the core patterns of the attention maps of the teacher group and the student group, avoiding the interference of redundant information when directly comparing matrices.

[0066] After the above steps, the first feature vector Second eigenvector The first feature vector, which has become a feature vector with the same dimension (n-dimensional), can be quantified using a cosine similarity measure. Second eigenvector The difference between the two, as a group The first redistillation loss, i.e. Figure 3 The first redistillation loss shown can be represented by formula (1).

[0067] (1)

[0068] in, It is the first eigenvector The transpose of , Represents the first eigenvector Length of the mold Represents the second eigenvector The length of the module.

[0069] By maximizing and The cosine similarity is used to minimize the first redistillation loss. , to align the attention patterns of the student network to the teacher network, and realize knowledge transfer at the attention level. By using the cosine similarity to quantify the difference between the first feature vector and the second feature vector , the smaller the complement, the higher the alignment of the core features of the teacher network and the student network.

[0070] According to the embodiments of the present application, by determining the first attention map to the second attention map for knowledge distillation according to the first feature vector and the second feature vector, the difference between the core feature vectors extracted by PCA can be focused on, which can reflect the macro feature direction consistency of the attention pattern, so as to determine the knowledge difference between the teacher and the student network from different aspects, ensure the comprehensiveness of knowledge transfer, and improve the knowledge distillation effect.

[0071] In some embodiments, Figure 3 The second re-distillation loss shown can be obtained according to the graph distribution features of the first attention matrix and the second attention matrix. For example, the graph distribution features can include skewness and kurtosis, the skewness is used to represent the degree of asymmetry of the attention graph in the attention matrix, and the kurtosis is used to represent the steepness of the peak value of the attention graph distribution in the attention matrix; according to the graph distribution features of the first attention matrix and the second attention matrix, the second re-distillation loss of the first attention graph and the second attention graph is determined, including: based on the difference between the first skewness of the first attention matrix and the second skewness of the second attention matrix, and the difference between the first kurtosis of the first attention matrix and the second kurtosis of the second attention matrix, the second re-distillation loss is determined.

[0072] The second re-distillation loss can be a loss based on high-order statistics. The embodiments of the present application further introduce high-order statistics such as skewness (Skewness) and kurtosis (Kurtosis). By quantifying the difference between the teacher network and the student network in these high-order statistics, the second re-distillation loss term can be constructed.

[0073] For the first attention matrix and the second attention matrix , the first skewness and the first kurtosis of the first attention matrix are calculated, and the second skewness and the second kurtosis of the second attention matrix are calculated. The skewness and kurtosis can be calculated using the corresponding statistical formula, for example, the skewness formula is , and the kurtosis formula is (wherein is the first attention matrix or the second attention matrix the elements in the matrix, is the first attention matrix or the second attention matrix mean, is the first attention matrix or the second attention matrix standard deviation, is the first attention matrix or the second attention matrix expectation). The first skewness and the first kurtosis are calculated, and the input matrix is the first attention matrix The second skewness and the second kurtosis are calculated, and the input matrix is the second attention matrix .

[0074] According to the first skewness and the second skewness , the first kurtosis and the second kurtosis , the second re-distillation loss determined can be shown in formula (2).

[0075] (2)

[0076] wherein, and are hyperparameters, and the default values are , , a weight used to balance the skewness and kurtosis difference. The second re-distillation loss can mine deeper distribution feature differences of the attention map. Compared with using first-order (such as mean) and second-order statistics (such as variance and covariance), high-order statistics (i.e., skewness and kurtosis) can capture the non-Gaussian characteristics of the distribution, provide more abundant learning information for the student network, and help it more comprehensively imitate the attention mode of the teacher network.

[0077] According to the embodiments of the present application, by determining the second re-distillation loss of knowledge distillation from the first attention map to the second attention map according to the skewness and the kurtosis, the differences between the teacher network and the student network in these high-order statistics can be quantified to determine the knowledge differences between the teacher and the student network from different aspects, ensure the comprehensiveness of knowledge transfer, and improve the knowledge distillation effect.

[0078] In some embodiments, Figure 3The third re-distillation loss shown can be obtained according to the attention distribution characteristics of the first attention matrix and the second attention matrix. For example, the first attention matrix and the second attention matrix can be normalized respectively to obtain a normalized first attention matrix and a normalized second attention matrix; and the third re-distillation loss is determined according to the relative entropy between the attention distribution of the normalized first attention matrix and the attention distribution of the normalized second attention matrix.

[0079] In some embodiments, on the basis of the first re-distillation loss and the second re-distillation loss, the embodiment of the application further introduces a third re-distillation loss to capture the knowledge difference between the teacher network and the student network from different angles. The first step of determining the third re-distillation loss is to perform a probability normalization (Softmax) operation on the dimension-unified matrix first attention matrix and the second attention matrix to obtain a normalized first attention matrix and a normalized second attention matrix .

[0080] The core function of the Softmax function is to convert the element values of each row in the first attention matrix or the second attention matrix into a probability distribution (i.e., the sum of each row element is 1, and each element value is in the range of [0, 1]). The significance of this conversion is that the essence of the attention map is the association strength between tokens (the greater the value, the closer the association), and by using Softmax, this strength can be normalized into an "attention allocation probability", so that the matrix element more intuitively represents the "model attention weight distribution for different tokens", laying a foundation for subsequent loss calculation based on distribution difference.

[0081] The third re-distillation loss aims to further constrain the student network to learn the attention pattern of the teacher network by quantifying the difference between the normalized first attention matrix and the normalized second attention matrix The third re-distillation loss function can be as shown in formula (3).

[0082] (3)

[0083] Wherein, the function is a loss function for measuring the difference between two feature matrices (or distributions), and the core function is to quantify the deviation of the teacher network and the student network in the attention distribution, and guide the student network to adjust the parameters through back propagation to reduce the deviation. Specifically, the following several ways can be used: the function is the relative entropy (Kullback-Leibler, KL divergence), which is used to quantify the difference between the feature maps of the teacher network and the student network to guide the learning process of the student network; the function The mean square error loss, cosine similarity, and the like can also be used. The KL divergence can be as shown in equation (4).

[0084] (4)

[0085] wherein z is an index, and there are Z (Z is a positive integer) source models and Z target models respectively, P(z) can be the attention distribution of the source model, such as the attention distribution of the teacher network, and Q(z) can be the attention distribution of the target model, such as the attention distribution of the student network. When the function is the Kullback-Leibler divergence, the third-order distillation loss can be as shown in equation (5).

[0086] (5)

[0087] According to the embodiments of the present application, the third-order distillation loss for determining the knowledge distillation from the first attention map to the second attention map by normalizing the first attention matrix and the second attention matrix can focus on the difference in the attention distribution of the first attention map and the second attention map after the Softmax conversion, and reflect the microscopic details of the attention allocation, so as to determine the knowledge difference between the teacher and student networks from different aspects, ensure the comprehensiveness of the knowledge transfer, and improve the knowledge distillation effect.

[0088] In some embodiments, the teacher network includes the first attention layer, and the student network includes the second attention layer, the first attention map is derived from the first attention layer, and the second attention map is derived from the second attention layer, which have been described above. The first-order distillation loss, the second-order distillation loss, and the third-order distillation loss according to the above operation can be combined with the task loss of the student network to adjust the model parameters of the student network. For example, the layer distillation loss between the first attention layer and the second attention layer can be determined according to the multi-order distillation loss; the target distillation loss from the teacher network to the student network can be determined according to the layer distillation loss; the target loss of the student network training process can be determined according to the target distillation loss and the task loss; and the model parameters of the student network can be adjusted according to the target loss.

[0089] In an embodiment, the complementary information of the three losses can be fused according to the first-order distillation loss, the second-order distillation loss, and the third-order distillation loss to determine the total loss in the group .

[0090] Since the single loss function has limitations in capturing the knowledge difference between the teacher network and the student network, it is difficult to comprehensively cover the knowledge characteristics in different dimensions, therefore, the first-order distillation loss , the second-order distillation loss and the third-order distillation loss Fusion is performed to achieve more comprehensive knowledge alignment. The group total loss The calculation process can be shown as formula (6).

[0091] (6)

[0092] wherein, is a coefficient, by default, , , .

[0093] The fusion method has the advantages that, focus on the difference of the core feature vectors extracted by PCA, which reflects the macroscopic feature direction consistency of the attention mode; quantify the difference between the teacher network and the student network in these high-order statistics; and focus on the difference of the attention distribution after Softmax conversion, which reflects the microscopic details of attention allocation. The three complement each other, can capture the knowledge difference between the teacher and the student network from different levels, and ensure the comprehensiveness of knowledge transfer.

[0094] In some embodiments, the in formula (6) can be dynamically adjusted. For example can be dynamically adjusted according to the polling progress of the training sample. For example, when the ratio of the number of times of polling of the training sample to the total number of times of polling of the predetermined training sample is less than a first predetermined ratio, it is determined that the current training stage of the student network is the initial training stage, the student network parameter in the initial training stage is random, and the gap with the teacher network is huge, and the first and third redistillation losses can be paid more attention to in the initial training stage, and can be greater than , for example are 0.4, 0.2, and 0.4, respectively. When the ratio of the number of times of polling of the training sample to the total number of times of polling of the predetermined training sample is greater than or equal to the first predetermined ratio and less than a second predetermined ratio, it is determined that the current training stage of the student network is the middle training stage, the student network can have preliminarily mastered the attention mode of the teacher in the middle training stage, and more refined adjustment is needed, so The three can be a balanced distribution, for example 0.3, 0.3, 0.4 respectively. In the case that the ratio of the number of times of being polled of the training sample to the predetermined total number of times of polling of the training sample is greater than or equal to the second predetermined ratio, the current training stage of the student network is determined to be the late training stage, in which the student network and the teacher network can have been highly aligned, bottleneck breakthrough is needed, and deeper differences can be captured, and more attention can be paid to the differences in high-order statistics (skewness, kurtosis), and may each be less than , for example 0.2, 0.5, 0.3 respectively.

[0095] According to an embodiment of the present application, by performing dynamic adjustment in stages on , the training accuracy and efficiency of the student network can be improved, the robustness of the student network can be improved, the workload of manual parameter adjustment can be reduced, different application scenarios can be adapted to, and the knowledge distillation effect can be improved.

[0096] In the case of determining the total loss in the group , the layer distillation loss of the "one-to-one" layer selected in the foregoing can be determined. .

[0097] For any "one-to-one" corresponding layer selected in the foregoing (the layer contains corresponding groups), the layer distillation loss of the corresponding layer is the average loss of all groups of the corresponding layer, which can be as shown in formula (7).

[0098] (7)

[0099] The groups in the same layer correspond to different subsets of the attention map respectively, and the knowledge information carried by each group is different. By taking the average, the extreme loss value of a certain group can be avoided from excessively affecting the overall layer loss, and the knowledge contribution of each group can be effectively balanced. After such processing, the single-layer loss can more stably and objectively reflect the knowledge alignment effect of the layer, and provide a reliable basis for the calculation of the subsequent overall distillation loss.

[0100] The target distillation loss of the entire knowledge distillation process is the average loss of all selected "one-to-one" corresponding layers. By averaging the layer distillation loss of all participating attention layers, the knowledge of different attention layers (such as low-level basic features and high-level semantic features) is integrated, so that the student network can learn from the multi-level knowledge of the teacher network, and finally the performance is improved. The target distillation loss of the entire knowledge distillation process can be as shown in formula (8).

[0101] ​ (8)

[0102] Note that constant factors can also be added in the formula of calculating and the formula of calculating .

[0103] In some embodiments, the following describes the determination process of the task loss function. In the embodiments of the present application, knowledge distillation can be performed on a large language model, and the task is to give a sequence of tokens and then predict the next token. Then, the task loss function adopts a cross-entropy loss function.

[0104] If there is a sequence containing tokens, for each position e, the model needs to predict the probability distribution of the next token belonging to the V words in the vocabulary. Let be a V-dimensional one-hot vector, representing the category corresponding to the true token at position e, be the V-dimensional probability distribution vector predicted by the model, then the cross-entropy loss function can be as shown in formula (9).

[0105] (9)

[0106] wherein is the jth element of the vector, is the jth element of the vector. Formula (9) calculates the average difference between the probability distribution predicted by the student network model and the true label, and adjusts the parameters of the model by minimizing this loss function, so that the model can better predict the next token, thereby improving the performance of the language model.

[0107] According to the embodiments of the present application, by adjusting the model parameters of the student network according to multiple knowledge distillation, layer-by-layer and refined knowledge transfer is realized, which is beneficial to the student network to better imitate the reasoning path and feature evolution process of the teacher network, and improves the knowledge distillation effect.

[0108] In some embodiments, the above process of determining the target loss of the student network training process according to the target distillation loss and the task loss can include the following operations: adjusting the target distillation loss by using a predetermined weight to obtain an updated distillation loss, the predetermined weight being determined according to the accuracy of the output result of the student network, the predetermined weight being used to describe the importance of the target distillation loss to the target loss; determining the target loss according to the updated distillation loss and the task loss.

[0109] In some embodiments, based on the target distillation loss and the task loss described above, the embodiments of the present application define a target loss , as a function of the target distillation loss and the task loss In the embodiments of the present application, a linear function can be used, and the target loss can be shown as formula (10).

[0110] (10)

[0111] wherein, is a predetermined weight, and is a coefficient for balancing and By dynamically adjusting during the training process, better accuracy can be achieved.

[0112] In some embodiments, the predetermined weight described above is determined by: obtaining a first characteristic value according to the difference between the accuracy of the output result of the student network and a predetermined accuracy threshold; obtaining a second characteristic value according to the difference between the first training constant and the predetermined accuracy threshold; determining a first candidate weight according to the ratio of the first characteristic value and the second characteristic value; and determining the predetermined weight according to the maximum value between the first candidate weight and a second training constant, the second training constant being less than the first training constant.

[0113] In some embodiments, a dynamic adjustment function can be defined, wherein may be the accuracy of the student network in the current training batch. When the accuracy of the output result of the student network in the current training batch is used as the input, (wherein is a preset accuracy threshold, the default value being , is a first characteristic value, 1 is a first training constant, is a second characteristic value, is a first candidate weight, and 0.5 is a second training constant), so that when the accuracy of the student network is low, for example, lower than a predetermined value, the weight of can be appropriately reduced, and when the accuracy of the student network is high, for example, higher than a predetermined value, the weight of can be appropriately increased, and the predetermined value can be the preset accuracy threshold.

[0114] According to the embodiment of the present application, by providing dynamic predetermined weights and dynamically adjusting the target distillation loss by using the dynamic predetermined weights, intelligent dynamic switching of the training center of gravity can be realized, automatic optimization of the model parameters can be realized, and the stability and efficiency of model training can be improved.

[0115] The embodiment of the present application first extracts core feature vectors from the attention feature matrix of the Transformer architecture teacher network and student network after dimension unification through principal component analysis (PCA), calculates the first loss based on the cosine similarity complement to align the macro feature direction, then calculates the second re-distillation loss of the teacher network and student network through skewness and kurtosis, and then performs Softmax probability normalization on the matrix after dimension unification, calculates the second loss based on the attention distribution difference by using KL divergence to match the microscopic attention mode; by fusing multiple losses, the comprehensive transfer of teacher network attention knowledge is realized, and the limitations of single loss on knowledge capture are avoided.

[0116] The embodiment of the present application focuses on the precise correspondence of the group-established attention graph in the Transformer architecture, realizes adaptive dimension conversion and unification by using a learnable matrix, and migrates attention knowledge by combining multiple losses. First, the attention graphs of the teacher and student are divided into the same number of groups by average allocation according to the sequence number to establish a corresponding relationship, then the different dimension matrices spliced in the group are unified into a square matrix of the same dimension by using a learnable matrix, and finally the knowledge distillation across the number of attention heads is realized by combining multiple losses, effectively solving the problems of number mismatch and dimension inconsistency caused by the difference in the number of attention heads between the teacher network and the student network. At the same time, this scheme can accurately capture the core knowledge of the attention layer under the Transformer architecture, significantly improve the efficiency and effectiveness of knowledge distillation, help the student network fully absorb the attention pattern of the teacher network, enhance the performance of the student network, further expand the adaptability and application value of the knowledge distillation method in the Transformer-based neural network model, and improve the performance of the student network. The above method is not only suitable for knowledge distillation in the scenario where the number of attention graphs of the Transformer architecture teacher network and student network is inconsistent, but also can be applied to natural language processing, computer vision or multi-modal understanding based on the Transformer architecture deep learning task to improve the reasoning performance and efficiency of the student network in the task.

[0117] Based on the above-described training process of the student network combined with multiple distillation losses and task losses, a specific embodiment is provided below to describe the training process of the student network.

[0118] For the training data set D, the source model (teacher network) and the target model (student model) in this embodiment are large language models. The data in the training data set is text data. The text data is input into the large language model, thereby generating intermediate state values in the model running (for example: query embedding, key embedding, value embedding, etc., and the final output text. In a large language model for question and answer, the input is a question, and the output is the answer to the question. The present application collects a large-scale text data set containing hundreds of millions of sentences, in which: the proportion of Chinese sentences is 30%, the proportion of English sentences is 60%, and the proportion of other languages is 10%. In the process of model training, the text data is sampled from the large-scale text data set in a random sampling manner. The above language proportions are exemplary values, and other numerical proportions do not affect the effect of the embodiments of the present application.

[0119] For the setting method of the input sample, in order to avoid the order of the sample in the training data set D affecting the performance of the model, at the beginning of the model training, all the samples in the training data set D are randomly arranged. In each round of training process, the model training algorithm reads a batch of samples, for example: 1024 samples.

[0120] For the model training algorithm, it can include training input, training output, loss function setting and model training steps.

[0121] The training input can be an input source model (teacher network) , for example: a large language model based on Transformer. The data is read from the training data set D, assuming that the data has been preprocessed as required and the format meets the model input requirements; the initial training period , the total training period ; the stochastic gradient descent algorithm is the Adam optimizer algorithm, and the related parameters are: the learning rate is , the momentum coefficient , , , the learning rate strategy is the cosine annealing strategy (cosine strategy), and the specific adjustment method is to dynamically adjust the learning rate according to the cosine function within the training period. The number of samples in a small batch is B = 1024.

[0122] The training output is a target model (student network) , wherein the target model is basically the same as the structure of the source model, and only the number of heads of the multi-head attention mechanism is different. In each layer of the target model The number of heads of the multi-head attention of the target model is not greater than the number of heads of the multi-head attention of the source model.

[0123] The loss function is set, and the training loss function is shown in formula (10) .

[0124] The model training steps are as follows:

[0125] Step 1: fix all parameter values of the source model ; construct the network structure of the target model according to the network structure of the source model ; randomly initialize the element values of the weight matrix in the target model and other parameters.

[0126] Step 2: when , the following operations are performed:

[0127] Step 2.1: .

[0128] Step 2.2: randomly shuffle the samples in the data set D.

[0129] Step 2.3: select a batch B of training samples from the data set D.

[0130] Step 2.4: according to the above training samples and parameter settings, calculate the loss value, and then update the parameters of the target model using the stochastic gradient descent algorithm (Adam algorithm). Before updating the parameters, the gradient is clipped, and the gradient threshold is set to 5 (which can be adjusted according to actual conditions) to prevent gradient explosion.

[0131] Step 2.5: repeat steps 2.3 and 2.4 until all samples in the data set D are used.

[0132] Step 2.6: end the current training period.

[0133] Step 3: end the training, and output the target model . Save the target model to a specified file path for subsequent use.

[0134] Based on the above model training method, the embodiment of the application also provides a method for applying the target model trained above, specifically, the embodiment of the application provides an image retrieval model.

[0135] Figure 4 An architecture diagram of an image retrieval model according to an embodiment of the application is shown.

[0136] As Figure 4As shown, the architecture of the image retrieval model can include a first image encoding model 401 and a first text encoding model 402 as teacher networks, and a second image encoder 403 and a second text encoding model 404 obtained by distilling the first image encoding model 401 and the first text encoding model 402 respectively using the knowledge distillation method described above as student networks. The second image encoding model 403 and the second text encoding model 404 can both be Transformer-based neural networks.

[0137] The image encoding model and the text encoding model can be a Contrastive Language-Image Pre-training (CLIP) model trained on a data set using images and texts.

[0138] An image database Db to be queried can be created in advance before applying the image retrieval model. The image database Db is an image database to be queried, which includes a large number of unlabeled images. The content of these images is unknown, i.e., the object classes in the images are unknown. The image database Db is large in size and covers images of multiple fields and scenes. Each image in the image database is an independent instance and has not been classified or labeled in advance. For the convenience of subsequent description, the total number of images of the image database Db is defined as .

[0139] According to the image database to be queried, an image feature database is created, each image in the image database to be queried Db is input into the second image encoding model to obtain an output feature embedding vector of each image, . In , the is taken as the feature of the image . Then, the image feature database includes image features, and the set composed of , the image features correspond one-to-one to the images.

[0140] In the case where the image database Db is created, the application process of the image retrieval model can be as follows.

[0141] The input query text T is a natural language text, and the language type should be consistent with the language type of the text in the training data set of the jointly trained image and text model. In the training data set of the jointly trained image and text model, the image and the corresponding text form an image-text pair, and the language type of the text determines the input text language type that the second text encoding model can process. In other words, if the text in the training data set of the jointly trained image and text model is Chinese, the input query text should also be Chinese; if the text in the training data set is English, the input query text should also be English; if the text in the training data set covers multiple languages, the input query text can contain multiple languages. Ensuring that the input query text is consistent with the language type of the text in the training data set ensures the accuracy and reliability of image retrieval.

[0142] The features of the input query text T are obtained, and the feature embedding vector obtained by inputting the query text T into the second text encoding model is , In the , take as the features of the input query text.

[0143] The similarity between the features of the input query text T and the image features in the image feature database is calculated, and the image feature database includes image features, and the set composed of the image features is . The features of the input query text T are . Then, the similarity between the features of the input query text T and each feature in the image feature data is as shown in formula (11).

[0144] (11)

[0145] Wherein, the operation process represented by “ ” takes as an example, wherein, and represent two column vectors, for example , represents the transpose vector of , and represents the vector length.

[0146] The image index is obtained, and the task of the image retrieval model of the present embodiment is to find one or more images corresponding to the input text content from the image database according to the input text. Therefore, a similarity threshold is set in the present embodiment.

[0147] ​The similarity between the features of the input query text T and the image features is as follows:

[0148]

[0149] So, based on the similarity threshold The resulting image index set is as follows:

[0150] , represents returning all A set of indices.

[0151] Returns a set of retrieved images, based on the image index set. It retrieves the corresponding image from the image database Db and returns it.

[0152] Based on the above-described knowledge distillation method, this invention also provides a knowledge distillation apparatus. The following will be combined with... Figure 5 The device is described in detail.

[0153] Figure 5 A structural block diagram of a knowledge distillation apparatus according to an embodiment of the present invention is shown.

[0154] like Figure 5 As shown, the knowledge distillation apparatus 500 of this embodiment includes a first determining module 510, a normalization module 520, a second determining module 530, and an adjustment module 540.

[0155] The first determining module 510 is used to determine multiple first attention maps for the teacher network and multiple second attention maps for the student network, respectively.

[0156] The normalization module 520 is used to normalize the dimensions of the first splicing matrix obtained by splicing multiple first attention maps and the second splicing matrix obtained by splicing multiple second attention maps, so as to obtain the first attention matrix and the second attention matrix with the same dimensions.

[0157] The second determining module 530 is used to determine the multiple distillation loss for knowledge distillation from the first attention map to the second attention map based on the matrix characteristics of the first attention matrix and the matrix characteristics of the second attention matrix.

[0158] The adjustment module 540 is used to adjust the model parameters of the student network during the training process of the student network based on the multiple distillation loss and the task loss of the student network, until all samples in the sample set used to train the student network are polled to obtain the target student network. The sample set includes training samples and sample labels, and the task loss is used to describe the difference between the output results of the student network based on the training samples and the sample labels.

[0159] According to an embodiment of the present application, any of the first determining module 510, the normalizing module 520, the second determining module 530 and the adjusting module 540 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present application, at least one of the first determining module 510, the normalizing module 520, the second determining module 530 and the adjusting module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging a circuit, etc. in hardware or firmware, or implemented in any one of software, hardware and firmware or in a proper combination of any of them. Alternatively, at least one of the first determining module 510, the normalizing module 520, the second determining module 530 and the adjusting module 540 can be at least partially implemented as a computer program module which, when executed, can perform the corresponding function.

[0160] It should be noted that the knowledge distillation device part in the embodiments of the present application corresponds to the knowledge distillation method part in the embodiments of the present application, and the description of the knowledge distillation device part is specifically referred to the knowledge distillation method part, which will not be repeated here.

[0161] Figure 6 A block diagram of an electronic device suitable for implementing the knowledge distillation method according to an embodiment of the present application is schematically shown.

[0162] As shown in Figure 6 The electronic device 600 according to an embodiment of the present application includes a processor 601 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded from a storage part 608 to a random access memory (RAM) 603. The processor 601 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset, and / or a special-purpose microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 can also include an on-board memory for cache use. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method processes according to embodiments of the present application.

[0163] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via the bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.

[0164] According to the embodiments of the present application, the electronic device 600 can further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the input / output (I / O) interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 608 including a hard disk, etc.; and a communication part 609 including a network interface card such as a LAN card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as necessary. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 610 as necessary, so that a computer program read out therefrom is installed in the storage part 608 as necessary.

[0165] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present application is implemented.

[0166] According to an embodiment of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer readable storage medium can include one or more memories of the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603.

[0167] Embodiments of the present application also include a computer program product comprising a computer program containing program code for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program code is used to make the computer system implement the knowledge distillation method provided by the embodiments of the present application.

[0168] The above functions defined in the system / device of the embodiments of the present application are performed when the computer program is executed by the processor 601. According to an embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by computer program modules.

[0169] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or installed from the detachable medium 611. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the foregoing.

[0170] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the detachable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiments of the present application are performed. According to an embodiment of the present application, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.

[0171] According to embodiments of the present application, program code for implementing the computer programs provided by embodiments of the present application can be written in any combination of one or more programming languages, and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming language can include, but is not limited to, Java, C++, python, "C" language, or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the remote computing device, or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider.

[0172] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0173] Those skilled in the art will appreciate that the features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways, even if such combinations or integrations are not expressly noted in the present application. In particular, the features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways without departing from the spirit and scope of the present application. All such combinations and / or integrations are within the scope of the present application.

[0174] The embodiments of the present application have been described above. However, these embodiments are merely for the purpose of illustration and are not intended to limit the scope of the present application. Although the embodiments are described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present application, and these substitutions and modifications should fall within the scope of the present application.

Claims

1. A knowledge distillation method, characterized in that, The method includes: Multiple first attention maps are identified for the teacher network and multiple second attention maps are identified for the student network. The dimensions of the first concatenation matrix obtained by concatenating the multiple first attention maps and the second concatenation matrix obtained by concatenating the multiple second attention maps are normalized to obtain first attention matrices and second attention matrices with the same dimensions. Based on the matrix features of the first attention matrix and the matrix features of the second attention matrix, a multiple distillation loss is determined for knowledge distillation from the first attention graph to the second attention graph; During the training of the student network, the model parameters of the student network are adjusted according to the multiple distillation loss and the task loss of the student network until all samples in the sample set used to train the student network are polled to obtain the target student network. The sample set includes training samples and sample labels. The task loss is used to describe the difference between the output results of the student network based on the training samples and the sample labels. The step of determining the multiple distillation loss for knowledge distillation from the first attention map to the second attention map based on the matrix features of the first attention matrix and the matrix features of the second attention matrix includes: The eigenvectors of the first attention matrix and the second attention matrix are determined by principal component analysis, and the first redistillation loss of the first attention map and the second attention map is determined based on the eigenvectors. The first redistillation loss is used to describe the difference in attention patterns between the first attention map and the second attention map. Based on the graph distribution characteristics of the first attention matrix and the second attention matrix, a second redistillation loss is determined for the first attention graph and the second attention graph. The second redistillation loss is used to describe the statistical difference between the first attention graph and the second attention graph. Based on the attention distribution characteristics of the first attention matrix and the second attention matrix, a third redistillation loss is determined for the first attention map and the second attention map. The third redistillation loss is used to describe the difference in attention distribution between the first attention map and the second attention map.

2. The method according to claim 1, characterized in that, The feature vector includes a first feature vector of a first attention matrix and a second feature vector of a second attention matrix. The first feature vector is related to the maximum eigenvalue of the first attention matrix, and the second feature vector is related to the maximum eigenvalue of the second attention matrix. The step of determining the first redistillation loss of the first attention map and the second attention map based on the feature vector includes: The first redistillation loss is obtained by processing the first feature vector and the second feature vector based on cosine similarity.

3. The method according to claim 1, characterized in that, The graph distribution features include skewness and kurtosis. Skewness is used to represent the degree of asymmetry of the attention matrix, and kurtosis is used to represent the steepness of the peak value of the attention matrix. The step of determining the second redistillation loss of the first attention map and the second attention map based on the graph distribution characteristics of the first attention matrix and the second attention matrix includes: The second redistillation loss is determined based on the difference between the first skewness of the first attention matrix and the second skewness of the second attention matrix, and the difference between the first kurtosis of the first attention matrix and the second kurtosis of the second attention matrix.

4. The method according to claim 1, characterized in that, The step of determining the third redistillation loss of the first attention map and the second attention map based on the attention distribution characteristics of the first attention matrix and the second attention matrix includes: The first attention matrix and the second attention matrix are normalized respectively to obtain the normalized first attention matrix and the normalized second attention matrix; The third redistillation loss is determined based on the relative entropy between the attention distribution of the normalized first attention matrix and the attention distribution of the normalized second attention matrix.

5. The method according to claim 1, characterized in that, The teacher network includes a first attention layer, and the student network includes a second attention layer. The first attention graph originates from the first attention layer, and the second attention graph originates from the second attention layer. The step of adjusting the model parameters of the student network based on the multiple distillation loss and the task loss of the student network includes: Based on the multiple distillation loss, determine the layer distillation loss between the first attention layer and the second attention layer; Based on the layer distillation loss, determine the target distillation loss from the teacher network to the student network; The target loss of the student network training process is determined based on the target distillation loss and the task loss. The model parameters of the student network are adjusted based on the target loss.

6. The method according to claim 5, characterized in that, Determining the target loss of the student network training process based on the target distillation loss and the task loss includes: The target distillation loss is adjusted using predetermined weights to obtain an updated distillation loss. The predetermined weights are determined based on the accuracy of the output results of the student network and are used to describe the importance of the target distillation loss to the target loss. The target loss is determined based on the updated distillation loss and the task loss.

7. The method according to claim 6, characterized in that, The predetermined weights are determined in the following manner: The first feature value is obtained based on the difference between the accuracy of the output result of the student network and a predetermined accuracy threshold. The second feature value is obtained based on the difference between the first training constant and the predetermined accuracy threshold; The first candidate weight is determined based on the ratio of the first feature value to the second feature value; The predetermined weight is determined based on the maximum value between the first candidate weight and the second training constant, wherein the second training constant is less than the first training constant.

8. The method according to claim 1, characterized in that, The step of normalizing the dimensions of the first concatenation matrix obtained by concatenating the multiple first attention maps and the second concatenation matrix obtained by concatenating the multiple second attention maps to obtain first attention matrices and second attention matrices with the same dimensions includes: The first concatenated matrix is ​​reduced in dimension using a first weight matrix, and the second concatenated matrix is ​​reduced in dimension using a second weight matrix, so that the first attention matrix obtained after reducing the dimension of the first concatenated matrix and the second attention matrix obtained after reducing the dimension of the second concatenated matrix are square matrices of a predetermined dimension. The first weight matrix and the second weight matrix are learnable matrices that are updated by the backpropagation algorithm as the student network is trained.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Knowledge distillation method, electronic equipment and computer readable storage medium

    CN120611768A