A Feature Distillation Method, System, Device and Medium Based on Multi-Model Fusion

Through the feature distillation method of multi-model fusion, the attention module and normalization layer are used to assign weights and similarity calculations to the teacher model features. Combined with PCA dimensionality reduction, the problems of huge dark knowledge matrix and feature redundancy of the teacher model are solved, and efficient student model learning and recognition performance improvements are achieved.

CN114462546BActive Publication Date: 2025-07-11SHANGHAI YUNCHONG ENTERPRISE DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210142194.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-07-11
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

In the prior art, the dark knowledge matrix of the teacher model for face recognition tasks is too large, the hardware resource consumption is high, and there is redundancy between features, which affects the learning and computing efficiency of the student model.

Method used

The multi-model fusion feature distillation method is adopted to obtain the features of the target data through multiple pre-trained teacher models, and feature distillation is performed through the backbone network and distillation subnet of the student model. The attention module and normalization layer are used to assign weights and similarity calculations to the features, and combined with the PCA dimensionality reduction algorithm to realize local and global learning of the features.

Benefits of technology

It effectively reduces the consumption of hardware resources, improves the recognition performance of student models, realizes more accurate learning of distillation features, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462546B_ABST
    Figure CN114462546B_ABST
Patent Text Reader

Abstract

The present invention provides a feature distillation method, system, device and medium based on multi-model fusion, including: respectively obtaining features of target data as first features through a plurality of pre-trained teacher models; obtaining second features of the target data through the backbone network of a student model, inputting the second features into a plurality of first distillation sub-networks respectively, and respectively outputting second features whose similarity with the first features reaches a set threshold through each of the first distillation sub-networks; fusing all the first features to obtain a first fused feature, and fusing the second features output by each distillation sub-network to obtain a second fused feature, inputting the first fused feature and the second fused feature into a second distillation sub-network to obtain the distillation feature of the target data; The present invention makes full use of the advantages of different teacher models, conducts distillation learning from both local and global directions, and improves the recognition performance of the student model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a feature distillation method, system, device, and medium based on multi-model fusion. Background Art

[0002] Model compression and knowledge extraction are key steps in model deployment. Among them, the training method mainly based on model distillation is widely used. The mainstream model distillation method will pre-train a large model (teacher model), calculate the probabilities of each category in the classification layer, use this probability distribution as "dark knowledge", and use the distance metric of KL divergence to guide the small model (student model) to learn the knowledge of the large model.

[0003] In the face recognition task, this method faces the following problems: The number of categories in the face recognition task is huge, which will cause the distribution of the dark knowledge matrix in the teacher model to be too large, not conducive to learning, and even consume a lot of hardware resources such as video memory; The feature fusion of multiple teacher models will form a more powerful teacher model, but improper training methods cannot fully obtain the benefits brought by multiple teachers, but instead increase the length of the features, bringing the burden of calculation and storage. Summary of the Invention

[0004] In view of the above problems existing in the prior art, the present invention proposes a feature distillation method, system, device, and medium based on multi-model fusion, mainly solving the problems that the dark knowledge matrix of the existing teacher model is too large, has high hardware requirements, and there are redundancies between features, which is not conducive to the learning and calculation of the student model.

[0005] In order to achieve the above object and other objects, the technical solution adopted by the present invention is as follows.

[0006] A feature distillation method based on multi-model fusion includes:

[0007] Obtaining the features of the target data as the first features through a plurality of pre-trained teacher models respectively;

[0008] Obtaining the second features of the target data through the backbone network of the student model, inputting the second features into a plurality of first distillation sub-networks respectively, and outputting the second features whose similarity with the first features reaches a set threshold through each of the first distillation sub-networks;

[0009] Fusing all the first features to obtain a first fusion feature, fusing the second features output by each distillation sub-network to obtain a second fusion feature, and inputting the first fusion feature and the second fusion feature into a second distillation sub-network to obtain the distillation features of the target data.

[0010] Optionally, the first distillation sub-network includes: an attention module, a normalization layer, a similarity calculation layer, and at least one fully connected layer.

[0011] The attention module obtains the weights corresponding to the features according to the magnitudes of the feature values of the features output by the fully connected layer and outputs them to the normalization layer.

[0012] The normalization layer completes the normalization of the corresponding features according to the features output by the fully connected layer and the weights output by the attention module.

[0013] The similarity calculation layer obtains the similarity between the normalized features and the first features output by the corresponding teacher model through a preset loss function.

[0014] Optionally, the attention module maps the feature values to the range between -1 and 1 through a mapping function.

[0015] Optionally, the mapping function includes: softmax function, sigmoid function.

[0016] Optionally, the second distillation sub-network has the same network structure as the distillation sub-network.

[0017] Optionally, before inputting the first fusion feature and the second fusion feature into the second distillation sub-network, it further includes:

[0018] Performing dimensionality reduction processing on the first fusion feature by using a dimensionality reduction algorithm.

[0019] Optionally, the number of the first distillation sub-networks corresponds to the number of the teacher models, and each first distillation sub-network receives the first feature of one of the teacher models respectively.

[0020] A feature distillation system based on multi-model fusion includes:

[0021] A first feature acquisition module, configured to respectively obtain the features of the target data as the first features through a plurality of pre-trained teacher models;

[0022] A student feature acquisition module, configured to obtain the second features of the target data through the backbone network of the student model, input the second features into a plurality of first distillation sub-networks respectively, and output the second features whose similarity with the first features reaches a set threshold through each of the first distillation sub-networks;

[0023] A fusion distillation module, configured to fuse all the first features to obtain a first fusion feature, fuse the second features output by each distillation sub-network to obtain a second fusion feature, input the first fusion feature and the second fusion feature into the second distillation sub-network, and obtain the distilled features of the target data.

[0024] An apparatus, comprising:

[0025] one or more processors; and

[0026] one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the apparatus to perform the multi-model fusion-based feature distillation method described above.

[0027] A machine-readable medium having instructions stored thereon, which, when executed by one or more processors, cause an apparatus to perform the multi-model fusion-based feature distillation method described above.

[0028] As described above, a multi-model fusion-based feature distillation method, system, apparatus, and medium of the present invention have the following beneficial effects.

[0029] After the second feature of the backbone network is distilled by the distillation sub-network and compared with the first feature obtained by the corresponding teacher model for similarity, the student model can learn the main features in the teacher model. Then, through the fusion of features for distillation, the student model is globally guided by the fusion features of the teacher model, realizing a learning process from local to global, completing sufficient learning, and obtaining more accurate distilled features. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic flowchart of a multi-model fusion-based feature distillation method according to an embodiment of the present invention.

[0031] Figure 2 It is a module diagram of a multi-model fusion-based feature distillation system according to an embodiment of the present invention.

[0032] Figure 3 It is a schematic structural diagram of an apparatus according to an embodiment of the present invention.

[0033] Figure 4 It is a schematic structural diagram of an apparatus according to another embodiment of the present invention.

[0034] Figure 5 It is a schematic structural diagram of a distillation sub-network according to an embodiment of the present invention.

[0035] Figure 6 It is a schematic network structure diagram of a student model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The following describes the implementation manners of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0038] Please refer to Figure 1 , the present invention provides a feature distillation method based on multi-model fusion, including the following steps:

[0039] Step S01, respectively obtain the features of the target data as the first features through multiple pre-trained teacher models;

[0040] Step S02, obtain the second features of the target data through the backbone network of the student model, input the second features into multiple first distillation sub-networks respectively, and output the second features whose similarity to the first features reaches a set threshold through each of the first distillation sub-networks;

[0041] Step S03, fuse all the first features to obtain a first fused feature, fuse the second features output by each distillation sub-network to obtain a second fused feature, and input the first fused feature and the second fused feature into a second distillation sub-network to obtain the distilled features of the target data.

[0042] In one embodiment, models with multiple different network structures can be used as teacher models, and each teacher model is pre-trained with labeled sample data to determine the model parameters of the teacher models. In order to ensure the performance of the features obtained by each teacher model during feature fusion, teacher models with recognition accuracies differing within a certain range can be selected. The specific selection of teacher models and the pre-training process are prior arts and will not be elaborated here.

[0043] Respectively extract the features of the target data as the first features through each pre-trained teacher model. Among them, the target data may include face images, vehicle images, etc. of the object to be recognized. Store the first features extracted by each teacher model for use by the student model for distillation.

[0044] The student model may include a backbone network, a local distillation module, and a global distillation module. The local distillation module includes a plurality of first distillation sub-networks. The global distillation module splices and fuses the first features of the target data obtained by each teacher model to obtain a fused feature matrix. Specifically, the first features of each teacher model can be concatenated to obtain a high-dimensional feature matrix. Since the feature length linearly increases at this time, limited by the learning ability of the student model and the mutual redundancy between the first features obtained by different teacher models, the PCA dimensionality reduction algorithm can be used to perform dimensionality reduction processing on the fused feature matrix, extract the main features, and obtain the first fused feature. Further, the features output by each first distillation sub-network are fused in the same way to obtain a second fused feature, and the first fused feature and the second fused feature are input into a second distillation sub-network to obtain the distillation feature of the target data. The second distillation sub-network can adopt the same network structure as the aforementioned distillation sub-network.

[0045] In one embodiment, the student model can extract features from the target data through the backbone network to obtain second features. The backbone network can adopt a conventional feature extraction network structure, such as extracting features through one or more convolutional layers. Other network structures capable of feature extraction can also be adopted according to needs, which are not limited here. The backbone network can establish connections with each first distillation sub-network respectively, and output the extracted second features to each first distillation sub-network for feature distillation.

[0046] In one embodiment, the first distillation sub-network includes: an attention module, a normalization layer, a similarity calculation layer, and at least one fully connected layer; the attention module obtains the weight of the corresponding feature according to the eigenvalue size of the feature output by the fully connected layer and outputs it to the normalization layer; the normalization layer completes the normalization of the corresponding feature according to the feature output by the fully connected layer and the weight output by the attention module; the similarity calculation layer obtains the similarity between the normalized feature and the first feature output by the corresponding teacher model through a preset loss function. For the specific network structure of the first distillation sub-network, reference can be made to Figure 5 , Figure 5The feature head represents a first distillation sub-network. The feature head includes two fully connected layers. The ReLU function is used as the activation function between the two fully connected layers. The node outputs from the previous fully connected layer to the next fully connected layer are selectively opened and closed through the activation function to prevent overfitting. The number of fully connected layers can be adjusted according to actual application requirements and is not limited here. Taking the feature head with two fully connected layers as an example, after the two fully connected layers, an attention module is connected. Through the attention module, the feature values of the features output by the fully connected layers are mapped to between -1 and 1, and the weights corresponding to each feature are calculated according to the magnitudes of the feature values. Since feature learning is not equally important for each dimension, the larger the absolute value of the feature dimension, the greater the impact of the weight on the final recognition score, and it should be focused on learning. Therefore, an attention module is added. After the fully connected layers are stacked through the activation function, features of a specific dimension are output. After normalizing the features of the specific dimension, the weight magnitude is restricted to -1 to 1. Then, the absolute value of the features of the specific dimension output by the fully connected layers is processed through a mapping function to obtain the weight magnitudes of the features of each dimension. Among them, the mapping function can adopt mapping functions that do not change monotonicity such as the softmax function, sigmoid function, etc., as long as it is ensured that larger feature values are assigned larger weights. The specific selection of the mapping function can be adjusted according to application requirements and is not limited here. The normalized features with weights are used to calculate the similarity with the first feature output by the teacher module.

[0047] In one embodiment, each distillation sub-network can be connected to the output of a teacher model. The teacher model and the distillation sub-network are combined in pairs, and the features output by the corresponding teacher model are distilled through the distillation sub-network. Please refer to Figure 6 , in the figure, the feature heads 1-3 respectively correspond to three distillation sub-networks. Each feature head calculates the similarity between the normalized weighted features in the feature head and the first feature (i.e., the teacher feature in the figure) output by the teacher model through a loss function. Among them, the loss function can adopt a cosine loss function, Euclidean distance loss function, KL divergence loss function, etc. The specific loss function can be selected according to actual application requirements and is not limited here. As Figure 6 shown, first, local distillation training is performed on a single feature head. Each single feature head has an attention mechanism, and the cosine distance is used to calculate the feature loss. The features output by each feature head are fused to obtain the student fusion feature. Then, the teacher features of each teacher model are concatenated to form a feature matrix, and the PCA algorithm is used to perform dimensionality reduction processing on the feature matrix to select the main features to obtain the teacher fusion feature. The similarity between the teacher fusion feature and the student fusion feature is calculated through the second distillation sub-network (i.e., the feature head 4 in the figure) to obtain the global distillation feature. The network structure of the second distillation sub-network can adopt the same network structure as the feature heads 1-3.

[0048] The student model aims at the features of the teacher model. The features obtained by the backbone network of the student model are used to calculate the cosine distance between the calculated features and the teacher features through the first distillation sub-network, and a cosine loss function is constructed for training. The lower the loss, the closer the feature distance between the student and the teacher, and the better the learned features. When training with the cosine loss function, feature learning is not equally important for each dimension. The larger the absolute value of the feature dimension weight, the greater the impact on the final recognition score, and such dimensions should be focused on for learning. Therefore, an attention module is proposed. After stacking activation functions of the fully connected layer, features of specific dimensions are output. After normalization, the weight size is restricted to -1 to 1, and then the weight size of each dimension feature is obtained through the absolute value and the softmax function. Combined with the cosine loss function, the student network can be tilted towards learning more important feature dimensions.

[0049] For multi-teacher features, multiple feature heads are designed to "tutor" the student model one by one, enabling the student model to more attentively master the feature knowledge of the teachers. From a global perspective, the features of multiple teachers will ultimately be concatenated into a fused feature, which has the characteristics of strong expressive ability but a long feature dimension. Considering the limited learning ability of the student model and the redundancy of the multi-teacher fused features, the PCA dimensionality reduction technique is used to extract the main global features of the specified dimension for the student model to learn.

[0050] For the face recognition task, distilling the recognition features can significantly save hardware resources. Adding a feature head with an attention mechanism makes the distillation learning of the student model more efficient and converges faster. The multi-model feature learning makes full use of the advantages of different teacher models and conducts distillation learning from both local and global directions, further improving the recognition performance of the student model.

[0051] Please refer to Figure 2 , in this embodiment, a feature distillation system based on multi-model fusion is provided for performing the feature distillation method based on multi-model fusion described in the foregoing method embodiment. Since the technical principle of the system embodiment is similar to that of the foregoing method embodiment, the same technical details will not be repeated.

[0052] In one embodiment, the feature distillation system based on multi-model fusion includes:

[0053] The first feature acquisition module 10 is configured to respectively acquire features of target data through multiple pre-trained teacher models as first features; the student feature acquisition module 11 is configured to acquire second features of the target data through the backbone network of the student model, input the second features into multiple first distillation sub-networks respectively, and output second features whose similarity to the first features reaches a set threshold through each of the first distillation sub-networks; the fusion distillation module 12 is configured to fuse all the first features to obtain a first fused feature, fuse the second features output by each distillation sub-network to obtain a second fused feature, and input the first fused feature and the second fused feature into a second distillation sub-network to acquire the distillation features of the target data.

[0054] Embodiments of the present application further provide a device, which may include: one or more processors; and one or more machine-readable media storing instructions thereon, which when executed by the one or more processors, cause the device to execute Figure 1 the method described above. In practical applications, the device may be used as a terminal device or a server. Examples of terminal devices may include: smart phones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, in-vehicle computers, desktop computers, set-top boxes, smart televisions, wearable devices, etc. Embodiments of the present application do not impose any restrictions on the specific device.

[0055] Embodiments of the present application further provide a non-volatile readable storage medium, which stores one or more modules (programs), and when the one or more modules are applied to a device, they can cause the device to execute the Figure 1 instructions for the steps included in the feature distillation method based on multi-model fusion in the present application.

[0056] Figure 3Schematic diagram of the hardware structure of a terminal device provided in an embodiment of the present application. As shown in the figure, the terminal device may include: an input device 1100, a first processor 1101, an output device 1102, a first memory 1103, and at least one communication bus 1104. The communication bus 1104 is used to implement communication connections between components. The first memory 1103 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory. Various programs may be stored in the first memory 1103 to complete various processing functions and implement the method steps of this embodiment.

[0057] Optionally, the above-mentioned first processor 1101 may be implemented, for example, as a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components. The processor 1101 is coupled to the above-mentioned input device 1100 and output device 1102 through a wired or wireless connection.

[0058] Optionally, the above-mentioned input device 1100 may include a variety of input devices. For example, it may include at least one of a user interface for users, a device interface for devices, a programmable interface for software, a camera, and a sensor. Optionally, the device interface for devices may be a wired interface for data transmission between devices, or may also be a hardware insertion interface for data transmission between devices (such as a USB interface, a serial port, etc.); Optionally, the user interface for users may be, for example, control buttons for users, a voice input device for receiving voice input, and a touch sensing device for users to receive user touch input (such as a touch screen with touch sensing function, a touchpad, etc.); Optionally, the above-mentioned programmable interface for software may be, for example, an entry for users to edit or modify programs, such as an input pin interface or an input interface of a chip, etc.; The output device 1102 may include output devices such as a display and a speaker.

[0059] In this embodiment, the processor of the terminal device includes functions for executing each module of the voice recognition device in each device. For the specific functions and technical effects, refer to the above-mentioned embodiment, and details are not described here again.

[0060] Figure 4 Schematic diagram of the hardware structure of a terminal device provided in another embodiment of the present application. Figure 4 is a specific embodiment in the Figure 3 implementation process. As shown in the figure, the terminal device of this embodiment may include a second processor 1201 and a second memory 1202.

[0061] The second processor 1201 executes the computer program code stored in the second memory 1202 to implement the method in the above embodiments Figure 1 described above.

[0062] The second memory 1202 is configured to store various types of data to support the operation of the terminal device. Examples of such data include instructions for any application or method operating on the terminal device, such as messages, pictures, videos, etc. The second memory 1202 may include a random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0063] Optionally, the first processor 1201 is disposed in the processing component 1200. The terminal device may further include: a communication component 1203, a power supply component 1204, a multimedia component 1205, an audio component 1206, an input / output interface 1207, and / or a sensor component 1208. The specific components included in the terminal device are set according to actual requirements, and this embodiment does not limit this.

[0064] The processing component 1200 generally controls the overall operation of the terminal device. The processing component 1200 may include one or more second processors 1201 to execute instructions to complete all or part of the steps of the method Figure 1 shown above. In addition, the processing component 1200 may include one or more modules to facilitate the interaction between the processing component 1200 and other components. For example, the processing component 1200 may include a multimedia module to facilitate the interaction between the multimedia component 1205 and the processing component 1200.

[0065] The power supply component 1204 provides power to various components of the terminal device. The power supply component 1204 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the terminal device.

[0066] The multimedia component 1205 includes a display screen that provides an output interface between the terminal device and the user. In some embodiments, the display screen may include a liquid crystal display (LCD) and a touch panel (TP). If the display screen includes a touch panel, the display screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.

[0067] The audio component 1206 is configured to output and / or input voice signals. For example, the audio component 1206 includes a microphone (MIC). When the terminal device is in an operating mode, such as a voice recognition mode, the microphone is configured to receive external voice signals. The received voice signals can be further stored in the second memory 1202 or transmitted via the communication component 1203. In some embodiments, the audio component 1206 further includes a speaker for outputting voice signals.

[0068] The input / output interface 1207 provides an interface between the processing component 1200 and a peripheral interface module, which can be a click wheel, buttons, etc. These buttons can include, but are not limited to: volume buttons, start buttons, and lock buttons.

[0069] The sensor component 1208 includes one or more sensors for providing status assessments of various aspects of the terminal device. For example, the sensor component 1208 can detect the on / off state of the terminal device, the relative positioning of components, the presence or absence of user contact with the terminal device. The sensor component 1208 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact, including detecting the distance between the user and the terminal device. In some embodiments, the sensor component 1208 can further include a camera, etc.

[0070] The communication component 1203 is configured to facilitate communication between the terminal device and other devices in a wired or wireless manner. The terminal device can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In one embodiment, the terminal device can include a SIM card slot for inserting a SIM card, enabling the terminal device to log in to the GPRS network and establish communication with a server via the Internet.

[0071] As can be seen from the above, in Figure 4 the embodiments, the communication component 1203, the audio component 1206, as well as the input / output interface 1207 and the sensor component 1208 can all be used as Figure 3 implementation manners of the input device in the embodiments.

[0072] The above embodiments merely illustrate the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A feature distillation method based on multi-model fusion, characterized in that, Including: Obtaining the features of the target data as the first features respectively through multiple pre-trained teacher models; the target data includes face images and vehicle images of the object to be recognized; The student model may include a backbone network, a local distillation module, and a global distillation module. Among them, the local distillation module includes multiple first distillation sub-networks. The global distillation module splices and fuses the first features of the target data obtained by each teacher model to obtain a first fused feature; obtaining the second feature of the target data through the backbone network of the student model, inputting the second feature into multiple first distillation sub-networks respectively, and outputting, through each of the first distillation sub-networks, a second feature whose similarity with the first feature reaches a set threshold; Fusing the second features output by each of the distillation sub-networks to obtain a second fused feature, inputting the first fused feature and the second fused feature into a second distillation sub-network, and obtaining the distillation feature of the target data; the second distillation sub-network has the same network structure as the first distillation sub-network.

2. The feature distillation method based on multi-model fusion according to claim 1, wherein The first distillation sub-network includes: an attention module, a normalization layer, a similarity calculation layer, and at least one fully connected layer; The attention module obtains the weight of the corresponding feature according to the feature value size of the feature output by the fully connected layer and outputs it to the normalization layer; The normalization layer completes the normalization of the corresponding feature according to the feature output by the fully connected layer and the weight output by the attention module; The similarity calculation layer obtains the similarity between the normalized feature and the first feature output by the corresponding teacher model through a preset loss function.

3. The feature distillation method based on multi-model fusion according to claim 2, wherein The attention module maps the feature value to between -1 and 1 through a mapping function.

4. The feature distillation method based on multi-model fusion according to claim 3, wherein The mapping function includes: softmax function, sigmoid function.

5. The feature distillation method based on multi-model fusion according to claim 1, wherein Before inputting the first fused feature and the second fused feature into the second distillation sub-network, it further includes: Performing dimensionality reduction processing on the first fused feature by using a dimensionality reduction algorithm.

6. The feature distillation method based on multi-model fusion according to claim 1, characterized in that The number of the first distillation sub-networks corresponds to the number of the teacher models, and each first distillation sub-network receives the first feature of one of the teacher models respectively.

7. A feature distillation system based on multi-model fusion for implementing the feature distillation method based on multi-model fusion according to any one of claims 1-6, characterized in that, Including: A first feature acquisition module, configured to obtain the features of the target data as the first features respectively through multiple pre-trained teacher models; A student feature acquisition module, configured to obtain the second feature of the target data through the backbone network of the student model, input the second feature into multiple first distillation sub-networks respectively, and output, through each of the first distillation sub-networks, a second feature whose similarity with the first feature reaches a set threshold; A fusion distillation module, configured to fuse all the first features to obtain a first fused feature, fuse the second features output by each of the distillation sub-networks to obtain a second fused feature, input the first fused feature and the second fused feature into a second distillation sub-network, and obtain the distillation feature of the target data; The second distillation sub-network has the same network structure as the first distillation sub-network.

8. An electronic device, characterized in that, Including: One or more processors; And One or more machine-readable media storing instructions which, when executed by the one or more processors, cause the device to perform the multi-model fusion-based feature distillation method according to any one of claims 1-7.

9. A machine-readable medium, characterized in that, One or more machine-readable media storing instructions which, when executed by one or more processors, cause the device to perform the multi-model fusion-based feature distillation method according to one or more of claims 1-7.

Citation Information

Patent Citations

  • Space-time double-flow segmented network behavior identification method and system based on knowledge distillation

    CN112446331A

  • Target detection knowledge distillation method for self-adaptive region refinement

    CN112766411A