Model training method based on knowledge distillation and electronic equipment
Through multiple teacher models extracting features and distilling prospects and background knowledge with student models, the problem of poor model migration performance is solved, and the concurrency ability and accuracy of the model is improved under limited samples.
Patent Information
- Application Number
- CN202510027202.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-17
AI Technical Summary
In scenarios where model migration is based on knowledge distillation, the model obtained by migration is poor.
Feature extraction of the first sample image through multiple teacher models is obtained, and the foreground supervision sub-features and background supervision sub-features of N local category objects are obtained. These features are distilled foreground knowledge and distilled background knowledge to generate a second student model, and image prediction of the second sample image through this model is used to adjust the model parameters.
This improves the generalization ability of students' models, so that they can distillate models with strong concurrency and high accuracy under the premise of a limited number of samples.
Smart Images

Figure CN120164075A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a model training method and electronic device based on knowledge distillation. Background Art
[0002] In the industry application of AI algorithms, small models have strong concurrency capabilities and low resource usage, which are suitable for real-time video streaming. However, small models have weak generalization capabilities and require the collection of a large number of private industry samples for model training, resulting in high data collection costs. In contrast, large models have weak concurrency capabilities and high resource usage, but they have the advantages of strong generalization capabilities and high precision. With the help of knowledge distillation technology, the migration of large model knowledge to small models can be completed at a low cost, thereby improving the accuracy of small models under the premise of limited industry samples. Therefore, how to improve the effect of migrating large model knowledge to small models has become one of the problems that need to be solved in the industry. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a model training method and electronic device based on knowledge distillation, so as to solve the technical problem that the performance of the migrated model is poor in the scenario of model migration based on knowledge distillation.
[0004] To solve the above technical problems, the embodiments of the present application are implemented as follows: On the one hand, an embodiment of the present application provides a model training method based on knowledge distillation, comprising: Performing feature extraction on the first sample image through multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; Performing feature extraction on the first sample image by using a first student model to obtain first foreground features and first background features of the N local category objects; According to the N foreground supervision sub-features and the first foreground feature, the first student model is subjected to foreground knowledge distillation, and according to the background supervision sub-features and the first background feature, the first student model is subjected to background knowledge distillation; the first student model after the foreground knowledge distillation and the background knowledge distillation is the second student model; The second sample image is predicted by the second student model, and the model parameters of the second student model are adjusted according to the prediction result of the image prediction.
[0005] On the other hand, an embodiment of the present application provides a model training device based on knowledge distillation, comprising: A first extraction module, configured to extract features from a first sample image through multiple teacher models, to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; A second extraction module, configured to extract features from the first sample image through a first student model, to obtain a first foreground feature and a first background feature of the N local category objects; A knowledge distillation module, configured to perform foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground feature, and to perform background knowledge distillation on the first student model according to the background supervision sub-features and the first background feature; the first student model after the foreground knowledge distillation and the background knowledge distillation is a second student model; An adjustment module, configured to perform image prediction on a second sample image through the second student model, and to adjust model parameters of the second student model according to a prediction result of the image prediction.
[0006] In another aspect, an embodiment of the present application provides an electronic device, including a processor and a memory electrically connected to the processor, where the memory stores a computer program, and the processor is configured to call and execute the computer program from the memory to implement the above-mentioned model training method based on knowledge distillation.
[0007] In another aspect, an embodiment of the present application provides a computer-readable storage medium, configured to store a computer program, and the computer program can be executed by a processor to implement the above-mentioned model training method based on knowledge distillation.
[0008] In another aspect, an embodiment of the present application provides a computer program product, including a computer program, and the computer program is executed by a processor to implement the above-mentioned model training method based on knowledge distillation.
[0009] Adopting the technical solution of the embodiment of the present application, multiple teacher models are used to extract features from the first sample image, obtaining foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; the first student model is used to extract features from the first sample image, obtaining the first foreground feature and the first background feature of N local category objects. Then, foreground knowledge distillation is performed on the first student model according to the N foreground supervision sub-features and the first foreground feature, and background knowledge distillation is performed on the first student model according to the background supervision sub-features and the first background feature, obtaining the second student model, and the second student model is used to perform image prediction on the second sample image, and the model parameters of the second student model are adjusted according to the prediction result of the image prediction. It can be seen that when performing model transfer based on knowledge distillation, foreground knowledge distillation and background knowledge distillation can be carried out synchronously, enabling the student model to not only learn the foreground knowledge in the sample image but also learn the low-heat background knowledge during the knowledge distillation process, thus helping to improve the generalization ability of the student model. In addition, since during the knowledge distillation process, features of multiple local category objects in the sample image can be extracted, including foreground supervision sub-features, background supervision sub-features, the first foreground feature, and the first background feature, and knowledge distillation is performed based on the features of multiple local category objects, the student model can learn multiple categories of knowledge simultaneously. That is to say, a student model can learn cross-category cross knowledge, so that the distilled student model has the ability to identify different category objects, achieving the effect of distilling a student model with strong concurrency and high accuracy under the premise of a limited number of samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions in one or more embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in one or more embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] Figure 1 is a schematic flowchart of a model training method based on knowledge distillation according to an embodiment of the present application; Figure 2 is a schematic application scenario diagram of a model training method based on knowledge distillation according to an embodiment of the present application; Figure 3 is a schematic scenario diagram of a model training method based on knowledge distillation according to an embodiment of the present application; Figure 4 is a schematic effect diagram of deleting redundant sub-features according to an embodiment of the present application; Figure 5 It is a schematic effect diagram of adding missing sub - features according to an embodiment of the present application; Figure 6 It is a schematic scenario diagram of a model training method based on knowledge distillation according to another embodiment of the present application; Figure 7 It is a schematic principle diagram of a model training method based on knowledge distillation according to an embodiment of the present application; Figure 8 It is a schematic block diagram of a model training device based on knowledge distillation according to an embodiment of the present application; Figure 9 It is a schematic block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0012] Embodiments of the present application provide a model training method and an electronic device based on knowledge distillation to solve the technical problem that in the scenario of model migration based on knowledge distillation, the performance of the migrated model is poor.
[0013] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0014] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein.
[0015] The model training method based on knowledge distillation provided by the embodiments of the present application can be executed by an electronic device or by software installed in the electronic device. Specifically, the electronic device can be a terminal device or a server device. Among them, the terminal device can include a smart phone, a laptop computer, a smart wearable device, a vehicle-mounted terminal, etc., and the server device can include an independent physical server, a server cluster composed of multiple servers, or a cloud server capable of performing cloud computing.
[0016] Figure 1 It is a schematic flowchart of a model training method based on knowledge distillation according to an embodiment of the present application. As Figure 1 shown, the method includes: Step S102: Use multiple teacher models to extract features from the first sample image, obtaining foreground supervision sub-features and background supervision sub-features of N local category objects.
[0017] Here, N is an integer greater than 1. The first sample image includes multiple local category objects, where a local category object refers to a partial category object among the multiple category objects in the first sample image. For example, if the first sample image includes a person and a vehicle, then the person and the vehicle are each a local category object in the first sample image. The foreground supervision sub-feature of a local category object is the foreground feature of the local category object, and the background supervision sub-feature of the local category object is the background feature of the local category object, both of which play a supervisory role in the knowledge distillation process.
[0018] The number of teacher models matches the number of categories of local category objects (i.e., the value of N). Optionally, the number of teacher models is N, and each teacher model is respectively used to extract the foreground supervision sub-feature and the background supervision sub-feature of a local category object. For example, if the first sample image includes two local category objects, a person and a vehicle, teacher model 1 has the ability to recognize a person but not a vehicle, and teacher model 2 has the ability to recognize a vehicle but not a person. Therefore, when performing knowledge distillation, both teacher model 1 and teacher model 2 need to be used, that is, the number of teacher models is 2.
[0019] Optionally, the number of teacher models is less than N, and one teacher model can be used to extract the foreground supervision sub-feature and the background supervision sub-feature of one or more local category objects. For example, if the first sample image includes three local category objects, a person, a vehicle, and a pet, teacher model 1 has the ability to recognize a person but not a vehicle or a pet, teacher model 2 has the ability to recognize a vehicle but not a person or a pet, and teacher model 3 has the ability to recognize a vehicle and a pet but not a person. Therefore, when performing knowledge distillation, teacher model 1 and teacher model 3 can be used simultaneously. Of course, teacher model 1, teacher model 2, and teacher model 3 can also be used simultaneously, as long as the local category objects recognized by the teacher models can cover all the local category objects required for model training.
[0020] The multiple teacher models can be models across industries. For example, the multiple teacher models include a campus teacher model, a traffic teacher model, etc. The campus teacher model can be used to extract people in the sample image, and the traffic teacher model can be used to extract vehicles in the sample image. This enables the student model to learn cross-industry knowledge, concentrating the object recognition capabilities of multiple different industries in one model, and greatly improving the distillation effect of the model.
[0021] Step S104: Extract features from the first sample image through the first student model to obtain the first foreground features and the first background features of N local category objects.
[0022] Among them, the N local category objects extracted by the first student model and the teacher model are the same.
[0023] Step S106: Perform foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground features, and perform background knowledge distillation on the first student model according to the background supervision sub-features and the first background features; the first student model after foreground knowledge distillation and background knowledge distillation is the second student model.
[0024] Among them, foreground knowledge distillation and background knowledge distillation are carried out synchronously. The second student model refers to the intermediate model obtained after knowledge distillation (including foreground knowledge distillation and background knowledge distillation) of the first student model. At this time, it is not yet the expected final student model, and the model parameters of the second student model need to be further optimized.
[0025] In supervised learning, background features are a supplement to the content of foreground features and play a key role in reducing model false alarms and improving model generalization. The teacher model itself has strong generalization ability. In the scenario of a small number of samples for distillation, the background feature distillation of the teacher model can achieve the transfer of the generalization ability from the teacher model to the student model.
[0026] Step S108: Perform image prediction on the second sample image through the second student model, and adjust the model parameters of the second student model according to the prediction result of the image prediction.
[0027] Optionally, the prediction result of the image prediction includes a foreground prediction result, that is, predicting the position of the local category object in the second sample image. Step S108 can be executed as the following steps: First, perform foreground prediction on the second sample image through the second student model to obtain a foreground prediction result containing N local category objects. The second sample image and the first sample image can be the same or different. Optionally, the pre-collected sample images are divided into a training set and a test set at a specific ratio. The sample images in the training set are used as the first sample images, and the sample images in the test set are used as the second sample images.
[0028] Second, adjust the model parameters of the second student model according to the foreground prediction result and the label information of the second sample image.
[0029] Figure 2It is a schematic application scenario diagram of a model training method based on knowledge distillation according to an embodiment of the present application. In this application scenario, N teacher models are used to perform knowledge distillation on a student model. The N teacher models include Teacher Model 1, Teacher Model 2, …… Teacher Model N. Due to space limitations, Figure 2 only Teacher Model 1 and Teacher Model N are exemplarily shown. As Figure 2 described, the first sample image is input into the N teacher models and the first student model at the same time. The N teacher models respectively extract features of N local category objects in the first sample image to obtain teacher features corresponding to each teacher model, such as Teacher Feature 1, Teacher Feature 2, …… Teacher Feature N. After these teacher features are combined, they include foreground supervision sub-features and background supervision sub-features of the N local category objects. At the same time, the first student model extracts features of the N local category objects in the first sample image to obtain student features, including the first foreground feature and the first background feature of the N local category objects. Then, foreground knowledge distillation is performed on the first student model according to the N foreground supervision sub-features and the first foreground feature, and background knowledge distillation is performed on the first student model according to the background supervision sub-features and the first background feature to obtain a second student model. Furthermore, the second student model performs image prediction on the second sample image, and according to the prediction result of the image prediction, the model parameters of the second student model are adjusted.
[0030] In this embodiment, the teacher model can be a pre-trained open-source large model. During the training process of the student model, the parameters of the teacher model can be frozen, so as to extract the features of each sample image and use them as the supervision information for the student model to extract features, enabling the student model to learn comprehensive and deep knowledge. In addition, during the knowledge distillation process of the student model, the training image labels are used to supervise the prediction results of the model, so as to complete the solution of the loss function and perform gradient backpropagation.
[0031] When collecting sample images in advance, industry sample images publicly available on the network can be collected. To ensure the diversity of sample types, images with different shooting angles, lighting, and background environments can be collected as sample images, and the sample images should contain all local category objects to be recognized. For example, if it is desired to train a student model with the ability to recognize N local category objects, then the sample images should include these N local category objects.
[0032] Adopting the technical solution of the embodiment of the present application, the first sample image is subjected to feature extraction by multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; the first student model is used to perform feature extraction on the first sample image to obtain the first foreground feature and the first background feature of N local category objects. Furthermore, foreground knowledge distillation is performed on the first student model according to the N foreground supervision sub-features and the first foreground feature, and background knowledge distillation is performed on the first student model according to the background supervision sub-features and the first background feature to obtain a second student model, and the second student model is used to perform image prediction on the second sample image, and the model parameters of the second student model are adjusted according to the prediction result of the image prediction. It can be seen that when performing model transfer based on knowledge distillation, foreground knowledge distillation and background knowledge distillation can be performed synchronously, so that the student model can not only learn the foreground knowledge in the sample image but also learn the low-heat background knowledge during the knowledge distillation process, which helps to improve the generalization ability of the student model. In addition, since multiple local category object features in the sample image can be extracted during the knowledge distillation process, including foreground supervision sub-features, background supervision sub-features, first foreground features, and first background features, and knowledge distillation is performed based on the features of multiple local category objects, the student model can learn multiple categories of knowledge at the same time. That is to say, a student model can learn cross-category cross knowledge, so that the distilled student model has the ability to recognize different category objects, achieving the effect of distilling a student model with strong concurrency and high accuracy under the premise of a limited number of samples.
[0033] Figure 3 FIG. is a schematic scenario diagram of a model training method based on knowledge distillation according to an embodiment of the present application. In this embodiment, taking N = 2 as an example, in the pre-constructed knowledge distillation network architecture, there are teacher model 1, teacher model 2, and student model. Among them, teacher model 1 provides the vehicle detection ability in the transportation field, that is, it has the ability to recognize vehicles. Teacher model 2 provides the human detection ability for park management and control, that is, it has the ability to recognize people. As Figure 3 shown, the first sample image includes people and vehicles, and the label information of the first sample image includes the position coordinates of people and vehicles. The first sample image is input into teacher model 1, teacher model 2, and student model respectively. Teacher model 1 extracts the foreground supervision sub-feature of the vehicle, and teacher model 2 extracts the foreground supervision sub-feature of the person. The student model extracts the first foreground features of people and vehicles. Knowledge distillation is performed on the student model based on the foreground supervision sub-feature of the vehicle and the foreground supervision sub-feature of the person, and finally the student model is enabled to have the ability to extract people and vehicles at the same time. Thus, during the knowledge distillation process, cross-domain cross-empowerment from the transportation field to the park management and control field is provided for the student model, and a high-precision and strong-concurrency student model can be quickly generated under the premise of a small number of samples.
[0034] In one embodiment, before performing step S102, multiple teacher models are pre-configured, and a teacher model that matches the training scenario of the student model can be selected from the multiple pre-configured teacher models. Optionally, the model categories corresponding to the pre-configured teacher models are compared with the label information of the first sample image to obtain a first comparison result, and based on the first comparison result, multiple teacher models that match the N local category objects are selected from the pre-configured teacher models. Each pre-configured teacher model of each model category is used to identify at least one local category object.
[0035] Among them, the label information of the first sample image includes: the position features of the N local category objects in the first sample image. The position features can be characterized in the form of position coordinates. By adding label information to the first sample image, it is possible to clearly know the position where the local category object in the first sample image is located. Optionally, the position features of the local category object in the first sample image are characterized as the position coordinates of the four key points (i.e., vertices) of the rectangular box, and the rectangular box can be a line box that contains the local category object and is close to the contour of the local category object. Of course, the local category object in the present application is not limited to the form of a rectangular box, and it can also be in any other type of form, such as a circle, a square, an ellipse, etc.
[0036] The multiple pre-configured teacher models can be open-source large models across industries or in the same industry. The pre-configured teacher models should cover the recognition capabilities of various types of objects as much as possible to ensure the diversity of object types. The label information of the first sample image can be labeled using any existing annotation software (such as labelImg).
[0037] There are various ways to annotate the label information in the first sample image. Optionally, in the first sample image, the N local category objects are annotated in the form of a rectangular box. Or, the position coordinates of each local category object are associated with the first sample image in a set form, and the set includes the position coordinates of the four vertices of the rectangular box corresponding to the local category object. For example, in the label information, the local category object is represented by the position coordinates [Xmin, Ymin, Xmax, Ymax]. Xmin and Xmax are used to determine the projection position of the rectangular box on the horizontal axis of the coordinate axis, and Ymin and Ymax are used to determine the projection position of the rectangular box on the vertical axis of the coordinate axis. Based on the label information, it is possible to realize the mapping of the corresponding feature regions of the local category objects in each hierarchical feature image, so as to find and peel off the foreground features of each local category object according to the mapped coordinates.
[0038] In this embodiment, by screening out multiple teacher models that match N local category objects from the pre-configured teacher models, the multiple screened teacher models can accurately extract the foreground supervision sub-features of the N local category objects, thereby providing accurate supervision information for the knowledge distillation of the student model.
[0039] In one embodiment, when multiple teacher models are used to extract features from the first sample image to obtain the foreground supervision sub-features and background supervision sub-features of N local category objects, the following method can be executed: in the case where the local category objects that match the model category of the teacher model are included in the N local category objects, the teacher model is used to extract features from the matching local category objects to obtain the foreground supervision sub-features of the matching local category objects.
[0040] Among them, the model category of the teacher model is divided according to the local category objects that the teacher model can recognize. For example, if teacher model 1 can recognize local category object A, then it can be considered that the model category of teacher model 1 is A. The local category objects that match the model category of the teacher model refer to the local category objects that the teacher model can recognize. For example, if teacher model 1 can recognize local category object A, then teacher model 1 matches local category object A.
[0041] In one embodiment, when multiple teacher models are used to extract features from the first sample image, the following steps can be executed: extract features from the local category objects that match the model category of the teacher model in the first sample image to obtain the foreground supervision sub-features of M local category objects, where M is an integer greater than 1. In the case where M is not equal to N, feature alignment processing is performed on the M foreground supervision sub-features according to the label information of the first sample image to obtain N foreground supervision sub-features. The label information of the first sample image includes: the position features of the N local category objects in the first sample image. That is, the label information of the first sample image is used to indicate the N local category objects.
[0042] In this embodiment, the number of local category objects that the multiple teacher models can cover may be different from the number of local category objects included in the first sample image. The local category objects that the multiple teacher models can cover refer to the number of local category objects that the multiple teacher models can recognize from the first sample image. According to the label information of the first sample image, the local category objects that need to be recognized can be determined. Therefore, by performing feature alignment processing on the M foreground supervision sub-features according to the label information of the first sample image, the foreground supervision sub-features corresponding to the N local category objects can be obtained.
[0043] The feature alignment processing method may include: deleting some foreground supervised sub-features, or adding some foreground supervised sub-features. When M is greater than N, the feature alignment processing method is to delete some foreground supervised sub-features. When M is less than N, the feature alignment processing method is to add some foreground supervised sub-features.
[0044] Optionally, when M is greater than N, redundant sub-features among the M foreground supervised sub-features are determined according to the M foreground supervised sub-features and the label information of the first sample image; the redundant sub-features are deleted from the M foreground supervised sub-features to obtain N foreground supervised sub-features.
[0045] Among them, by comparing the local category objects corresponding to the M foreground supervised sub-features with the N local category objects indicated by the label information of the first sample image, the redundant sub-features among the M foreground supervised sub-features can be determined. For example, multiple teacher models extract foreground supervised sub-features of three local category objects: person, vehicle, and pet, that is, M = 3. The label information of the first sample image indicates two local category objects: person and vehicle, that is, N = 2. By comparing the local category objects corresponding to the three foreground supervised sub-features extracted by the multiple teacher models with the two local category objects indicated by the label information of the first sample image, it can be determined that the redundant sub-feature is the foreground supervised sub-feature of the local category object "pet". In addition to the foreground supervised sub-features of redundant local category objects, the redundant sub-features may also include the foreground supervised sub-features of overlapping local category objects. For example, teacher model 1 and teacher model 2 have the ability to extract features of the same local category object, and both extract the foreground supervised sub-feature of this local category object during the distillation process. Then, during the feature alignment processing, one of the foreground supervised sub-features can be deleted to avoid the foreground supervised sub-feature of the same local category object being calculated multiple times.
[0046] Figure 4 It is a schematic effect diagram of deleting redundant sub-features according to an embodiment of the present application, as Figure 4 shown, the first sample image includes a person and a vehicle, and multiple teacher models extract the foreground supervised sub-features of the person and the vehicle. Assume that the label information of the first sample image only includes the position coordinates of the person, that is, it is expected that the extracted features only include the features of the local category object "person". Therefore, the foreground supervised sub-features extracted by the multiple teacher models include redundant sub-features, that is, the features of the local category object "vehicle". By deleting the foreground supervised sub-feature of the local category object "vehicle", the feature alignment processing can be achieved, ensuring that the student model learns the object features consistent with the label information, that is, only learns the detection ability of people.
[0047] When M is less than N, based on the M foreground supervised sub-features and the label information of the first sample image, determine the missing sub-features among the N foreground supervised sub-features; extract features from the first sample image through a teacher model that matches the missing sub-features to obtain the missing sub-features.
[0048] Among them, by comparing the local category objects corresponding to the M foreground supervised sub-features with the N local category objects indicated by the label information of the first sample image, the missing sub-features among the M foreground supervised sub-features can be determined. For example, multiple teacher models extract the foreground supervised sub-features of two local category objects, namely a vehicle and a pet, that is, M = 2. The label information of the first sample image indicates three local category objects: a person, a vehicle, and a pet, that is, N = 3. By comparing the local category objects corresponding to the two foreground supervised sub-features extracted by the multiple teacher models with the three local category objects indicated by the label information of the first sample image, it can be determined that the missing sub-feature is the foreground supervised sub-feature of the local category object "person".
[0049] When extracting features from the first sample image through a teacher model that matches the missing sub-features to obtain the missing sub-features, if the pre-constructed knowledge distillation network framework includes a teacher model that matches the missing sub-features, use this teacher model to extract features from the first sample image to obtain the missing sub-features. If the pre-constructed knowledge distillation network framework does not include a teacher model that matches the missing sub-features, supervised learning can be performed according to the label information of the first sample image, so as to improve the missing foreground supervised sub-features through multiple rounds of iteration.
[0050] Figure 5 is a schematic effect diagram of adding missing sub-features according to an embodiment of the present application. As Figure 5 shown, the first sample image includes a person and a vehicle, and multiple teacher models only extract the foreground supervised sub-feature of the person. Assume that the label information of the first sample image includes the position coordinates of the person and the vehicle, that is to say, the features to be extracted include the features of the local category object "person" and the features of the local category object "vehicle". Therefore, there are missing sub-features among the foreground supervised sub-features extracted by the multiple teacher models, that is, the features of the local category object "vehicle". By adding the foreground supervised sub-feature of the local category object "vehicle", feature alignment processing can be achieved, ensuring that the student model learns object features consistent with the label information, that is, learning the detection capabilities of both the person and the vehicle at the same time.
[0051] In this embodiment, regardless of whether there are redundant or missing foreground supervision sub-features extracted by multiple teacher models, the M foreground supervision sub-features can be aligned into N foreground supervision sub-features through feature alignment processing, so as to ensure that the student model can learn the feature information of N local category objects during knowledge distillation, which is beneficial to distilling a student model with strong concurrency ability and high accuracy.
[0052] In one embodiment, when multiple teacher models are used to extract features from the first sample image to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, the following steps A1 - A2 can be executed: Step A1, use the teacher model to extract features from the first sample image to obtain foreground supervision sub-features of local category objects that match the model category of the teacher model.
[0053] Step A2, remove the first image region corresponding to the foreground supervision sub-feature and the second image region corresponding to the redundant foreground feature from the first sample image to obtain the background supervision sub-feature. The second image region is the image region corresponding to the local category object that the teacher model did not extract.
[0054] In this embodiment, for each teacher model, the foreground supervision sub-feature and the background supervision sub-feature of the corresponding local category object can be obtained by executing steps A1 - A2.
[0055] Figure 6 It is a schematic scenario diagram of a model training method based on knowledge distillation according to another embodiment of the present application. Figure 6 Compared with Figure 3 the embodiment shown, a background knowledge distillation module is added. The background feature region for distillation is the image region obtained after removing the first image region corresponding to the foreground supervision sub-feature and the second image region corresponding to the redundant foreground feature. The first image region corresponding to the foreground supervision sub-feature is the image region where the local category object that matches the model category of the teacher model is located. The second image region corresponding to the redundant foreground feature is the image region where the local category object that does not match the model category of the teacher model is located. Optionally, if the first sample image also includes local category objects that do not match the label information, that is, local category objects that are not expected to be extracted, then when determining the background feature region, the image region where the local category objects that are not expected to be extracted are located needs to be removed, that is, the image region where the local category objects that are not expected to be extracted are located is also considered as the background feature region.
[0056] For example, the first sample image includes a person and a vehicle. The teacher model 1 provides the vehicle detection ability in the transportation field, that is, the ability to recognize vehicles. Then, the local category object matching the model category of the teacher model 1 is the vehicle, and the local category object not matching the model category of the teacher model 1 is the person. When determining the background supervision sub-feature corresponding to the teacher model 1, the first image region corresponding to the foreground supervision sub-feature, that is, the image region where the vehicle is located, is removed from the first sample image, and the second image region corresponding to the redundant foreground feature, that is, the image region where the person is located, is removed, and then the background feature region can be obtained. The teacher model extracts features from the background feature region to obtain the background supervision sub-feature corresponding to the teacher model. Each teacher model extracts the background supervision sub-feature in the same way, so as to obtain N background supervision sub-features.
[0057] As Figure 6 shown, the teacher model 1 extracts the foreground supervision sub-feature and the background supervision sub-feature of the vehicle, the teacher model 2 extracts the foreground supervision sub-feature and the background supervision sub-feature of the person, and the student model extracts the first foreground feature and the first background feature of the vehicle and the person. All the foreground features and background features extracted by the teacher model 1 and the teacher model 2 are used as the supervision information of the student model, so as to realize the synchronous foreground feature distillation and background feature distillation. It can be seen that when the teacher model 1, the teacher model 2 and the student model extract the background features, the corresponding background feature regions are the same, which are the image regions obtained after deleting the full foreground feature region from the first sample image, and the full foreground feature region is the image region where all the expected local category objects are located.
[0058] In one embodiment, when performing foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground feature, first, the N foreground supervision sub-features are concatenated to obtain the foreground supervision feature of N local category objects; secondly, according to the difference degree between the foreground supervision feature and the first foreground feature, the distillation foreground loss function of the first student model is calculated, and the model parameters of the first student model are adjusted according to the distillation foreground loss function.
[0059] Optionally, the foreground distillation loss function can be expressed by the following formula: (1) where represents the pixel value of a certain pixel point in the foreground supervision feature extracted by the teacher model, represents the pixel value of a certain pixel point in the first foreground feature extracted by the student model, C represents the number of channels of the foreground supervision feature and the first foreground feature. n represents the feature level of the feature extraction network in the teacher model and the student model.
[0060] Similarly, when performing background knowledge distillation on the first student model based on the background supervision sub-features and the first background feature, first, the N background supervision sub-features are concatenated to obtain the background supervision features of N local class objects; second, according to the difference degree between the background supervision features and the first background feature, the distillation background loss function of the first student model is calculated, and the model parameters of the first student model are adjusted according to the distillation background loss function.
[0061] The representation of the background distillation loss function and the foreground distillation loss function is similar, as shown in the following formula (2): (2) Among them, represents the pixel value of a certain pixel point in the background supervision feature extracted by the teacher model, represents the pixel value of a certain pixel point in the first background feature extracted by the student model, C represents the number of channels of the background supervision feature and the first background feature. n represents the feature level of the feature extraction network in the teacher model and the student model.
[0062] In one embodiment, when performing foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground feature, the following steps B1 - B3 are executed: Step B1, for each foreground supervision sub-feature, determine the first foreground sub-feature in the first foreground feature that matches the foreground supervision sub-feature.
[0063] Step B2, according to the difference degree between the foreground supervision sub-feature and the first foreground sub-feature, calculate the local foreground loss function of the first student model. Since there are N foreground supervision sub-features, N local foreground loss functions can be obtained through calculation.
[0064] Step B3, according to the N local foreground loss functions, determine the distillation foreground loss function of the first student model, and adjust the model parameters of the first student model according to the distillation foreground loss function.
[0065] Optionally, the distillation foreground loss function of the first student model is equal to the sum of the N local foreground loss functions. When executing step B2, each local foreground loss function can be calculated in the manner of the above formula (1).
[0066] When performing background knowledge distillation on the first student model according to the background supervision sub-features and the first background feature, the following steps C1 - C3 are executed: Step C1, for each background supervision sub-feature, determine the first background sub-feature in the first background feature that matches the background supervision sub-feature.
[0067] Step C2: Calculate the local background loss function of the first student model according to the difference degree between the background supervision sub-features and the first background sub-features. Since there are N background supervision sub-features, N local background loss functions can be obtained through calculation.
[0068] Step C3: Determine the distilled background loss function of the first student model according to the N local background loss functions, and adjust the model parameters of the first student model according to the distilled background loss function.
[0069] Optionally, the distilled background loss function of the first student model is equal to the sum of the N local background loss functions. When performing Step C2, each local background loss function can be calculated in the manner of the above formula (2).
[0070] In this embodiment, by calculating the local foreground loss function and the local background loss function corresponding to each local category object respectively, and then determining the total distilled foreground loss function according to the N local foreground loss functions, and determining the total distilled background loss function according to the N local background loss functions, the effective calculation of the loss function in the scenario of distilling one student model by using multiple teacher models is realized. After calculating the distilled foreground loss function and the distilled background loss function, gradient backpropagation can be performed based on the distilled foreground loss function and the distilled background loss function to realize the knowledge transfer from multiple teacher models to one student model.
[0071] In one embodiment, the prediction result of the image prediction includes a foreground prediction result, and the foreground prediction result includes: the predicted position information of N local category objects in the second sample image. The label information of the second sample image includes: the reference position information of N local category objects in the second sample image. The reference position information is the correct position information of the local category object in the second sample image, which can be represented as position coordinates.
[0072] When adjusting the model parameters of the second student model according to the foreground prediction result and the label information of the second sample image, the supervision loss function of the second sample image can be determined first according to the difference degree between the predicted position information and the reference position information. Then, the model parameters of the second student model are adjusted according to the supervision loss function.
[0073] In one embodiment, since the foreground feature and the background feature are separated, an imbalance may occur between the foreground feature and the background feature. To alleviate this imbalance problem, it can be eliminated by preprocessing the global features (including foreground features and background features) of the first sample image. Optionally, by guiding the student model to capture and enhance the relationship between image channels and the position relationship between pixel points, and performing knowledge distillation on the enhanced feature image. Specifically, it includes the following steps D1-D4: Step D1: Determine the global supervision features of multiple teacher models based on N foreground supervision sub - features and background supervision sub - features. Also, determine the global student features of the second student model based on the first foreground feature and the first background feature.
[0074] Step D2: Perform feature enhancement processing on the global supervision features to obtain enhanced global supervision features.
[0075] Step D3: Perform feature enhancement processing on the global student features to obtain enhanced global student features.
[0076] Step D2 and Step D3 can be executed synchronously, and there is no sequential execution relationship.
[0077] Step D4: Adjust the model parameters of the second student model according to the difference degree between the enhanced global supervision features and the enhanced global student features.
[0078] Taking the feature enhancement processing of the global supervision features as an example. First, perform feature compression on the global supervision features in a specified dimension to obtain the global supervision compressed features in the specified dimension. Assume the global supervision features are 2D features , perform feature mean pooling in the H and W directions respectively to obtain the coordinate features in the H and W directions, which are used to capture the long - distance dependencies in the spatial direction and the position information of the space respectively, and perform feature concatenation to obtain the global supervision compressed features with the dimension of C×1×(W + H). The acquisition method of the global supervision compressed features can be expressed by the following formula (3): (3) where i and j respectively represent the pixel values of a certain pixel point on the feature image in the H and W directions.
[0079] Secondly, perform feature activation on the global supervision compressed features to obtain the activated global supervision compressed features. Based on the specified dimension, perform compression restoration on the activated global supervision compressed features to obtain the activated global supervision features.
[0080] The purpose of feature activation is to reduce the complexity of the network and improve the generalization ability of the model. An optional activation method is: First, use a fully - connected layer to compress the channel dimension of the global supervision features with a compression ratio of r to obtain the compressed features ; then perform batch normalization and feature activation on the features to reduce over - fitting of the model and introduce non - linearity, and then complete the restoration in the channel dimension at a ratio of 1 / r through a fully - connected layer to obtain the activated global supervision features . The process of the feature activation network can be expressed by the following formula:
[0081]
[0082] Among them, indicates that the global supervision feature is processed by the fully connected layer for processing. indicates that for the feature obtained after that is batch-normalized. indicates feature activation.
[0083] Again, the activated global supervision feature is used for feature enhancement processing to obtain the enhanced global supervision feature.
[0084] Optionally, the feature enhancement processing method is: weighting the activated global supervision feature and the original feature (i.e., the global supervision feature) to achieve feature enhancement in the channel coordinates. First, the activated global supervision feature is split in the H and W directions to obtain the feature weights and and the feature weights are normalized to enable them to weight the original feature, as shown in the following formula:
[0085]
[0086]
[0087] Among them, is the enhanced global supervision feature.
[0088] The process of feature enhancement processing for the global student feature is as follows: feature compression of the global student feature in the specified dimension is performed to obtain the globally compressed student feature in the specified dimension; feature activation of the globally compressed student feature is performed to obtain the activated globally compressed student feature; based on the specified dimension, decompression restoration of the activated globally compressed student feature is performed to obtain the activated global student feature; the activated global student feature is used for feature enhancement processing to obtain the enhanced global student feature. The specific process of feature enhancement processing for the global student feature is the same as that of the global supervision feature and will not be repeated here.
[0089] Finally, the enhanced global supervision feature and the global student feature output are enhanced in terms of channels and coordinates, and the loss function is solved by the loss function solving method provided in the foregoing embodiments. By performing feature enhancement processing on the global supervision feature and the global student feature, the problem of feature imbalance caused by stripping foreground features and background features during the knowledge distillation process can be eliminated, and the distillation effect of the model can be improved.
[0090] The following uses a specific embodiment to illustrate the model training method based on knowledge distillation provided by this application.
[0091] First, the implementation principle of the model training method based on knowledge distillation is described. Figure 7 It is a schematic principle diagram of a model training method based on knowledge distillation according to an embodiment of this application. As Figure 7 shown, a network model architecture is pre-built, including multiple teacher models and one student model. Due to space limitations, Figure 7 only one teacher model and one student model are shown. The teacher model includes a backbone network and a feature pyramid, and the student model includes a lightweight backbone network, a feature pyramid, and a network head. The sample images are respectively input into the teacher model and the student model, and foreground supervision sub-features and background supervision sub-features of N local category objects are extracted through the teacher model. The foreground supervision sub-features of N local category objects are concatenated to obtain foreground supervision features. The background supervision sub-features of N local category objects are concatenated to obtain background supervision features. Foreground features and background features of N local category objects are extracted through the student model. The feature levels of the features extracted by the teacher model and the student model are the same. Then, the foreground supervision sub-features, background supervision sub-features, foreground features, and background features are input into a multi-teacher feature distillation module for knowledge distillation, so that the student model can learn the recognition capabilities of multiple teacher models for N local category objects. The network head connected to the feature pyramid of the student model is used to output the prediction result of the student model for the sample image, and the prediction result includes the position prediction of N local category objects. Then, the loss function is calculated according to the prediction result and the label information of the sample image, and the model parameters of the student model are adjusted according to the loss function, so as to obtain the trained student model.
[0092] Based on Figure 7 the network model architecture shown, the following takes the training of a port monitoring model as an example for illustration.
[0093] First, sample images are obtained. The port includes scenarios of personnel management and dangerous operations, and the local category objects to be detected include personnel, vehicles, and operation equipment. The offline video in this scenario is collected, and the frame-extracted images of the offline video are obtained, and the effective images are screened out as sample images. The effective image refers to an image whose clarity meets the preset requirements and contains at least one local category object to be detected. In addition, public data samples can be obtained as a supplement to the sample images. The collected image set can be divided into a training set and a test set according to a certain ratio. Among them, the training set is used for knowledge distillation, and the test set is used for performance testing and optimization of the distilled second student model. After obtaining the sample images, the label information of the sample images needs to be labeled. The label information includes the position coordinates of the local category objects to be detected in the sample images.
[0094] Secondly, construct a pre-trained large model (i.e., the teacher model) and the student model. Since the local category objects to be detected include personnel, vehicles, and operation equipment, multiple teacher models need to be pre-configured so that the multiple teacher models have the ability to identify personnel, vehicles, and operation equipment. In this embodiment, 2 teacher models are configured, including a human detection model and a vehicle equipment detection model. The human detection model is used to identify personnel in the sample image, and the vehicle equipment detection model is used to identify vehicles and operation equipment in the sample image. Initialize the model parameters of the human detection model and the vehicle equipment detection model. In addition, according to the resource conditions of the port, construct a network model structure that meets the on-site real-time performance, including basic structures such as a backbone network, a feature pyramid, and a network head, and initialize the parameters of the network model structure, so as to generate a student model for distillation.
[0095] After that, perform knowledge distillation based on the constructed network model structure. Input the sample images containing label information in the training set into the human detection model, the vehicle equipment detection model, and the student model for feature extraction, and output multi-level feature images through the feature pyramids of each model. For the human detection model and the vehicle equipment detection model, according to the label information, obtain the position coordinates mapped by the label information from the multi-level feature images, so as to separate the foreground supervised sub-features. Then splice multiple foreground supervised sub-features to obtain the foreground supervised feature. Among them, after obtaining the foreground supervised feature, the image area obtained by removing the image area where all the foreground supervised features in the sample image are located is the background feature area. In this embodiment, the image area remaining after removing the areas where personnel, vehicles, and operation equipment are located in the sample image is the background image area, and the background supervised feature can be obtained by extracting the image features on the background image area.
[0096] During knowledge distillation, the distillation of the foreground supervised feature and the background supervised feature is carried out synchronously. For the foreground supervised feature, first perform feature alignment processing on the foreground supervised feature and the foreground feature extracted by the student model. The feature alignment processing methods may include: deleting some foreground supervised sub-features, or adding some foreground supervised sub-features. The specific feature alignment processing methods have been described in detail in the above embodiments and will not be elaborated here. After that, use the MSE loss function shown in formula (1) as the foreground distillation loss function, and perform loss solution and gradient backpropagation on the aligned teacher model and student model to achieve the foreground knowledge distillation of the multi-teacher model. For the background supervised feature, the MSE loss function shown in formula (2) can be used as the background distillation loss function, and perform loss solution and gradient backpropagation on the aligned teacher model and student model to achieve the background knowledge distillation of the multi-teacher model.
[0097] Before solving the loss, feature enhancement processing can be performed on the features respectively extracted by the teacher model and the student model. Optionally, obtain the full features of the human detection model, the vehicle device detection model, and the student model at a certain feature layer, and then perform feature enhancement on the obtained full features. The specific feature enhancement method has been described in detail in the above embodiments and will not be elaborated here. When solving the loss, according to the prediction result output by the network head of the student model and the label information of the sample image, loss solving and gradient backpropagation are performed.
[0098] It can be seen that by adopting the technical solution provided in the embodiment of the present application, when performing model migration based on knowledge distillation, foreground knowledge distillation and background knowledge distillation can be performed synchronously, so that the student model can not only learn the foreground knowledge in the sample image but also learn the low-heat background knowledge during the knowledge distillation process. Therefore, it helps to improve the generalization ability of the student model. In addition, since in the knowledge distillation process, the features of multiple local category objects in the sample image can be extracted, including the foreground supervision sub-features, background supervision sub-features, foreground features, and background features of personnel, vehicles, and operation equipment, and knowledge distillation is performed based on the features of multiple local category objects, the student model can learn multiple categories of knowledge at the same time. That is to say, a student model can learn cross-category cross knowledge, so that the distilled student model has the ability to identify different category objects, achieving the effect of distilling a student model with strong concurrency and high accuracy under the premise of a limited number of samples.
[0099] In summary, specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.
[0100] The above is the model training method based on knowledge distillation provided by the embodiment of the present application. Based on the same idea, the embodiment of the present application also provides a model training device based on knowledge distillation.
[0101] Figure 8 is a schematic block diagram of a model training device based on knowledge distillation according to an embodiment of the present application. As Figure 8 shown, the device includes: A first extraction module 81, configured to extract features from a first sample image through multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; The second extraction module 82 is configured to extract features from the first sample image through the first student model to obtain the first foreground features and the first background features of the N local category objects; The knowledge distillation module 83 is configured to perform foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground features, and perform background knowledge distillation on the first student model according to the background supervision sub-features and the first background features; the first student model after the foreground knowledge distillation and the background knowledge distillation is the second student model; The adjustment module 84 is configured to perform image prediction on the second sample image through the second student model, and adjust the model parameters of the second student model according to the prediction result of the image prediction.
[0102] In one embodiment, the apparatus further includes: The comparison module is configured to compare the model category of the pre-configured teacher model with the label information of the first sample image before extracting the foreground supervision sub-features and the background supervision sub-features of the N local category objects from the first sample image through multiple teacher models to obtain a first comparison result; each pre-configured teacher model of the model category is used to identify at least one local category object; the label information of the first sample image includes: the position features of the N local category objects in the first sample image; The screening module is configured to screen out the multiple teacher models that match the N local category objects from the pre-configured teacher models according to the first comparison result.
[0103] In one embodiment, when the first extraction module 81 extracts the foreground supervision sub-features and the background supervision sub-features of the N local category objects from the first sample image through multiple teacher models, the following steps are performed: In the case where the N local category objects include local category objects that match the model category of the teacher model, extract features from the matching local category objects through the teacher model to obtain the foreground supervision sub-features of the matching local category objects.
[0104] In one embodiment, when the first extraction module 81 extracts the foreground supervision sub-features and the background supervision sub-features of the N local category objects from the first sample image through multiple teacher models, the following steps are performed: Extract features from the local category objects in the first sample image that match the model category of the teacher model to obtain the foreground supervision sub-features of M local category objects; M is an integer greater than 1; When M is not equal to N, perform feature alignment processing on the M foreground supervised sub-features according to the label information of the first sample image to obtain the N foreground supervised sub-features; the label information of the first sample image includes: the position features of the N local category objects in the first sample image.
[0105] In one embodiment, when the first extraction module 81 performs feature alignment processing on the M foreground supervised sub-features according to the label information of the first sample image to obtain the N foreground supervised sub-features when M is not equal to N, the following steps are executed: When M is greater than N, determine the redundant sub-features among the M foreground supervised sub-features according to the M foreground supervised sub-features and the label information of the first sample image; delete the redundant sub-features from the M foreground supervised sub-features to obtain the N foreground supervised sub-features; When M is less than N, determine the missing sub-features among the N foreground supervised sub-features according to the M foreground supervised sub-features and the label information of the first sample image; perform feature extraction on the first sample image through a teacher model that matches the missing sub-features to obtain the missing sub-features.
[0106] In one embodiment, when the first extraction module 81 performs feature extraction on the first sample image through multiple teacher models to obtain the foreground supervised sub-features and background supervised sub-features of N local category objects, the following steps are executed: Perform feature extraction on the first sample image through the teacher model to obtain the foreground supervised sub-features of the local category objects that match the model category of the teacher model; Remove the first image area corresponding to the foreground supervised sub-feature and the second image area corresponding to the redundant foreground feature from the first sample image to obtain the background supervised sub-feature; the second image area is the image area corresponding to the local category object that the teacher model did not extract.
[0107] In one embodiment, when the knowledge distillation module 83 performs foreground knowledge distillation on the first student model according to the N foreground supervised sub-features and the first foreground feature, the following steps are executed: Concatenate the N foreground supervised sub-features to obtain the foreground supervised feature of the N local category objects; Calculate the distillation foreground loss function of the first student model according to the difference degree between the foreground supervised feature and the first foreground feature, and adjust the model parameters of the first student model according to the distillation foreground loss function.
[0108] In one embodiment, when the knowledge distillation module 83 performs foreground knowledge distillation on the first student model according to N foreground supervision sub-features and the first foreground feature, the following steps are executed: For each foreground supervision sub-feature, determine a first foreground sub-feature in the first foreground feature that matches the foreground supervision sub-feature; According to the degree of difference between the foreground supervision sub-feature and the first foreground sub-feature, calculate the local foreground loss function of the first student model; According to the N local foreground loss functions, determine the distillation foreground loss function of the first student model, and adjust the model parameters of the first student model according to the distillation foreground loss function.
[0109] In one embodiment, the prediction result of the image prediction includes: a foreground prediction result; When the adjustment module 84 performs image prediction on the second sample image through the second student model and adjusts the model parameters of the second student model according to the prediction result of the image prediction, the following steps are executed: Perform foreground prediction on the second sample image through the second student model to obtain the foreground prediction result including the N local category objects; Adjust the model parameters of the second student model according to the foreground prediction result and the label information of the second sample image.
[0110] In one embodiment, the foreground prediction result includes: prediction position information of the N local category objects in the second sample image; the label information of the second sample image includes: reference position information of the N local category objects in the second sample image; When the adjustment module 84 adjusts the model parameters of the second student model according to the foreground prediction result and the label information of the second sample image, the following steps are executed: Determine the supervision loss function of the second sample image according to the degree of difference between the prediction position information and the reference position information; Adjust the model parameters of the second student model according to the supervision loss function.
[0111] In one embodiment, the device further includes: A determination module, configured to determine the global supervision features of the multiple teacher models according to the N foreground supervision sub-features and the background supervision sub-feature; and determine the global student features of the second student model according to the first foreground feature and the first background feature; A feature enhancement module, configured to perform feature enhancement processing on the global supervised features to obtain enhanced global supervised features; and perform feature enhancement processing on the global student features to obtain enhanced global student features. A second adjustment module, configured to adjust the model parameters of the second student model according to the difference degree between the enhanced global supervised features and the enhanced global student features.
[0112] By using the device according to the embodiment of the present application, foreground supervised sub-features and background supervised sub-features of N local category objects are obtained by performing feature extraction on a first sample image through multiple teacher models, where N is an integer greater than 1; and a first foreground feature and a first background feature of N local category objects are obtained by performing feature extraction on the first sample image through a first student model. Furthermore, foreground knowledge distillation is performed on the first student model according to the N foreground supervised sub-features and the first foreground feature, and background knowledge distillation is performed on the first student model according to the background supervised sub-features and the first background feature to obtain a second student model, and the second student model is used to perform image prediction on a second sample image, and the model parameters of the second student model are adjusted according to the prediction result of the image prediction. It can be seen that when performing model transfer based on knowledge distillation, foreground knowledge distillation and background knowledge distillation can be performed synchronously, so that the student model can not only learn the foreground knowledge in the sample image but also learn the low-heat background knowledge during the knowledge distillation process, which helps to improve the generalization ability of the student model. In addition, since multiple local category object features in the sample image, including foreground supervised sub-features, background supervised sub-features, first foreground features, and first background features, can be extracted during the knowledge distillation process, and knowledge distillation is performed based on the features of multiple local category objects, the student model can learn multiple categories of knowledge at the same time. That is to say, a student model can learn cross-category cross knowledge, so that the distilled student model has the ability to identify different category objects, achieving the effect of distilling a student model with strong concurrency and high accuracy under the premise of a limited number of samples.
[0113] Those skilled in the art should understand that Figure 8 the model training device based on knowledge distillation in can be used to implement the model training method based on knowledge distillation described above, and the detailed description thereof should be similar to the method part described above. To avoid redundancy, it will not be elaborated here.
[0114] Based on the same idea, the embodiment of the present application also provides an electronic device, such as Figure 9As shown. Electronic devices can vary significantly due to different configurations or performances. They can include one or more processors 901 and a memory 902. One or more application programs or data can be stored in the memory 902. Among them, the memory 902 can be transient storage or persistent storage. The application programs stored in the memory 902 can include one or more modules (not shown in the figure). Each module can include a series of computer-executable instructions for the electronic device. Further, the processor 901 can be set to communicate with the memory 902 and execute a series of computer-executable instructions in the memory 902 on the electronic device. The electronic device can also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input / output interfaces 905, and one or more keyboards 906.
[0115] Specifically, in this embodiment, the electronic device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs can include one or more modules. Each module can include a series of computer-executable instructions for the electronic device and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions for: Extract features from the first sample image through multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; Extract features from the first sample image through the first student model to obtain the first foreground feature and the first background feature of the N local category objects; Perform foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground feature, and perform background knowledge distillation on the first student model according to the background supervision sub-features and the first background feature; the first student model after the foreground knowledge distillation and the background knowledge distillation is the second student model; Perform image prediction on the second sample image through the second student model, and adjust the model parameters of the second student model according to the prediction result of the image prediction.
[0116] Adopting the technical solution of the embodiment of the present application, multiple teacher models are used to extract features from the first sample image, obtaining foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; the first student model is used to extract features from the first sample image, obtaining the first foreground feature and the first background feature of the N local category objects. Furthermore, foreground knowledge distillation is performed on the first student model according to the N foreground supervision sub-features and the first foreground feature, and background knowledge distillation is performed on the first student model according to the background supervision sub-features and the first background feature, obtaining a second student model, and the second student model is used to perform image prediction on the second sample image, and the model parameters of the second student model are adjusted according to the prediction result of the image prediction. It can be seen that when performing model transfer based on knowledge distillation, foreground knowledge distillation and background knowledge distillation can be performed synchronously, enabling the student model to not only learn foreground knowledge in the sample image but also learn low-heat background knowledge during the knowledge distillation process, thus helping to improve the generalization ability of the student model. In addition, since during the knowledge distillation process, features of multiple local category objects in the sample image can be extracted, including foreground supervision sub-features, background supervision sub-features, the first foreground feature, and the first background feature, and knowledge distillation is performed based on the features of multiple local category objects, the student model can learn multiple categories of knowledge simultaneously. That is to say, a student model can learn cross-category cross knowledge, so that the distilled student model has the ability to recognize different category objects, achieving the effect of distilling a student model with strong concurrent ability and high accuracy under the premise of a limited number of samples.
[0117] The embodiment of the present application also proposes a computer-readable storage medium, which stores one or more computer programs. The one or more computer programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute each process of the above-mentioned model training method embodiment based on knowledge distillation, and are specifically used to execute: Using multiple teacher models to extract features from the first sample image, obtaining foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; Using the first student model to extract features from the first sample image, obtaining the first foreground feature and the first background feature of the N local category objects; Performing foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground feature, and performing background knowledge distillation on the first student model according to the background supervision sub-features and the first background feature; the first student model after the foreground knowledge distillation and the background knowledge distillation is the second student model; Performing image prediction on the second sample image through the second student model, and adjusting the model parameters of the second student model according to the prediction result of the image prediction.
[0118] Adopting the technical solution of the embodiment of the present application, extracting feature of the first sample image through multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; extracting feature of the first sample image through the first student model to obtain first foreground features and first background features of N local category objects. Furthermore, performing foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground features, and performing background knowledge distillation on the first student model according to the background supervision sub-features and the first background features to obtain the second student model, and performing image prediction on the second sample image through the second student model, and adjusting the model parameters of the second student model according to the prediction result of the image prediction. It can be seen that when performing model transfer based on knowledge distillation, foreground knowledge distillation and background knowledge distillation can be carried out synchronously, so that the student model can not only learn the foreground knowledge in the sample image but also learn the low-heat background knowledge during the knowledge distillation process, which helps to improve the generalization ability of the student model. In addition, since multiple local category objects' features in the sample image can be extracted during the knowledge distillation process, including foreground supervision sub-features, background supervision sub-features, first foreground features and first background features, and knowledge distillation is performed based on the features of multiple local category objects, the student model can learn multiple categories of knowledge at the same time. That is to say, a student model can learn cross-category cross knowledge, so that the distilled student model has the ability to recognize different category objects, achieving the effect of distilling a student model with strong concurrency and high accuracy under the premise of a limited number of samples.
[0119] The embodiment of the present application provides a computer program product, including a computer program, the computer program is executed by a processor to implement each process of the above-mentioned method embodiment of model training based on knowledge distillation, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0120] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0121] For the convenience of description, when describing the above device, various units are described separately according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.
[0122] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0123] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or blocks.
[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one or more of the flows Figure 1 or blocks.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or blocks.
[0126] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0127] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0128] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0129] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0130] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0131] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0132] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A model training method based on knowledge distillation, characterized in that: include: Performing feature extraction on the first sample image through multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, where N is an integer greater than 1; Performing feature extraction on the first sample image by using a first student model to obtain first foreground features and first background features of the N local category objects; According to the N foreground supervision sub-features and the first foreground feature, the first student model is subjected to foreground knowledge distillation, and according to the background supervision sub-features and the first background feature, the first student model is subjected to background knowledge distillation; the first student model after the foreground knowledge distillation and the background knowledge distillation is the second student model; The second sample image is predicted by the second student model, and the model parameters of the second student model are adjusted according to the prediction result of the image prediction.
2. The method according to claim 1, characterized in that Before extracting features from the first sample image using multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects, the method further includes: Comparing the model category corresponding to the preconfigured teacher model with the label information of the first sample image to obtain a first comparison result; the preconfigured teacher model of each model category is used to identify at least one local category object; the label information of the first sample image includes: position features of the N local category objects in the first sample image; According to the first comparison result, the multiple teacher models matching the N local category objects are screened out from the preconfigured teacher models.
3. The method according to claim 1, characterized in that The method of extracting features from the first sample image by using multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects includes: In the case where the N local category objects include a local category object that matches the model category of the teacher model, the teacher model is used to perform feature extraction on the matching local category object to obtain a foreground supervision sub-feature of the matching local category object.
4. The method according to claim 1, characterized in that The method of extracting features from the first sample image by using multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects includes: Extracting features of local category objects that match the model category of the teacher model in the first sample image to obtain foreground supervision sub-features of M local category objects, where M is an integer greater than 1; When M is not equal to N, feature alignment processing is performed on the M foreground supervision sub-features according to the label information of the first sample image to obtain the N foreground supervision sub-features; the label information of the first sample image includes: position features of the N local category objects in the first sample image.
5. The method according to claim 4, characterized in that In the case where M is not equal to N, performing feature alignment processing on the M foreground supervision sub-features according to the label information of the first sample image to obtain the N foreground supervision sub-features includes: When M is greater than N, determining redundant sub-features among the M foreground supervision sub-features according to the M foreground supervision sub-features and the label information of the first sample image; deleting the redundant sub-features from the M foreground supervision sub-features to obtain the N foreground supervision sub-features; When M is less than N, the missing sub-features among the N foreground supervision sub-features are determined based on the M foreground supervision sub-features and the label information of the first sample image; and the missing sub-features are obtained by performing feature extraction on the first sample image through a teacher model matching the missing sub-features.
6. The method according to claim 1, characterized in that The method of extracting features from the first sample image by using multiple teacher models to obtain foreground supervision sub-features and background supervision sub-features of N local category objects includes: Performing feature extraction on the first sample image by using the teacher model to obtain foreground supervision sub-features of local category objects that match the model category of the teacher model; The first image area corresponding to the foreground supervision sub-feature and the second image area corresponding to the redundant foreground feature are removed from the first sample image to obtain the background supervision sub-feature; the second image area is the image area corresponding to the local category object not extracted by the teacher model.
7. The method according to claim 1, characterized in that The step of performing foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground feature includes: The N foreground supervision sub-features are concatenated to obtain the foreground supervision features of the N local category objects; According to the difference between the foreground supervision feature and the first foreground feature, a distilled foreground loss function of the first student model is calculated, and model parameters of the first student model are adjusted according to the distilled foreground loss function.
8. The method according to claim 1, characterized in that: The step of performing foreground knowledge distillation on the first student model according to the N foreground supervision sub-features and the first foreground feature includes: For each foreground supervisory sub-feature, determining a first foreground sub-feature in the first foreground features that matches the foreground supervisory sub-feature; Calculating a local foreground loss function of the first student model according to the difference between the foreground supervision sub-feature and the first foreground sub-feature; According to the N local foreground loss functions, a distilled foreground loss function of the first student model is determined, and model parameters of the first student model are adjusted according to the distilled foreground loss function.
9. The method according to claim 1, characterized in that: The prediction results of the image prediction include: foreground prediction results; The performing image prediction on the second sample image by the second student model, and adjusting the model parameters of the second student model according to the prediction result of the image prediction, comprises: Performing foreground prediction on the second sample image by using the second student model to obtain the foreground prediction result including the N local category objects; According to the foreground prediction result and the label information of the second sample image, the model parameters of the second student model are adjusted.
10. The method according to claim 9, characterized in that The foreground prediction result includes: predicted position information of the N local category objects in the second sample image; the label information of the second sample image includes: reference position information of the N local category objects in the second sample image; The adjusting the model parameters of the second student model according to the foreground prediction result and the label information of the second sample image includes: Determining a supervision loss function of the second sample image according to a difference between the predicted position information and the reference position information; The model parameters of the second student model are adjusted according to the supervised loss function.
11. The method according to claim 7, characterized in that The method further comprises: Determine the global supervisory features of the multiple teacher models according to the N foreground supervisory sub-features and the background supervisory sub-features; determine the global student features of the second student model according to the first foreground feature and the first background feature; Performing feature enhancement processing on the global supervisory feature to obtain an enhanced global supervisory feature; Performing feature enhancement processing on the global student feature to obtain an enhanced global student feature; The model parameters of the second student model are adjusted according to the difference between the enhanced global supervisory features and the enhanced global student features.
12. An electronic device comprising a processor and a memory electrically connected to the processor, the memory storing a computer program, and the processor being used to call and execute the computer program from the memory to implement the model training method based on knowledge distillation as described in any one of claims 1-11.
13. A computer-readable storage medium, wherein the storage medium is used to store a computer program, wherein the computer program can be executed by a processor to implement the model training method based on knowledge distillation as described in any one of claims 1-11.
14. A computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the model training method based on knowledge distillation as described in any one of claims 1 to 11.
Citation Information
Cited By
Image feature extraction method and device, equipment and storage medium
CN121415083A
Knowledge distillation-based model training method and electronic device
WO2026149279A1