A sample-based selective knowledge distillation method and system
Through the sample-based selective knowledge distillation method, the stability and accuracy of RGB features are strengthened, and the problem of difficulty in taking into account accuracy and resource consumption in visual position recognition tasks is solved, achieving efficient, real-time and robust visual position recognition.
Patent Information
- Application Number
- CN202211441887.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-11-17
AI Technical Summary
The prior art is difficult to reduce time and resource consumption while improving accuracy in visual position recognition tasks, especially when processing RGB images, which cannot effectively deal with light and seasonal changes, and the re-ranking process consumes a large amount of computing resources.
The sample-based selective knowledge distillation method is used to extract RGB and SEG global features through the MobileNetV2 network, and the network is trained using a ternary loss function, combining the sample set division strategy and weight function to enhance the stability and accuracy of RGB features.
Without adding additional network inference, the accuracy of visual position recognition tasks is improved, time and resource consumption is reduced, and real-time and environmental robustness are taken into account.
Smart Images

Figure CN115761403B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and robotics, and specifically relates to a sample-based selective knowledge distillation method and system. Background Art
[0002] With the vigorous development of computer vision, deep learning-based retrieval and positioning have shown great development potential in the field of robotics. The current high-precision retrieval and positioning solution that can be processed in real time on robots is to use RGB images for two-stage retrieval: first, retrieval ranking is performed based on global features, and then re-ranking is performed based on local features in the selected top-N. Among them, the global retrieval method that only uses RGB images for image-level supervision does not work well. It is unable to extract sufficiently stable and robust features from RGB images to cope with changes in light, seasons, etc.; and the re-ranking process, while improving accuracy, often consumes a lot of time and computing resources.
[0003] In order to achieve better global retrieval results, some methods propose to use segmented images or depth images as auxiliary, together with RGB images as network input, and build a larger network structure to make up for the deficiency of only RGB images as input for global retrieval. However, in actual testing, this method needs to first generate the corresponding segmented images or depth images, and perform reasoning on a larger network, which still consumes a lot of time and resources and cannot guarantee the real-time requirements of visual position recognition tasks.
[0004] Therefore, a new training method is needed that can reduce time and resource consumption while enhancing the stability of RGB features. Summary of the invention
[0005] In order to solve the problem that time resource consumption and accuracy cannot be achieved simultaneously in visual location recognition tasks, the present invention provides a global retrieval method based on sample selective knowledge distillation, which is dedicated to improving the accuracy of visual location recognition tasks without adding additional network reasoning.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is: a sample-based selective knowledge distillation method, which includes two stages. The first stage includes the following steps:
[0007] The MobileNetV2 network is used as the RGB branch network in the first training stage. The multi-layer feature maps output by the MobileNetV2 network are merged and processed to obtain the RGB global feature.
[0008] The simplified MobileNetV2 network is used as the SEG branch network in the first training stage. The multi-layer feature maps output by the simplified MobileNetV2 network are merged and processed to obtain the SEG global feature.
[0009] The second phase includes the following steps:
[0010] For encoding the segmented image, the SEG image is converted into a tensor using one-hot encoding, and the one-hot encoding is converted into a weighted encoding, increasing the initial encoding value of static objects and reducing the initial encoding value of dynamic objects;
[0011] Use the triplet loss as the supervision of the network to train the SEG branch network and the RGB branch network respectively. Based on the trained RGB branch network and SEG branch network, calculate the global distance of each training set sample pair respectively.
[0012] A sample set partitioning strategy is designed to obtain a sample pair partitioning result, wherein a sample pair is composed of a query image and a positive sample; in the sample set partitioning strategy, a rectangular coordinate system is used for quantization, and the sample set is divided into different groups according to set conditions, wherein the horizontal axis x represents the ranking of p in the recall result of q when the sample pair (p, q) is tested with the trained SEG branch; and the vertical axis y represents the performance of the sample pair on the RGB branch;
[0013] Based on the sample pair division results, a weight function is set to give each sample pair a degree of importance in the distillation process;
[0014] The training is completed through the RGB branch network and the final loss function, and then the final test inference is performed through the RGB branch network. The final loss function is obtained by summing the ternary loss function and the distillation loss function of each sample pair; the MobileNetV2 network is used as the RGB branch network in the second training stage; the multi-layer feature maps output by the RGB branch network are merged and processed to obtain the final global features.
[0015] The SEG image is converted into a tensor, and the H×W array is converted into a C×H×W tensor using one-hot encoding. The C class labels are clustered into 6 categories, and the specific categories are: sky, ground, vegetation, dynamic objects, static objects and other objects.
[0016] The weight function is as follows:
[0017]
[0018] The weight function is proportional to (yx) and inversely proportional to x, N m is a constant hyperparameter.
[0019] The distillation loss function for each sample pair is:
[0020]
[0021] in represents the global features extracted by the first stage SEG network, G I represents the global features extracted by the second-stage RGB network, and T(·) is a fully connected layer that adjusts the dimensions of the two global features to be consistent.
[0022] The final loss function is:
[0023]
[0024] The sample set is divided into groups and defined as follows:
[0025]
[0026]
[0027]
[0028]
[0029] Where N is a constant hyperparameter.
[0030] The SEG branch network and the RGB branch network are trained separately using the ternary loss as the supervision of the network. Based on the trained RGB branch network and the SEG branch network, the global distance of each training set sample pair is calculated respectively. The distance between the global features adopts the Euclidean distance. The global distance between the image q and the image p is:
[0031] d(G q ,G p )=‖G q -G p ‖.
[0032] For the triple (q, p, n), the ternary loss function is used. The goal is to minimize the distance between the positive sample p and the query image q in the feature space, while making the negative sample n and the query image q as far away as possible in the feature space. The specific calculation formula is as follows:
[0033]
[0034] Where m is the distance threshold, which prevents the features of the samples from being aggregated into a very small space, G q ,G p ,G nThey are the features corresponding to the query image q, the positive sample image p and the negative sample image n respectively. When the distance between the query image q and the positive sample p is smaller than the distance between the query image q and the negative sample n, and smaller than the threshold m, there is no loss and no action is required; when the distance between the query image q and the positive sample p is larger than the distance between the query image q and the negative sample n, there is a non-zero loss.
[0035] Based on the above concept, the present invention provides a system for strengthening structural information based on sample-based selective knowledge distillation, including an RGB global feature acquisition module, a SEG global feature acquisition module, a segmented image encoding module, a sample pair global distance calculation module, a sample set partitioning module, a training module test reasoning module;
[0036] The RGB global feature acquisition module is used to use the MobileNetV2 network as the RGB branch network in the first training phase, merge the multi-layer feature maps output by the MobileNetV2 network, and obtain the RGB global feature.
[0037] The SEG global feature acquisition module is used to use the simplified MobileNetV2 network as the SEG branch network in the first training phase, merge the multi-layer feature maps output by the simplified MobileNetV2 network, and obtain the SEG global feature.
[0038] The segmented image encoding module uses one-hot encoding to convert the SEG image into a tensor, and the one-hot encoding is converted into a weighted encoding, which increases the initial encoding value of static objects and reduces the initial encoding value of dynamic objects;
[0039] The sample pair global distance calculation module uses the ternary loss as the supervision of the network to train the SEG branch network and the RGB branch network respectively. Based on the trained RGB branch network and SEG branch network, the global distance of each training set sample pair is calculated respectively;
[0040] The sample set partitioning module is used to design a sample set partitioning strategy and obtain a sample pair partitioning result, wherein a sample pair is composed of a query image and a positive sample; in the sample set partitioning strategy, a rectangular coordinate system is used for quantization, and the sample set is divided into different groups according to set conditions, wherein the horizontal axis x represents the ranking of p in the recall result of q when the sample pair (p, q) is tested with the trained SEG branch; the vertical axis y represents the performance of the sample pair on the RGB branch;
[0041] The training module test reasoning module is used to set a weight function to give each sample pair a degree of importance in the distillation process based on the sample pair division results; the training is completed through the RGB branch network and the final loss function, and then the final test reasoning is performed through the RGB branch network. The final loss function is obtained by summing the ternary loss function and the distillation loss function of each sample pair; the MobileNetV2 network is used as the RGB branch network in the second training stage; the multi-layer feature maps output by the RGB branch network are merged and processed to obtain the final global features.
[0042] In addition, the present invention also provides a computer device, including a processor and a memory, the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and when the processor executes the computer executable program, the sample-based selective knowledge distillation method can be implemented.
[0043] At the same time, a computer-readable storage medium can be provided, in which a computer program is stored. When the computer program is executed by a processor, the sample-based selective knowledge distillation method can be implemented.
[0044] Compared with the existing training schemes for visual position recognition, the present invention introduces additional modal data (segmented (SEG) images) in training, and avoids additional consumption during test reasoning while strengthening the stability of RGB features through sample-based selective knowledge distillation. The traditional scheme that only uses RGB images for backbone network training has low global retrieval accuracy. Although the accuracy is improved after re-ranking, the time and resource consumption after adding re-ranking will also be greatly increased. The scheme that introduces additional modal information and combines RGB to train a large network is of great help in improving the accuracy of global retrieval, but at the same time, the generation and reasoning of additional modal data cannot be avoided during testing, and still does not meet the real-time requirements of robot visual position recognition tasks. The present invention strengthens the invariant features in the additional modal information SEG in the RGB features through knowledge distillation without retaining additional branch networks. The segmented image is input into the training network through weighted one-hot encoding to extract the scene structure information for the visual position recognition task. Through sample-based selective knowledge distillation, high-quality knowledge is directionally strengthened to greatly improve the network accuracy. Finally, the network only uses RGB images for global retrieval during testing, taking into account real-time, high-precision and environmental robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 Schematic diagram of global feature extraction in the present invention.
[0046] Figure 2 Schematic diagram of sample-based selective knowledge distillation in the present invention.
[0047] Figure 3 Schematic diagram of the sample partitioning strategy in the present invention. DETAILED DESCRIPTION
[0048] The specific implementation modes of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0049] This paper designs a sample-based selective knowledge distillation method for visual place recognition (VPR) tasks. First, the network for segmented (SEG) images and RGB images is trained to complete the VPR task respectively, and then the high-quality knowledge contained in the SEG features is transferred to the RGB features, thereby improving the positioning accuracy without additional time resource consumption. The training network is based on triples, consisting of 1 query, 1 positive sample and J negative samples, denoted as (q, p, n j ). The specific feature extraction is implemented through steps A1 to A3, and the distillation method is implemented through steps B1 to B7.
[0050] Step A1: Use the MobileNetV2 network as the RGB branch network in the first training phase. Merge the multi-layer feature maps output by the MobileNetV2 network to obtain the RGB global feature; the i-th layer feature map is denoted as The RGB features of image I are denoted as The present invention selects the last three layers of feature maps To merge, the merging process is as follows Figure 1 shown.
[0051] Step A2, use the simplified MobileNetV2 network shown in Table 1 as the SEG branch network in the first training phase. The bottleneck operation in the table is the aggregation of a series of basic network operations. t is the expansion factor. The number of network channels is expanded by t times during the processing and then scaled to the required number of output channels. c is the number of output channels. n represents the number of repetitions of the layer. s is the convolution stride. Figure 1 As shown, the multi-layer feature map output by the simplified MobileNetV2 network The SEG global feature is obtained by merging. The SEG feature of image I is recorded as
[0052] Step A3, use the MobileNetV2 network as the RGB branch network in the second training phase, and also the network used in the final test reasoning; Figure 1 As shown in the figure, the multi-layer feature maps output by the network are merged and processed to obtain the final global features; the final RGB features of image I are recorded as
[0053] Table 1 Simplified MobileNetV2 network structure
[0054]
[0055] Step B1, encoding of segmented images.
[0056] Using the open source segmentation model, we can get the SEG image corresponding to the RGB image. Assuming that the number of segmentation categories of the model is C, the size of the RGB image is H×W, the SEG image can be represented by an H×W array, where the value of each position represents the label of the pixel, and the value range is [0, C-1].
[0057] The SEG image is used as the network input. Before completing the VPR task, the SEG image needs to be converted into a tensor. The traditional method uses one-hot encoding to convert the H×W array into a C×H×W tensor. Specifically, if the category label of a pixel is j, the C-dimensional vector corresponding to the position is 1 only in the jth channel, and the values of other channels are 0.
[0058] Furthermore, the present invention considers that when the value of the original label number C is too large, the SEG image encoding and the corresponding network will be very large, and the overly fine-grained segmentation has no additional advantages for the VPR task, but is prone to introduce noise. The present invention clusters the C-type labels into 6 categories, and the specific categories are: sky, ground, vegetation, dynamic objects, static objects and other objects. At this time, the SEG image encoding size is 6×H×W.
[0059] Finally, the present invention combines the human attention mechanism, that is, humans will automatically focus on static objects and other landmarks and ignore dynamic objects when performing scene recognition, and converts the one-hot encoding form into weighted encoding, that is, increasing the initial encoding value of static objects and reducing the initial encoding value of dynamic objects, so as to better guide the model convergence.
[0060] Step B2, first training stage.
[0061] In the first stage of training, the SEG branch network and the RGB branch network are trained separately, using the ternary loss as the supervision of the network. The distance between global features adopts the Euclidean distance, and the global distance between image q and image p is:
[0062] d(G q ,G p )=‖G q -G p ‖. (1)
[0063] For the triplet (q, p, n), the VPR task often uses a ternary loss function, the goal is to minimize the distance between the positive sample p and the query image q in the feature space, while making the negative sample n and the query image q as far away as possible in the feature space. The specific calculation formula is as follows:
[0064]
[0065] Where m is the distance threshold to avoid the sample features from being aggregated into a very small space. Generally, 0.1 is used. q ,G p ,G n They are the features corresponding to the query image q, the positive sample image p and the negative sample image n respectively. It can be seen that when the distance between the query image q and the positive sample p is smaller than the distance between the query image q and the negative sample n (and smaller than the threshold m), there is no loss and no action is required; and when the distance between the query image q and the positive sample p is larger than the distance between the query image q and the negative sample n, there is a non-zero loss.
[0066] Step B3: training set sample division method.
[0067] According to the result of step B1, in the VPR task, the SEG network is not completely superior to the RGB network. This is different from the setting of "teacher network is stronger than student network" in traditional knowledge distillation. The present invention screens and weights samples before knowledge distillation; considering the concept of triples in the VPR task, the division object is not a single image sample, but a sample pair consisting of 1 query image and 1 positive sample, i.e. (p, q).
[0068] Different from the general method of using only a single network for sample division, the present invention calculates the recall rate of each training set sample pair based on the RGB branch network and the SEG branch network trained in the first stage, that is, for a given sample pair (p, q), query the ranking of the positive sample p in the recall result of q. Obviously, the higher the ranking of the positive sample p in the recall result, the smaller the recall rate, indicating better performance. In order to make the performance of the sample pair on the two branch networks more intuitive, the present invention uses a rectangular coordinate system to quantify it, and divides the sample set into different groups according to the set conditions, such as Figure 3 As shown in the figure. The horizontal axis x represents the ranking of p in the recall result of q when the sample pair (p, q) is tested with the trained SEG branch; the vertical axis y represents the performance of the sample pair on the RGB branch. The divided groups reflect the performance of different sample pairs on the two branches. For example, sample pairs that perform better on the SEG branch network than the RGB network are more beneficial and important for the distillation from SEG to RGB, while sample pairs that perform poorly on the SEG branch network will damage the effect of distillation because these sample pairs do not contain high-quality effective knowledge. Figure 3 The group partitions shown are defined as follows:
[0069]
[0070] Where N is a constant hyperparameter, usually 5 or 10. The smaller x is, the stronger the SEG branch's ability to distinguish sample pairs is, and the better the performance is. For the sample set, the processing capability of the SEG branch is poor.
[0071] Step B4. Weighted Distillation Loss Function
[0072] Based on the sample pair division results of step B3, different distillation weights can be applied to different groups. In order to distill more accurately, a more reasonable way is to assign each sample pair a degree of importance in the distillation process. For this purpose, the present invention further designs the following weight function:
[0073]
[0074] The weight function is proportional to (yx) and inversely proportional to x, that is, the worse the SEG branch network performs, the lower the sample pair weight; the better the SEG branch network performs than the RGB branch network, the higher the sample pair weight. m It is a constant hyperparameter, usually 2N.
[0075] Step B5. Final loss function calculation.
[0076] The distillation loss for each sample pair is as follows:
[0077]
[0078] in represents the global features extracted by the SEG branch network in the first stage, G I represents the global features extracted by the second-stage RGB network, and T(·) is a fully connected layer that adjusts the dimensions of the two global features to be consistent.
[0079] The two parts of the loss function calculated in formula (2) and formula (5) are summed to further constrain the optimization space and achieve joint optimization:
[0080]
[0081] Formula (6) realizes the aggregation of weighted distillation loss and original VPR loss, as Figure 2As shown in the figure, in the second training stage, the training is completed through the RGB branch network and the final loss function, and then the final test reasoning is performed through the RGB branch network. It should be noted that although the same RGB branch network as the first training stage is used here, the network parameters are different from those of the first training stage.
[0082] First, the network that segmented the SEG image and RGB image to complete the VPR task is trained, and then the high-quality knowledge contained in the SEG feature is transferred to the RGB feature. The invariant features in the additional modal information are strengthened in the RGB feature by knowledge distillation without retaining additional branch networks. The present invention inputs the segmented image into the training network through weighted unique hot encoding to extract scene structure information for the visual position recognition task; through sample-based selective knowledge distillation, high-quality knowledge is directionally strengthened to further improve the network accuracy; the final network only uses RGB images for global retrieval during testing, taking into account real-time, high-precision and environmental robustness.
[0083] Based on the technical concept of the present invention, a sample-based selective knowledge distillation method and system are provided, including an RGB global feature acquisition module, a SEG global feature acquisition module, a segmented image encoding module, a sample pair global distance calculation module, a sample set partitioning module, a training module test reasoning module;
[0084] The RGB global feature acquisition module is used to use the MobileNetV2 network as the RGB branch network in the first training phase, merge the multi-layer feature maps output by the MobileNetV2 network, and obtain the RGB global feature.
[0085] The SEG global feature acquisition module is used to use the simplified MobileNetV2 network as the SEG branch network in the first training phase, merge the multi-layer feature maps output by the simplified MobileNetV2 network, and obtain the SEG global feature.
[0086] The segmented image encoding module uses one-hot encoding to convert the SEG image into a tensor, and the one-hot encoding is converted into a weighted encoding, which increases the initial encoding value of static objects and reduces the initial encoding value of dynamic objects;
[0087] The sample pair global distance calculation module uses the ternary loss as the supervision of the network to train the SEG branch network and the RGB branch network respectively. Based on the trained RGB branch network and SEG branch network, the global distance of each training set sample pair is calculated respectively;
[0088] The sample set partitioning module is used to design a sample set partitioning strategy and obtain a sample pair partitioning result, wherein a sample pair is composed of a query image and a positive sample; in the sample set partitioning strategy, a rectangular coordinate system is used for quantization, and the sample set is divided into different groups according to set conditions, wherein the horizontal axis x represents the ranking of p in the recall result of q when the sample pair (p, q) is tested with the trained SEG branch; the vertical axis y represents the performance of the sample pair on the RGB branch;
[0089] The training module test reasoning module is used to set a weight function to give each sample pair a degree of importance in the distillation process based on the sample pair division results; the training is completed through the RGB branch network and the final loss function, and then the final test reasoning is performed through the RGB branch network. The final loss function is obtained by summing the ternary loss function and the distillation loss function of each sample pair; the MobileNetV2 network is used as the RGB branch network in the second training stage; the multi-layer feature maps output by the RGB branch network are merged and processed to obtain the final global features.
[0090] In addition, the present invention can also provide a computer device, including a processor and a memory, the memory is used to store a computer executable program, the processor reads part or all of the computer executable program from the memory and executes it, and when the processor executes part or all of the computer executable program, it can implement the sample-based selective knowledge distillation method described in the present invention.
[0091] On the other hand, the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the sample-based selective knowledge distillation method described in the present invention can be implemented.
[0092] The computer device may be a laptop computer, a desktop computer or a workstation.
[0093] The processor can be a central processing unit (CPU), a graphics processing unit (GPU) / digital signal processor (DSP), an application specific integrated circuit (ASIC), or an off-the-shelf field programmable gate array (FPGA).
[0094] The memory of the present invention may be an internal storage unit of a laptop computer, a desktop computer or a workstation, such as a memory or a hard disk; or an external storage unit, such as a mobile hard disk or a flash memory card.
[0095] Computer-readable storage media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media may include: read-only memory (ROM), random access memory (RAM), solid-state drive (SSD) or optical disk, etc. Among them, random access memory may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM).
[0096] The above is a detailed introduction to a visual position recognition system and method based on selective knowledge distillation of samples disclosed in an embodiment of the present invention. The description of the above embodiment is only used to help understand the method of the present invention and its core idea, but the present invention is not limited to the above implementation. For those of ordinary skill in the art, within the scope of their knowledge, according to the idea of the present invention, changes can be made in the specific implementation or scope of application. In summary, the content of this specification should not be understood as a limitation of the present invention.
Claims
1. A sample-based selective knowledge distillation method, characterized in that: It consists of two stages. The first stage includes the following steps: The MobileNetV2 network is used as the RGB branch network in the first training stage. The multi-layer feature maps output by the MobileNetV2 network are merged and processed to obtain the RGB global feature. The simplified MobileNetV2 network is used as the SEG branch network in the first training stage. The multi-layer feature maps output by the simplified MobileNetV2 network are merged and processed to obtain the SEG global feature. The second phase includes the following steps: For encoding the segmented image, the SEG image is converted into a tensor using one-hot encoding, and the one-hot encoding is converted into a weighted encoding, increasing the initial encoding value of static objects and reducing the initial encoding value of dynamic objects; Use the triplet loss as the supervision of the network to train the SEG branch network and the RGB branch network respectively. Based on the trained RGB branch network and SEG branch network, calculate the global distance of each training set sample pair respectively. A sample set partitioning strategy is designed to obtain a sample pair partitioning result, wherein a sample pair is composed of a query image and a positive sample; in the sample set partitioning strategy, a rectangular coordinate system is used for quantization, and the sample set is divided into different groups according to set conditions, wherein the horizontal axis x represents the ranking of p in the recall result of q when the sample pair (p, q) is tested with the trained SEG branch; and the vertical axis y represents the performance of the sample pair on the RGB branch; Based on the sample pair division results, a weight function is set to give each sample pair a degree of importance in the distillation process; The training is completed through the RGB branch network and the final loss function, and then the final test reasoning is performed through the RGB branch network. The final loss function is obtained by summing the ternary loss function and the distillation loss function of each sample pair; the MobileNetV2 network is used as the RGB branch network in the second training stage; the multi-layer feature maps output by the RGB branch network are merged and processed to obtain the final global features; The weight function is as follows: The weight function is proportional to (yx) and inversely proportional to x, N m is a constant hyperparameter; The distillation loss function for each sample pair is: in represents the global features extracted by the SEG network in the first stage, represents the global features extracted by the second-stage RGB network, T(·) is a fully connected layer that adjusts the dimensions of the two global features to be consistent; The sample set is divided into groups as follows: Where N is a constant hyperparameter.
2. The sample-based selective knowledge distillation method according to claim 1, characterized in that: The SEG image is converted into a tensor, and the H×W array is converted into a C×H×W tensor using one-hot encoding. The C class labels are clustered into 6 categories, and the specific categories are: sky, ground, vegetation, dynamic objects, static objects and other objects.
3. The sample-based selective knowledge distillation method according to claim 1, characterized in that: The final loss function is: G q ,G p ,G n are the features corresponding to the query image q, the positive sample image p and the negative sample image n respectively. represents the ternary loss function.
4. The sample-based selective knowledge distillation method according to claim 1, characterized in that: The SEG branch network and the RGB branch network are trained separately using the ternary loss as the supervision of the network. Based on the trained RGB branch network and the SEG branch network, the global distance of each training set sample pair is calculated respectively. The distance between the global features adopts the Euclidean distance. The global distance between the image q and the image p is: d(G q ,G p )=‖G q -G p ‖. For the triple (q, p, n), the ternary loss function is used. The goal is to minimize the distance between the positive sample p and the query image q in the feature space, while making the negative sample n and the query image q as far away as possible in the feature space. The specific calculation formula is as follows: Where m is the distance threshold, which prevents the features of the samples from being aggregated into a very small space, G q ,G p ,G n They are the features corresponding to the query image q, the positive sample image p and the negative sample image n respectively. When the distance between the query image q and the positive sample p is smaller than the distance between the query image q and the negative sample n, and smaller than the threshold m, there is no loss and no action is required; when the distance between the query image q and the positive sample p is larger than the distance between the query image q and the negative sample n, there is a non-zero loss.
5. A system for selective knowledge distillation based on samples, characterized in that: It includes RGB global feature acquisition module, SEG global feature acquisition module, segmentation image encoding module, sample pair global distance calculation module, sample set division module, training module test inference module; The RGB global feature acquisition module is used to use the MobileNetV2 network as the RGB branch network in the first training phase, merge the multi-layer feature maps output by the MobileNetV2 network, and obtain the RGB global feature. The SEG global feature acquisition module is used to use the simplified MobileNetV2 network as the SEG branch network in the first training phase, merge the multi-layer feature maps output by the simplified MobileNetV2 network, and obtain the SEG global feature. The segmented image encoding module uses one-hot encoding to convert the SEG image into a tensor, and the one-hot encoding is converted into a weighted encoding, which increases the initial encoding value of static objects and reduces the initial encoding value of dynamic objects; The sample pair global distance calculation module uses the ternary loss as the supervision of the network to train the SEG branch network and the RGB branch network respectively. Based on the trained RGB branch network and SEG branch network, the global distance of each training set sample pair is calculated respectively; The sample set partitioning module is used to design a sample set partitioning strategy and obtain a sample pair partitioning result, wherein a sample pair is composed of a query image and a positive sample; in the sample set partitioning strategy, a rectangular coordinate system is used for quantization, and the sample set is divided into different groups according to set conditions, wherein the horizontal axis x represents the ranking of p in the recall result of q when the sample pair (p, q) is tested with the trained SEG branch; the vertical axis y represents the performance of the sample pair on the RGB branch; The training module test reasoning module is used to set a weight function to give each sample pair its importance in the distillation process based on the sample pair division results; the training is completed through the RGB branch network and the final loss function, and then the final test reasoning is performed through the RGB branch network. The final loss function is obtained by summing the ternary loss function and the distillation loss function of each sample pair; the MobileNetV2 network is used as the RGB branch network in the second training stage; the multi-layer feature maps output by the RGB branch network are merged and processed to obtain the final global features; the weight function is as follows: The weight function is proportional to (yx) and inversely proportional to x, N m is a constant hyperparameter; The distillation loss function for each sample pair is: in represents the global features extracted by the SEG network in the first stage, represents the global features extracted by the second-stage RGB network, T(·) is a fully connected layer that adjusts the dimensions of the two global features to be consistent; The sample set is divided into groups as follows: Where N is a constant hyperparameter.
6. A computer device, characterized in that: It includes a processor and a memory, the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and when the processor executes the computer executable program, it can implement the sample-based selective knowledge distillation method described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: A computer program is stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the sample-based selective knowledge distillation method as described in any one of claims 1 to 4.