A Focused Distillation Method and System for 3D Object Detection

Through the focus distillation method for three-dimensional object detection, the target area is located using the query vector and mask generation distillation is performed, and the focus distillation is searched for representative feature positions in combination with the variability attention mechanism, which solves the problems of high computational complexity and noise interference of the existing three-dimensional object detection model, and achieves efficient and accurate three-dimensional object detection.

CN116343149BActive Publication Date: 2025-05-30SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310246290.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-14
Publication Date
2025-05-30
Estimated Expiration
2043-03-14

AI Technical Summary

Technical Problem

The existing three-dimensional object detection model based on bird's eye view has high computational complexity, its operating efficiency cannot meet the requirements of industrial deployment, and its characteristics are ray-shaped and contain a large amount of noise interference information.

Method used

A focal distillation method for three-dimensional object detection is proposed. By positioning the area where the three-dimensional object is located based on the query vector, the representative characteristic location of the three-dimensional object is searched for focal distillation to avoid noise areas.

Benefits of technology

Effectively screen out meaningful areas for distillation, avoid noise areas, improve the operating efficiency and accuracy of the model, and is suitable for three-dimensional object detection models based on bird's eye view and other types of three-dimensional object detection models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343149B_ABST
    Figure CN116343149B_ABST
Patent Text Reader

Abstract

This case involves a focus distillation method and system for three-dimensional object detection, belonging to the field of autonomous driving in computer vision, and is used to solve the problems in the prior art that the calculation complexity of the surround-view perception model is high and the operation efficiency cannot meet the requirements of industrial deployment, and to overcome the problem that the features of the three-dimensional object detection model based on the bird's-eye view are ray-shaped and contain a large amount of noise interference information. The purpose of this case is to achieve model compression by means of knowledge distillation technology. The specific implementation is target region mask generation and focus region search-based distillation. First, the coarse-grained target regions are located based on the query vector, and mask generation-based distillation is performed within these target regions. Then, using the offsets generated by the query vector, the representative feature positions of the target are searched, and focus distillation is performed at these representative feature positions, so as to achieve knowledge transfer between the teacher network and the student network, and achieve the effects of both high accuracy and high operation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving in computer vision, and particularly to a focus distillation method and system for three-dimensional object detection. Background Art

[0002] Perceiving the surrounding environment in three-dimensional space is the basis for downstream tasks of autonomous driving. The method based on camera images is widely used in autonomous driving vehicles due to its rich semantic information, low cost, and strong deployability. To achieve 360-degree omnidirectional perception of the surrounding environment, multiple cameras are usually deployed around the vehicle to detect surrounding vehicles, pedestrians, etc. The current state-of-the-art solution is to convert the image features of multiple cameras around the vehicle into a bird's-eye view perspective and then perform object detection from the bird's-eye view perspective, which is called bird's-eye view surround perception.

[0003] Although the current three-dimensional object detection models based on bird's-eye view can achieve high accuracy, the computational complexity of the models is high, and the running efficiency far fails to meet the requirements of industrial deployment. By compressing the depth of the model, the resolution of the input image, and the resolution of the bird's-eye view grid, the running speed of the model can be significantly improved, but its accuracy also has a non-negligible attenuation. The knowledge distillation technology can be used to guide and supervise the training process of a small model by relying on the two-dimensional image features and bird's-eye view features of a large model, so as to improve the accuracy of the small model.

[0004] The knowledge distillation technology usually sets a large model with high accuracy and low computational efficiency as the teacher network to supervise and guide the training process of another student network with relatively low accuracy but high computational efficiency. During the distillation process, the output or intermediate layer features of the teacher network are used as additional training targets for the student network to imitate and learn, so as to achieve the effect of transferring the knowledge of the teacher network to the student network. In the field of object detection, feature distillation is widely used due to its generality.

[0005] The difficulty in performing knowledge distillation on three-dimensional object detection based on bird's-eye view is that due to factors such as inaccurate depth estimation and occlusion in the teacher network model of pure vision three-dimensional object detection, the features for the student network to learn are not ideal enough. Specifically, the bird's-eye view features are in a ray shape, and false detections are likely to occur in the ray area. How to avoid these false detection areas and selectively perform distillation is a major difficulty that cannot be bypassed. Summary of the Invention

[0006] In view of the problems in the prior art that the calculation complexity of the surround view perception model is high and the operation efficiency cannot meet the requirements of industrial deployment, and the features of the 3D object detection model based on the bird's-eye view are ray-shaped and contain a large amount of noise interference information, this case designs a focus distillation method for 3D object detection. This method can effectively screen out meaningful regions for distillation and avoid noise regions. It is applicable not only to the 3D object detection model based on the bird's-eye view, but also to other types of 3D object detection models, with wide applicability and ease of use. The technical solution of this case is as follows.

[0007] In the first aspect, this case proposes a focus distillation method for 3D object detection, which can effectively distill meaningful target regions and avoid noise regions. The method not only focuses on the selection of the distillation region, but also considers the alignment method of features, including the following steps:

[0008] Locate the region where the 3D object is located based on the query vector, and perform mask generation distillation in the target region;

[0009] Through the deformable attention mechanism, use the position offset generated by the query vector to search for the representative feature positions of the 3D object, and perform focus distillation at the representative feature positions.

[0010] In the above technical solution, locating the region where the 3D object is located based on the query vector is specifically implemented through the following steps:

[0011] Use the aligned BEV (Bird’s Eye View) query vector to sample and generate 3D detection boxes from the BEV features;

[0012] Project the generated 3D detection boxes into the PV (Perspective View) and BEV views, and use the projected regions as the target regions.

[0013] In the above technical solution, the aligned BEV (Bird’s Eye View) query vector is obtained through the following steps:

[0014] Pass the query vector Q for distillation through the deformable self-attention layer and the feed-forward network pair (DeformAttn) and the feed-forward network Sample the teacher network BEV feature F E to generate y E , y E represents the detection boxes and their confidence levels output by the teacher network;

[0015] According to the ground truth of the detection box annotation and y E Calculate the detection loss based on the bipartite matching method, so that the query vector for distillation corresponds to the foreground target region;

[0016] Filter out the query vectors corresponding to the foreground target regions and remove the query vectors corresponding to the false positive recognition results. The remaining vectors are the aligned BEV query vectors, which are used as the focus query vectors.

[0017] In the above technical solution, mask generation-based distillation is performed in the target region, and the implementation process is as follows:

[0018] Generate a random mask I within the target region m ;

[0019] Multiply the student network feature F A by I m to obtain a masked student feature with some pixels masked;

[0020] Use a generator including two convolutional layers to restore the masked student feature to obtain the restored student feature

[0021] Calculate the per-channel distribution loss between the student feature E and the teacher network feature F Through optimize the student network feature, that is, complete the mask generation-based distillation in the target region.

[0022] In the above technical solution, The formula is as follows:

[0023]

[0024] Where: C is the total number of feature channels between the student feature and the teacher network feature F E ; c is the channel identifier, H and W are the height and width of the teacher network feature image respectively, i is the pixel identifier, τ is the temperature coefficient when calculating SoftMax, F is the channel feature, is the convolution function, The expression of

[0025]

[0026] is as follows: In the formula: x is the feature to be convolved.

[0027] In the above technical solution, if the dimensions of the student network feature and the teacher network feature are different, then switch the convolutional layer corresponding to the student network to a transposed convolutional layer.

[0028] In the above technical solution, through the variable attention mechanism, the representative feature positions of the three-dimensional target are searched using the position offsets generated by the query vectors, and the implementation process is as follows:

[0029] For each focus query vector q in the focus query vector set Q f based on the feature z of the focus query vector q q , a reference point p is generated through a linear transformation : q :

[0030]

[0031] The reference point p is used q to represent the projection center of the estimated bounding box on the sampled feature plane;

[0032] For each focus query vector q, using m to represent the index of the attention head and k to represent the sampling point, K sampling offsets Δp are generated through a linear transformation : mqk :

[0033]

[0034] In the above technical solution, focus distillation is performed at the representative feature positions, and the implementation process is as follows:

[0035] Denote the reference point corresponding to each focus query vector q in the focus query vector set Q f as p q , and denote the offset at the k-th sampling point under the attention head index m as Δp mqk , where m ∈ M, k ∈ K, M is the total number of attention head indices, and K is the total number of sampling points;

[0036] Obtain the sampled features of the teacher network at the representative feature positions and the sampled features of the student network as follows:

[0037]

[0038]

[0039] Normalize to and

[0040] respectively according to the maximum and minimum values, and calculate the L2 loss between the normalized representative features of the teacher network and the student network:

[0041]

[0042] Where: N f is the total number of focus query vectors.

[0043] In a second aspect, the present case proposes a focus distillation device for three-dimensional object detection, including a memory and a processor. A computer program capable of being loaded and executed by the processor for any of the above methods is stored on the memory.

[0044] In a third aspect, the present case proposes a focus distillation system for three-dimensional object detection. The system includes a target region mask generation module and a focus region search-based distillation module;

[0045] The target region mask generation module is configured to locate the region where the three-dimensional object is located based on the query vector and perform mask generation-based distillation in the target region;

[0046] The focus region search-based distillation module is configured to search for the representative feature position of the three-dimensional object through the variable attention mechanism using the position offset generated by the query vector and perform focus distillation at the representative feature position. Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 、 One Comparison of schematic diagrams of various methods for object detection knowledge distillation in a certain implementation;

[0049] Figure 2 、 One Schematic diagrams of the target region mask generation-based and focus region search-based distillation frameworks in a certain implementation;

[0050] Figure 3 、 One Schematic diagrams of the offset generation based on the query vector, the feature sampling of the teacher and student networks, and the query vector parameter update process in a certain implementation. Detailed Implementation Modes

[0051] Accurate 3D object detection is an important foundation for autonomous driving. Pure vision solutions based on surround view multi-cameras are widely used in autonomous driving vehicles because of their rich semantic information, low cost, and strong deployability. However, the current surround view perception model has high computational complexity and its operating efficiency is far from meeting the requirements of industrial deployment. To address this problem, this technical solution aims to achieve model compression with the help of knowledge distillation technology, and to achieve knowledge transfer between the teacher network that uses a large network and inputs high-resolution images and the student model that uses a small network and inputs low-resolution images, so as to achieve both high accuracy and high operating efficiency.

[0052] exist Figure 1 The following types of object detection knowledge distillation methods are illustrated in Figure 2.

[0053] Figure 1 (a) Global feature distillation, which performs imitation learning equally on the global region of the feature, but this method usually does not work well in object detection that focuses more on local features.

[0054] Figure 1 (b) Foreground region feature distillation: extracting the foreground region and only focusing on the foreground region or paying special attention to the foreground region. However, the disadvantage of this method is that the region selection is not precise, and there is a domain difference between the teacher network features and the student network features, so direct alignment has problems.

[0055] Figure 1 (c) Random mask generative distillation does not directly imitate the teacher network features, but randomly masks the pixels of the student network features, restores the masked features through a generator, and then aligns them with the teacher network features. This method eliminates the domain difference between the teacher network features and the student network features to a certain extent, but does not consider the different importance of different regions in detection.

[0056] Figure 1 (d) is the target area mask generation and focal area search distillation proposed in this case. The coarse-grained target area is located based on the query vector, and mask generation distillation is performed in these target areas. In addition, the offset generated by the query vector is used to search for the representative feature positions of the target, and focal distillation is performed at these representative feature positions. In this way, both the selection of the distillation area and the alignment of the features are taken into consideration. The proposed technical solution first roughly locates the target area through the query vector, performs mask generation distillation in these areas, and then uses the position offset generated by the query vector to search for the representative feature positions of the target, and performs refined focal distillation at these positions.

[0057] Figure 2Figure 1 shows a schematic diagram of the proposed target region mask generation and focus region search-based distillation framework. The framework aims to transfer the dark knowledge of the focus region from the teacher network to a compact student network, where feature distillation is performed on the perspective view (PV) and bird's-eye view (BEV). Aligned BEV query vectors are used. Through these query vectors, masks for the target region in PV and BEV are generated, and mask generation distillation is performed within the masked region to enhance the feature representation. In addition, these query vectors simultaneously utilize the deformable attention mechanism to search for representative features in the focus BEV region, which are used for focus distillation at these representative feature positions.

[0058] During the implementation of the technical solution of this case, the distillation framework has two important parts: target region mask generation distillation and focus region search distillation.

[0059] (I) Distillation Framework

[0060] Although the BEV perception 3D object detection network based on multi-camera input achieves high accuracy, it has the disadvantage of high computational overhead. This disadvantage mainly stems from the deep network layers and high input resolution. By compressing the network depth or input resolution, we obtain a lightweight network with faster inference speed. However, the lightweight network will inevitably be affected by accuracy degradation. Therefore, a large network model with high accuracy but low operating efficiency is used as the teacher network (Expert Network), and the lightweight model obtained by the above method is used as the student network (Apprentice Network). The parameters of the teacher network are frozen, and the intermediate features of the teacher network are used as auxiliary supervision for the student network.

[0061] As Figure 2 shown in Figure 2, a distillation framework between the teacher network and the student network is constructed. For the BEV perception model, distillation is performed on both PV features and BEV features simultaneously.

[0062] In the distillation framework, an additional distillation head is set. The distillation head is a network structure similar to the head of the deformable DETR network. The query vector Q for distillation is updated through the operation of the distillation head, that is, by interacting with the feature values of the objects to be detected. Based on the deformable attention mechanism, N query vectors sample key features from the BEV features and generate 3D detection boxes. The generated 3D detection boxes are projected onto the PV and BEV perspectives. The projected regions are used for mask generation distillation, and the BEV features sampled during the deformable attention process are used for fine distillation. To ensure that the projected regions are in the foreground region and the sampled BEV features come from the representative feature positions of the target region, the distillation head is also optimized through the auxiliary detection loss.

[0063] The distillation head is built on the BEV features of the teacher network and the student network. The query vector Q passes through the self-attention layer, and then through the deformable attention layer (DeformAttn) and the feed-forward network Sample the BEV features F of the teacher network E to generate y E , where y E represents the detection boxes output by the teacher network and their confidence levels. This process is expressed by the formula:

[0064] y E = FFN(DeformAttn(Q, F E ))

[0065] After obtaining y representing the output of the teacher network E , calculate the detection loss based on the ground truth of the detection box annotations and y E in a bipartite matching manner and optimize it, so that the query vectors for distillation are corresponding to the foreground target regions. Select the query vectors corresponding to the foreground target regions and remove the query vectors corresponding to the false positive recognition results. The remaining vectors are called focal query vectors, denoted as Q f .

[0066] (II) Target Region Mask Generation-based Distillation

[0067] Project the 3D boxes generated by the focal query vector Q f onto the PV view and the BEV view. The 3D boxes are projected onto the PV view according to the internal and external parameters of the camera and projected onto the BEV view from top to bottom.

[0068] The following described masking process is the same for both the PV view and the BEV view, so the PV view and the BEV view are not distinguished when describing this process.

[0069] Randomly set some pixels to 0 within the projection area of the 3D boxes generated by the focal query vector Q f to form a mask I m . Multiply the student network features F A by I m to obtain the masked student features with some pixels masked, and then pass it through a generator for restoration to obtain the restored features

[0070]

[0071] The generator here It includes two convolutional layers: If the sizes of the student network features and the teacher network features are the same, the two convolutional layers are the same and both play the role of restoring the mask features; if the dimensions of the student network features and the teacher network features are different, that is, the dimension of the student network features is smaller than that of the teacher network features, the convolutional layer corresponding to the student network is switched to a transposed convolutional layer. After restoring the mask features, the student network features are upsampled to obtain a feature image of the same size as the output of the teacher network.

[0072] The student features obtained by getting the target region mask and restoring through the generator After that, at Calculate the per-channel distribution loss between the student network features and the teacher network feature F E which is denoted as The calculation process is as follows:

[0073]

[0074] where: C is the total number of feature channels between the student features and the teacher network feature F E c is the channel identifier, H and W are the height and width of the teacher network feature image respectively, i is the pixel identifier, τ is the temperature coefficient when calculating SoftMax, F is the channel feature, is the convolution function, The expression of

[0075]

[0076] In the formula: x is the feature to be convolved.

[0077] Through Optimize the student network features, that is, complete the target region mask generation-based distillation.

[0078] (III) Focus region search-based distillation

[0079] The goal in this part is to automatically search for the representative positions of the object using a set of distillation-oriented query vectors and perform distillation on the teacher features and the student features at the same positions. We use the deformable attention module to achieve this goal. The deformable attention module generates a set of offsets around the reference point. Based on the offsets to the reference point, features at the key sampling points are sampled for object detection. We use this mechanism to perform refined distillation on the representative feature positions, as Figure 3 shown.

[0080] Figure 3It illustrates the process of offset generation based on query vectors, feature sampling of the teacher and student networks, and query vector parameter update. Using the deformable attention mechanism, local teacher network features and student network features are sampled at the same position according to the offsets generated from the query vectors. In addition, the query vector parameters are updated using the collected teacher network features.

[0081] For each focus query vector q in the set of focus query vectors Q f the reference point p q is generated from the feature z of the focus query vector q q through a linear transformation as follows:

[0082]

[0083] The reference point p q represents the projection center of the estimated bounding box on the sampling feature plane.

[0084] For each focus query vector q, M attention heads are attached to z q Each attention head generates K sampling offsets through another linear projection The offsets are calculated as follows:

[0085]

[0086] where: m represents the index of the attention head and k represents the sampling point.

[0087] Focus distillation based on query vectors is performed at these key representative feature positions. Let represent the sampling features of the teacher network and the student network respectively. They are calculated as follows:

[0088]

[0089]

[0090] The sampled teacher network features and student network features are normalized to and respectively according to the maximum and minimum values. Finally, the L2 loss is calculated between the normalized key sampling features of the teacher network and the student network:

[0091]

[0092] where: N f is the total number of focus query vectors.

[0093] By calculating the loss of fine-grained local regions, the focus distillation of these representative feature positions is completed. Query-based focus distillation aims to optimize the local patterns at the instance representative positions.

[0094] The above implementation is a distillation method for a pure-vision surround-view perception 3D detection model. It performs mask generation distillation based on query vectors to locate the target region, and at the same time searches for the representative feature region of the target according to the position offset generated by the query vectors for focus distillation. This method selectively distills meaningful regions and avoids noisy regions. The implementation of this case is convenient and easy to use, and can be used as a plug-and-play module for different types of 3D object detection models.

[0095] Through the description of the above implementation manners, those skilled in the art can clearly understand that the present disclosure can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, in most cases for the present disclosure, software program implementation is a better implementation manner.

[0096] In another implementation manner, according to the above content, it is implemented as a focus distillation system for 3D object detection. The system includes a target region mask generation module and a focus region search-based distillation module;

[0097] The target region mask generation module is configured to locate the region where the 3D target is located based on the query vector and perform mask generation distillation in the target region;

[0098] The focus region search-based distillation module is configured to search for the representative feature positions of the 3D target through a variability attention mechanism using the position offset generated by the query vector and perform focus distillation at the representative feature positions.

[0099] In summary, the technical solution of this case not only pays attention to the selection of the distillation region but also considers the feature alignment method. First, the region where the target is located is roughly located through the query vector, mask generation distillation is performed in these regions, and then the position offset generated by the query vector is used to search for the representative feature positions of the target, and refined focus distillation is performed at these positions, so as to effectively distill meaningful target regions and avoid noisy regions.

[0100] Although the embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and all of these fall within the scope of protection of the present invention.

Claims

1. A focus distillation method for 3D object detection, characterized in that, the method comprises the following steps: Locate the region where the 3D object is located based on the query vector, and perform mask generation-based distillation in the target region; Through the variability attention mechanism, use the position offset generated by the query vector to search for the representative feature positions of the 3D object, and perform focus distillation at the representative feature positions; Performing mask generation-based distillation in the target region, the implementation process is as follows: Generate a random mask within the target area ; Multiply the student network features by to obtain the masked student features with some pixels masked out; Use a generator including two convolutional layers to restore the masked student features and obtain the restored student features ; Among student features and teacher network features compute the per-channel distribution loss , and optimize the student network features through to complete the target region mask generative distillation; The formula is as follows: Wherein: is the student feature and the total number of feature channels between the teacher network features ; is the channel identifier , are the height and width of the teacher network feature image respectively is the pixel identifier is the temperature coefficient when calculating SoftMax is the channel feature is the convolution function The expression of is as follows: In the formula: is the feature to be convolved.

2. The method according to claim 1, characterized in that, Locating the region where the 3D object is located based on the query vector, the implementation steps are as follows: Use the aligned BEV query vector to sample and generate 3D detection boxes from the BEV features; Project the generated 3D detection boxes into the PV and BEV perspectives, and use the projected regions as the target regions.

3. The method according to claim 2, characterized in that, The aligned BEV query vector is obtained through the following steps: The query vector for distillation Sample the teacher network BEV features through the deformable self-attention layer and the feed-forward network to generate , representing the detection boxes and their confidence levels output by the teacher network; According to the true value marked by the detection box and calculate the detection loss based on the bipartite matching method, so that the query vector for distillation corresponds to the foreground target area; Select the query vectors corresponding to the foreground target regions, and remove the query vectors corresponding to the false positive recognition results. The remaining vectors are the aligned BEV query vectors, which are used as the focus query vectors.

4. The method according to claim 1, characterized in that, If the dimensions of the student network features and the teacher network features are different, then switch the convolutional layer corresponding to the student network to a transposed convolutional layer.

5. The method according to claim 3, characterized in that, Through the variability attention mechanism, use the position offset generated by the query vector to search for the representative feature positions of the 3D object, the implementation process is as follows: For each focus query vector in the focus query vector set , based on the characteristics of the intersection query vector , a reference point is generated through a linear transformation : ​​​ Using a reference point represents the projection center of the estimated bounding box on the sampled feature plane; For each focus query vector , use to represent the index of the attention head, to represent the sampling point, and generate through a linear transformation sampling offsets : 。 6. The method according to claim 1, characterized in that, Performing focus distillation at the representative feature positions, the implementation process is as follows: Set the focus query vector Each focus query vector in The corresponding reference point is denoted as , index the attention head No. The offset at each sampling point is recorded as , , , M is the total number of attention head indices, K is the total number of sampling points; Obtain the sampled features of the teacher network at the representative feature positions and the sampled features of the student network as follows: Will , be normalized according to the maximum value and the minimum value respectively into and ; Calculate the L2 loss between the normalized representative features of the teacher network and the student network; In the formula: is the total number of focus query vectors.

7. A focus distillation system for 3D object detection, characterized in that: It includes a memory and a processor, and a computer program capable of being loaded and executed by the processor, such as any one of the methods according to claims 1 to 6, is stored on the memory.

8. A focus distillation system for 3D object detection, characterized in that: The system includes a target region mask generation module and a focus region search-based distillation module; The target region mask generation module is configured to locate the region where the 3D object is located based on the query vector, and perform mask generation-based distillation in the target region; The focus region search-based distillation module is configured to use the variability attention mechanism to use the position offset generated by the query vector to search for the representative feature positions of the 3D object, and perform focus distillation at the representative feature positions; Wherein: Performing mask generation-based distillation in the target region, the implementation process is as follows: Generate a random mask within the target area ; Multiply the student network features by to obtain the masked student features with some pixels masked out; Use a generator including two convolutional layers to restore the masked student features and obtain the restored student features ; Among student features and teacher network features compute the per-channel distribution loss , and optimize the student network features through to complete the target region mask generative distillation; The formula is as follows: Wherein: is the student feature and the total number of feature channels between the teacher network features is the channel identifier is the channel identifier and are the height and width of the teacher network feature image respectively is the pixel identifier is the temperature coefficient when calculating SoftMax is the channel feature is the convolution function The expression of is as follows: In the formula: is the feature to be convolved.

Citation Information

Patent Citations

  • Image target detection method and detector based on knowledge distillation and training method thereof

    CN112164054A

  • Method for automatically compressing multitask-oriented pre-trained language model and platform thereof

    US20220188658A1