Method and apparatus for learning efficient detection transducer
By applying knowledge distillation techniques, particularly query-guided knowledge distillation strategies, to the detection transformer, the feature distillation of the transformer encoder is optimized, solving the problem of high computational overhead in the DETR model and achieving efficient object detection.
Patent Information
- Application Number
- CN202380096935.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-11-04
AI Technical Summary
Existing Detection Transformer (DETR) models suffer from high computational overhead due to their heavy trunks and transformer heads, making them difficult to operate efficiently in practical applications.
The knowledge distillation technique is used to transfer learned knowledge from a large teacher model to a smaller student model. Through the query-guided knowledge distillation (QGD) strategy, the focus is on feature distillation on the transformer encoder. The student model's performance is optimized by weighting the average attention map and the attention map of the object query.
It significantly improves the efficiency and performance of the detection transformer, reduces computational overhead, and maintains detection accuracy, making it suitable for real-world object detection tasks.
Smart Images

Figure CN120898231A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Aspects of the disclosure generally relate to artificial intelligence, and more particularly to methods and apparatus provided for learning efficient detection transformers via query-guided knowledge distillation. BACKGROUND
[0002] The goal of object detection is to predict a set of bounding boxes and class labels for each object of interest. Detection transformers (DETR) and its variants have become the new paradigm for object detection due to their direct set prediction capability. However, existing DETR-like models (DETRs) suffer from heavy backbone and transformer heads, which limit their practical applications. To address this issue, possibilities to improve their efficiency are being explored.
[0003] Knowledge distillation is a well-known popular practice to learn efficient models, which can significantly improve efficiency while maintaining performance by transferring learned knowledge from a large teacher model to a smaller student model. Therefore, there is a motivation to apply knowledge distillation to DETRs. SUMMARY
[0004] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements of all aspects nor delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.
[0005] Disclosed herein is how to effectively apply knowledge distillation to DETRs. To achieve this goal, at least two main aspects need to be considered: 1) which components of a DETR model should be reduced; and 2) how to utilize knowledge distillation to improve the performance of the reduced DETR model, in other words, where to distill and how to effectively distill.
[0006] In an aspect, a computer-implemented method for machine learning is disclosed. The method comprises: training a first detection transformer (DETR) model; training a second DETR model based on a loss that weights feature differences between layers of an encoder of the trained first DETR model and an encoder of the second DETR model with an average attention map, wherein the encoder of the second DETR model has fewer layers than the encoder of the first DETR model; and wherein the average attention map is based on one or more of: an average of attention maps of multi-heads of a decoder of the trained first DETR model and one or more object queries, or an average of attention maps of multi-heads of a decoder of the second DETR model and one or more object queries.
[0007] In another aspect, the average attention map is based on a ratio of an average of the attention maps of the multi-head of the decoder of the trained first DETR model and the one or more object queries to an average of the attention maps of the multi-head of the decoder of the second DETR model and the one or more object queries.
[0008] In another aspect, the attention map of the trained first DETR model and the attention map of the second DETTR model are attention maps from a last layer of the respective decoder.
[0009] In another aspect, the one or more object queries of the decoder of the trained first DETR model and the one or more object queries of the decoder of the second DETR model are deterministic object queries.
[0010] In another aspect, the feature difference is computed between the output features of each layer of the encoder of the second DETR model and the output features of a final layer of the encoder of the trained first DETR model.
[0011] In another aspect, the feature difference is computed between the output features of each layer of the encoder of the second DETR model and the output features of each corresponding intermediate layer of the encoder of the trained first DETR model.
[0012] In another aspect, training the second DETR model is further based on a loss that evaluates a difference between an output of the decoder of the second DETR model and a ground truth label.
[0013] In another aspect, the first DETR model and the second DETR model are trained with image data, and the outputs of the first DETR model and the second DETR model are class labels and bounding boxes corresponding to the one or more object queries.
[0014] In an aspect, a computer-implemented method for object detection using a DETR model trained with one of the methods disclosed herein is disclosed. The method includes obtaining image data captured during movement of a vehicle; inputting the image data into the DETR model; outputting object class labels and bounding boxes corresponding to one or more object queries of the DETR model; and assisting travel of the vehicle based on the outputted object class labels and bounding boxes.
[0015] In an aspect, a computer system is disclosed. The computer system includes one or more processors; and one or more storage devices storing computer-executable instructions that, when executed, cause the one or more processors to perform operations of one of the methods disclosed herein.
[0016] In an aspect, one or more computer-readable storage media storing computer-executable instructions that, when executed, cause one or more processors to perform operations of one of the methods disclosed herein are disclosed.
[0017] In an aspect, a computer program product comprising computer-executable instructions that, when executed, cause one or more processors to perform operations of one of the methods disclosed herein is disclosed.
[0018] In an aspect, a vehicle comprising one or more apparatuses for performing one of the methods disclosed herein is disclosed. BRIEF DESCRIPTION OF DRAWINGS
[0019] The disclosed aspects will be described with respect to the following drawings, which are provided for illustration and not limitation.
[0020] Figure 1 An example structure of a detector transformer is shown in accordance with aspects of the disclosure.
[0021] Figure 2 An example structure of knowledge distillation is shown in accordance with aspects of the disclosure.
[0022] Figure 3 An example of inference time of DETR is shown in accordance with aspects of the disclosure.
[0023] Figure 4 An example distillation strategy is shown in accordance with aspects of the disclosure.
[0024] Figure 5 An example structure of query-guided knowledge distillation is shown in accordance with aspects of the disclosure.
[0025] Figure 6 Two example distillation options are shown in accordance with aspects of the disclosure.
[0026] Figure 7 An example flow diagram for learning efficient DETR is shown in accordance with aspects of the disclosure.
[0027] Figure 8 An example flow diagram for using a trained DETR is shown in accordance with aspects of the disclosure.
[0028] Figure 9 An example computer system is shown in accordance with aspects of the disclosure. DETAILED DESCRIPTION
[0029] The present disclosure will now be discussed with reference to a number of example implementations. It should be appreciated that discussing these implementations is for illustrative purposes only and is not intended to limit the scope of this disclosure in any way. Further, it should be appreciated that the implementations can each be implemented in numerous ways, including as a process, article of manufacture, or a system in the form of a machine.
[0030] Various implementations will now be described in detail with reference to the drawings. Wherever possible, the same or like reference numerals are used in the drawings and the following description. References to example and implementations are for illustrative purposes only and are not intended to limit the scope of this disclosure in any way.
[0031] Object detection is an important computer vision task that involves detecting instances of a particular class of visual objects (such as people, animals, vehicles, or traffic signs) in digital images. The goal of object detection is typically to predict a set of bounding boxes and class labels for each object of interest. Object detection tasks can be widely used in real-world applications, such as autonomous driving, video surveillance, and robot vision, among others.
[0032] The Detection Transformer (DETR) and its variants (collectively referred to herein as DETR or DETR models) have evolved into a new paradigm for object detection. DETR uses a Transformer as a detection head to directly predict a set of objects by interacting between object queries and visual features, without multiple hand-designed components such as anchors and NMS.
[0033] However, existing DETRs suffer from large computational overhead due to their heavy backbone (e.g., ResNet-101) and Transformer-based head. Directly shrinking the model (e.g., using a lightweight backbone or fewer encoder / decoder layers) can significantly degrade the detection performance.
[0034] Knowledge distillation is an effective technique to boost the performance of low-capacity models by transferring “dark knowledge” from larger models. To transfer the knowledge of a complex model as a teacher model to a simplified model as a student model that is friendly for deployment, the output of the trained teacher model can be provided to the student model for training.
[0035] Due to their specialized design, existing works are often not applicable or difficult to apply to DETRs (sparse object detectors). Unlike previous works, this disclosure discloses how to apply knowledge distillation to learn efficient DETRs. To achieve this goal, two aspects can be considered: 1) how to design a reasonable student network, i.e., which components should be shrunk; 2) how to leverage knowledge distillation to improve the performance of the student network, i.e., where to distill and how to distill effectively.
[0036] For the first aspect, based on the observation that the transformer-based heads of DETR usually take up a large portion of the computational overhead, in this disclosure, additional directions for distilling DETR are mainly considered: distillation for transformer-based heads. And for the second aspect, through comparison between different distillation strategies, feature distillation on the transformer encoder is considered to be a more effective way to learn well-performing student models.
[0037] Based on the above, in order to further improve the effectiveness of encoder distillation, a query-guided knowledge distillation (QGD) strategy is disclosed herein, which is driven by the idea that where a model pays attention is more important. Specifically, attention maps corresponding to multi-head and object queries are used to guide the imitation of encoder features, which can facilitate students to learn discriminative features effectively. Since DETR takes a lot of time to learn reasonable attention maps, the guidance of query attention allows the student network to bypass rough attention to quickly converge. More details will be discussed below in conjunction with the drawings.
[0038] Figure 1 An example structure of a detection transformer is shown, in accordance with aspects of the present disclosure.
[0039] Existing object detection methods can be roughly divided into two-stage methods and one-stage methods. Two-stage methods usually generate dense proposals via RPN, and then extract RoI features to refine the prediction results. One-stage methods directly predict bounding boxes and class probabilities using dense anchor boxes or anchor points. In these methods, duplicate predictions are removed by NMS.
[0040] Detection transformer (DETR) and its variants treat object detection as a set prediction problem, directly predicting the bounding boxes and classes of objects. The absence of redundant designs such as anchors and NMS makes DETR more concise. As shown in Figure 1 An example DETR model will consist of a backbone 101, a transformer encoder 102, a transformer decoder 103, and a plurality of prediction heads 105, as shown. Additionally, the transformer decoder 103 can take as input a small fixed number of learned positional embeddings, referred to as object queries 104, and the number of prediction heads 105 is equal to the fixed number of object queries 104.
[0041] As an example, DETR uses a regular CNN as backbone 101 to learn a 2D representation of the input image, which in real-world applications can be an image captured by a camera mounted on a vehicle. The model flattens it and supplements it with positional encodings before passing it to the transformer encoder 102. The transformer decoder 103 then takes a fixed number of object queries 104 as input and additionally focuses on the transformer encoder output. Each output embedding of the decoder 103 will be passed to the prediction head 105, which predicts either a detection (class label and bounding box) or the “no object” class. Figure 1 An example is shown in which the second and fourth object queries are detected as “no object”.
[0042] However, existing DETR-like methods still suffer from heavy backbones and transformer heads. Therefore, knowledge distillation is considered as a method to improve the efficiency of DETR.
[0043] Figure 2 An example structure of knowledge distillation is shown in accordance with aspects of the present disclosure.
[0044] Knowledge distillation is a method for learning a compact model, which can be referred to as a student model, with supervision of a large and accurate model, which can be referred to as a teacher model. As Figure 2 As shown, the model 202 on the top side is the teacher model, while the model 205 on the bottom side is the student model, which is also referred to as the distilled model. Both the teacher model and the student model have multiple layers, in an aspect, the student model has fewer layers than the teacher model.
[0045] It is generally known that the objective function for training should reflect the user’s real objective as closely as possible. Nonetheless, models are often trained to optimize performance on training data, while the real objective is to generalize well to test data that is new to the model. The goal is to train a model to generalize well, which requires information about the right way to generalize, and this information is often not available.
[0046] A promising approach to transfer the generalization capability of the teacher model to the student model is to train the student model using the classification probabilities produced by the teacher model as “soft targets” 210. It is more beneficial to introduce, in addition to the soft targets 210, hard targets 211 computed based on the difference between the output of the student model and the ground truth labels as part of the objective function, which encourages the student model to predict the real objective as well as the labels provided by the teacher model.
[0047] Returning to Figure 2teacher model 202 is first trained, and the outputs of the teacher model are distilled by temperature to produce a set of soft targets 204. Then the student model 205 is trained, and the outputs of the student model are distilled by the same temperature to produce a set of soft predictions 207, except that the outputs of the student model have a temperature set to 1 to produce a set of hard predictions 208. A soft loss 210 is generated based on the set of soft targets 204 and the set of soft predictions 207, and a hard loss 211 is generated based on the set of hard predictions 208 and the ground truth labels 209. The soft loss 210 and the hard loss 211 can be combined to train the student model 205.
[0048] To apply knowledge distillation to DETR, first an empirical study is conducted to explore the above two aspects. As an example, conditional DETR is mainly considered in this section due to its simplicity and faster convergence. It should be noted that those skilled in the art can easily understand that other DETRs will have similar characteristics to conditional DETR.
[0049] First, from the inference time analysis of the key components of DETR and conditional DETR, including the backbone, transformer encoder, and decoder. Figure 3 An example of the inference time of DETR is shown according to aspects of the present disclosure.
[0050] As Figure 3 shown, the inference time of each key component of DETR is shown, where (a) represents DETR and (b) represents conditional DETR, where 301 and 301' indicate the inference time of the backbone, 302 and 302' indicate the inference time of the transformer encoder, 303 and 303' indicate the inference time of the transformer decoder, and 304 and 304' indicate the time cost of pre-processing and post-processing. In Figure 3 It can be observed in that the inference time of the transformer-based head (including both the encoder and the decoder) is as large as the inference time of the backbone. For example, in conditional DETR, the transformer head accounts for 40% of the total inference time (302' + 303'), which is close to the time of the backbone (301') (42%). This indicates that, in addition to the backbone, the compression of the transformer head is also essential to build an efficient student network.
[0051] Then, an ablation experiment is conducted to explore the impact of each component of conditional DETR on the final detection performance. The number of layers of the transformer encoder and decoder is gradually reduced to observe their impact on mAP. As shown in Table 1 below, the number of decoder layers has a more significant impact on detection performance than the number of encoder layers. R18-x-y indicates that the model (ResNet-18 as the backbone) contains x layers of encoder and y layers of decoder, and AP change relative to R50-6-6. Table 1 - Ablation study of each component in DETR under condition D
[0052] As can be seen from Table 1, retaining half of the encoder and decoder layers (i.e., R18-3-3) still achieves reasonable performance (4.6% AP drop), but further reducing the number of decoder layers results in significant performance drop. The above analysis shows that it is more reasonable to compress the transformer encoder than the decoder in DETR.
[0053] Based on the above, further study is conducted on where to distill DETR. Existing knowledge distillation methods can be roughly divided into two types: distillation on features and distillation on logits. Following these two paradigms, feature distillation for backbone and transformer encoder and logit distillation for transformer decoder are used respectively.
[0054] Figure 4 An example distillation strategy is shown in accordance with aspects of the present disclosure. In Figure 4 the above DETR model T represents a teacher model, and the following DETR model S represents a student model, where each model has a similar structure to Figure 1 For further explanation, 101 T and 101 S represent the backbone of the teacher model and the student model, 102 T and 102 S represent the transformer encoder of the teacher model and the student model, 103 T and 103 S represent the transformer decoder of the teacher model and the student model, and 105 T and 105 S represent the output object query of the teacher model and the student model. Figure 4 Three distillation strategies 401-403 are shown, which will be discussed in more detail below.
[0055] Now begin, distillation strategy 401 can be referred to as backbone knowledge distillation (KD), which means that the learned knowledge can be transferred from the teacher's backbone to the student's backbone through feature distillation. As an example, distillation is performed directly on the single-scale output features of the backbone (C5 features of ResNet). As an example, the difference between the features of the teacher and student backbone can be calculated as an L2 loss: (1) where , are the features from the teacher and student backbone respectively, and is a 1X1 convolutional layer for aligning channels.
[0056] Since the transformer encoder in DETR is usually a module to enhance the features generated from the backbone and takes considerable computational overhead, a simple solution is to apply the feature distillation strategy directly to the encoder, which is shown as distillation strategy 402 in Figure 4 , referred to as encoder-KD. Like backbone distillation, as an example, L2 loss is used to transfer the learned knowledge of the teacher encoder to the student encoder: (2) where , and denote the output features of the teacher and student encoders and the fully connected layer, respectively.
[0057] Another way is to extract the prediction results of the transformer decoder, referred to as query-KD, and is shown as extraction strategy 403 in Figure 4 . As an example, KL divergence is used to transfer classification knowledge. In addition, the attention map of the decoder is also extracted, which can further improve the performance. The class probabilities and attention maps from the student and teacher are denoted as , and , , respectively. The loss is: (3) where and are the set of target queries and the correspondence between the teacher queries and the student queries, respectively. An important problem of query-KD is how to determine the correspondence . Two types of correspondences are explored: direct one-to-one matching of all queries (referred to as slots) and matching of positive queries determined by assigned ground truth instances (referred to as GT).
[0058] In the case of matching by slots, the student queries and the teacher queries are matched in a direct one-to-one manner (i.e., ). All queries are used for knowledge distillation (i.e., contains all queries). Since the object queries of the student network are randomly initialized, this method will align the student's queries with the teacher's queries. While in the case of matching by GT, this method matches queries by assigned ground truth instances. The binary matching training in DETR establishes a one-to-one correspondence between queries and ground truth instances (GT). This correspondence is reused to match the student queries and the teacher queries that correspond to the same GT. In this case, only the queries that match the GT (i.e., positive queries) are used.
[0059] Table 2 below shows a comparison of the three distillation strategies. Empirical experiments are conducted on the MSCOCO dataset and detection AP is reported on the validation set. The teacher network is trained for 50 epochs and ResNet-50 is used as the backbone. The student network is trained for 24 epochs for a quick comparison. From Table 2, we can conclude that (1) for DETR, feature distillation (on the backbone and encoder) outperforms logit distillation (on the decoder); (2) for query-KD, aligning with the ground truth labels leads to a 1.2% AP improvement; (3) applying feature distillation to the encoder works better than to the backbone. Table 2 - Comparison of distillation strategies
[0060] Therefore, based on empirical findings, combining query-KD with encoder-KD will bring modest gains on DETR. Thus, the present disclosure focuses on encoder-KD rather than the backbone and decoder due to its effectiveness and simplicity. How to effectively extract learning knowledge from the encoder features will be discussed in more detail below.
[0061] First, briefly review the cross-attention mechanism of the transformer decoder in DETR. Given a set of object queries (e.g., 104 in Figure 1 , DETR aggregates features from image features (obtained from the encoder, e.g., 102 in ) using cross-attention Figure 1 , , (4) where denotes the head (in multi-head attention), and , and are obtained by linear projections of and , respectively. This can be interpreted as weighting image features from the encoder with an attention map . After that, output features are obtained by concatenating features from different heads and applying a linear transformation to obtain output features , (5) where denotes the number of multi-heads.
[0062] Motivated by the idea that where a model pays attention matters, a query-guided knowledge distillation (QGD) strategy for encoder-KD is proposed in the present disclosure. Figure 5An example structure of query-guided knowledge distillation is shown.
[0063] In Figure 5 the same reference signs are used to refer to parts that are the same or similar in Figure 4 the same or similar to those in
[0064] The attention maps corresponding to different object queries to emphasize valuable encoder features in the teacher model are reused to implement the idea of emphasizing valuable regions. Intuitively, regions with larger attention weights have a stronger impact on the final prediction and are therefore more important. For Figure 5 In the example in
[0065] Therefore, query-guided knowledge distillation can be derived as a loss weighted by the multi-headed attention map via the query. As an example, it can be derived as an L2 loss weighted by the multi-headed attention map via the query, derived as: (6) (7)
[0066] In an example, may represent a set of all object queries, such as the per-slot matching case described earlier. In another example, to further improve the AP, may represent a set of object queries determined (i.e., queries matching with ground truth instances), such as the matching by assigned ground truth instances case described earlier. For the original transformer head, is equivalent to weighting feature distillation with the average attention map using different heads and object queries, as follows: (8)
[0067] In the present disclosure, in an example, feature distillation can be applied to the last encoder layer, but in another example, it can also be applied to previous layers of the encoder.
[0068] In the present disclosure, as an example, with the attention map from the last layer of the decoder, in another example, the attention maps from other layers of the decoder can be used.
[0069] To combine the attention maps from different models and balance the ratio of the two attention maps, a hyperparameter is introduced to balance the ratio of the two attention maps from the teacher model and the student model, resulting in the following general form: (9) where and represent attention maps from student and teacher, respectively. With different values , the attention map in equation (6) can come from the teacher model or the student model, or both. As another example, for deformable attention, it can refer to equation (6) and apply the same sampling points as the query to the L2 loss.
[0070] In this disclosure, it is further disclosed how to effectively transfer knowledge among different layers of the encoder, as the encoder is usually composed of multiple transformer layers. The function of a transformer encoder can be viewed as an iterative process of refining image features. By guiding the learning of earlier layers of the student encoder, rather than just the last layer, the quality of its final output features can be further improved. Two distillation options are proposed in this disclosure, as shown in Figure 6 .
[0071] In an example, the shallow features of the student can be guided by the deep features of the teacher, as shown in (a) of Figure 6 . As an example, the features of each layer of the encoder of the student model can be guided with the output features of the last layer of the encoder of the teacher model as the shared target for the different student layers.
[0072] In another example, the shallow features of the student can be guided by the corresponding shallow features of the teacher, as shown in (b) of Figure 6 . As an example, the features of each layer of the encoder of the student model can be guided with the features of the corresponding intermediate layers of the encoder of the teacher model in a progressive manner.
[0073] The total distillation loss is defined as follows: (10) where is the number of student encoder layers, is the multiplicity of the number of teacher layers relative to the number of student layers (for deep supervision, is replaced by the number of teacher encoder layers), and denotes the output of the l encoder layer, is the loss weight of the l layer. Then is added to the original DETR loss for end-to-end training.
[0074] Figure 7An example flow diagram for learning an efficient DETR is shown in accordance with aspects of the present disclosure. As described below, some or all of the illustrated features can be omitted in particular implementations within the scope of the present disclosure, and some illustrated features can not be required for implementation of all embodiments. Moreover, some of the blocks can be performed in parallel or in a different order. In some examples, the method can be performed by any suitable means for performing the functions or algorithm steps described below.
[0075] The method begins at block 701, where a first detection transformer (DETR) model is trained. As an example, the first DETR model includes at least a backbone, a transformer encoder, a transformer decoder, where the encoder includes a plurality of transformer layers.
[0076] The method then proceeds to block 702, where a second DETR model is trained based on a loss that weights a difference in features between layers of the encoder of the trained first DETR model and the encoder of the second DETR model with an average attention map. In an aspect of the present disclosure, the method can be embodied in equation (7) above. As an example, the second DETR model includes at least a backbone, a transformer encoder, a transformer decoder, where the encoder includes a plurality of transformer layers. As an example, the encoder of the second DETR model has fewer layers than the encoder of the first DETR model.
[0077] As an example, the average attention map is based on one or more of an average of the attention maps of the multiple heads of the decoder of the trained first DETR model and one or more object queries, or an average of the attention maps of the multiple heads of the decoder of the second DETR model and one or more object queries. In an aspect of the present disclosure, the method can be embodied in equation (8) above.
[0078] As another example, the average attention map is based on a ratio combination of an average of the attention maps of the multiple heads of the decoder of the trained first DETR model and one or more object queries and an average of the attention maps of the multiple heads of the decoder of the second DETR model and one or more object queries. In an aspect of the present disclosure, the method can be embodied in equation (9) above.
[0079] As an example, the attention map of the trained first DETR model and the attention map of the second DETR model are attention maps from a last layer of the respective decoder.
[0080] As an example, the one or more object queries of the decoder of the trained first DETR model and the one or more object queries of the decoder of the second DETR model are deterministic object queries.
[0081] As another example, one or more object queries of the decoder of the first DETR model and one or more object queries of the decoder of the second DETR model are all object queries.
[0082] As an example, the feature difference is calculated between the output features of each layer of the encoder of the second DETR model and the output features of the final layer of the encoder of the trained first DETR model. In one aspect of this disclosure, the method can be described above... Figure 6 The deep supervision method described in (a) is implemented.
[0083] As an example, the feature difference is calculated between the output features of each layer of the encoder of the second DETR model and the output features of each corresponding intermediate corresponding layer of the encoder of the trained first DETR model. In one aspect of this disclosure, the method can be used with the above... Figure 6 The gradual approach described in (b) is implemented.
[0084] As another example, training the second DETR model is also based on a loss that evaluates the difference between the output of the decoder in the second DETR model and the ground truth label. In one aspect of this disclosure, the loss of equation (6) or (10) can be added to the original DETR loss for end-to-end training.
[0085] Taking real-world applications as an example, both the first and second DETR models can be trained for object detection in autonomous driving. The training sets for both models can be collections of images, each containing various objects such as traffic signs, road surfaces, pedestrians, and vehicles. The expected outputs of both models are the categories and bounding boxes corresponding to the objects in the images.
[0086] Figure 8 Exemplary flowcharts for using a trained DETR according to various aspects of this disclosure are shown. As described below, some or all of the shown features may be omitted in specific implementations within the scope of this disclosure, and some shown features may not be required for implementations of all embodiments. Furthermore, some blocks may be executed in parallel or in different orders. In some examples, the method may be executed by any suitable means or unit for performing the functions or algorithms described below.
[0087] The method begins at box 801, where image data captured during vehicle movement is obtained. As an example, image data can be captured using devices such as cameras, lidar, etc., mounted on the vehicle.
[0088] The method proceeds to box 802, where the image data is input into the DETR model. The DETR model can be...Figure 7 One of the described methods trains a second DETR model.
[0089] The method proceeds to block 803, where object class labels and bounding boxes corresponding to one or more object queries of the DETR model are output.
[0090] The method proceeds to block 804, where the vehicle is assisted in traveling based on the output object class labels and bounding boxes.
[0091] Figure 9 An example computer system is shown in accordance with various aspects of the present disclosure. The computer system can include at least one processor 910. The computer system can also include at least one storage device 920. It should be understood that the storage device 920 can store computer-executable instructions that, when executed, cause the processor 910 to perform any of the operations described in accordance with embodiments of the present disclosure. Figures 1-8
[0092] Embodiments of the present disclosure can be embodied in one or more computer-readable media, such as non-transitory computer-readable media. The non-transitory computer- readable media can store instructions that, when executed, cause one or more processors to perform any of the operations described in accordance with embodiments of the present disclosure. Figures 1-8
[0093] Embodiments of the present disclosure can be embodied in a computer program product that includes computer-executable instructions that, when executed, cause one or more processors to perform any of the operations described in accordance with embodiments of the present disclosure. Figures 1-8
[0094] Embodiments of the present disclosure can be embodied in a vehicle that includes one or more devices for performing any of the operations described in accordance with embodiments of the present disclosure. Figure 8
[0095] It should be understood that all of the operations in the above-described methods are merely exemplary and the present disclosure is not limited to any of the operations in the methods or the order of the operations, and should cover all other equivalents under the same or similar concepts.
[0096] It should also be understood that all of the modules in the above-described devices can be implemented in various methods. The modules can be implemented as hardware, software, or a combination thereof. Furthermore, any of the modules can be further divided functionally into sub-modules or combined together.
[0097] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. In view of the many possible embodiments to which the principles defined herein can be applied, it should be recognized that the elements recited in the claims are not to be construed as limiting. Rather, the scope of the disclosure is to be determined by the following claims.
Claims
1. A computer-implemented method for machine learning, comprising: Train the first Detection Transformer (DETR) model; A second DETR model is trained based on a loss that uses an average attention map to weight the feature differences between layers of the encoders of the trained first DETR model and the second DETR model, wherein the encoder of the second DETR model has fewer layers than the encoder of the first DETR model; and The average attention map is based on one or more of the following: the average of the attention maps of the multi-head decoder of the trained first DETR model and one or more object queries, or the average of the attention maps of the multi-head decoder of the second DETR model and one or more object queries.
2. The method according to claim 1, wherein, The average attention map is a combination of the ratio of the average of the attention maps of the multi-head and one or more object queries of the decoder of the first DETR model to the average of the attention maps of the multi-head and one or more object queries of the decoder of the second DETR model.
3. The method according to claim 2, wherein, The attention maps of the first and second DETR models are attention maps from the last layer of their respective decoders.
4. The method of claim 1, wherein the one or more object queries of the decoder of the first DETR model trained thereon and the one or more object queries of the decoder of the second DETR model are determined object queries.
5. The method according to claim 1, wherein, The feature difference is calculated between the output features of each layer of the encoder of the second DETR model and the output features of the final layer of the encoder of the trained first DETR model.
6. The method according to claim 1, wherein, The feature difference is calculated between the output features of each layer of the encoder of the second DETR model and the output features of each corresponding intermediate layer of the encoder of the trained first DETR model.
7. The method according to claim 1, wherein, Training the second DETR model is also based on a loss that evaluates the difference between the output of the decoder of the second DETR model and the ground truth label.
8. The method according to claim 1, wherein, The first DETR model and the second DETR model are trained using image data, and the outputs of the first DETR model and the second DETR model are category labels and bounding boxes corresponding to one or more object queries.
9. A computer-implemented method for object detection using a DETR model trained according to any one of claims 1-8, comprising: Obtain image data captured while the vehicle is in motion; The image data is input into the DETR model; The DETR model outputs object category labels and bounding boxes corresponding to one or more object queries; as well as The output object category labels and bounding boxes are used to assist vehicle driving.
10. A computer system, comprising: One or more processors; as well as One or more storage devices storing computer-executable instructions, which, when executed, cause the one or more processors to perform the operation of the method according to any one of claims 1-9.
11. One or more computer-readable storage media storing computer-executable instructions, which, when executed, cause one or more processors to perform the operation of the method according to any one of claims 1-9.
12. A computer program product comprising computer-executable instructions, which, when executed, cause one or more processors to perform the operation of the method according to any one of claims 1-9.
13. A vehicle comprising one or more means for performing the method according to claim 9.