End-to-end target detection framework based on representative points

By using an end-to-end detection framework based on representative points in object detection, the problem of rough surrounding frame representation in the prior art is solved, and a more refined target representation and more efficient detection effect are achieved.

CN120070839APending Publication Date: 2025-05-30HUBEI PROVINCE TOBACCO CO TIANMEN CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311599359.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the enclosure box is used to represent the target, with a coarse particle size, and the enclosure box often contains a large number of background areas, which makes the learned representation not fine enough, thereby affecting the final detection effect.

Method used

An end-to-end object detection framework based on representative points is proposed, including detectors, encoders and decoders. The encoder is used to extract the predicted dense representative points and target categories on the refinement feature map. The decoder updates the query and representative points coordinates through the self-attention mechanism and the cross-attention mechanism, and finally outputs the minimum enclosure box of the representative points.

Benefits of technology

Through fine-grained representative dot representation, the semantic significant areas of the target can be captured more precisely, thereby improving the accuracy and efficiency of target detection and reaching the industry-leading level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070839A_ABST
    Figure CN120070839A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image detection, and particularly relates to an end-to-end target detection framework based on representative points, which comprises a detector, an encoder and a decoder, the encoder is used for extracting representative points with dense prediction and target categories on the refined feature map; the detector is used for taking out the first N targets in the representative points as the query of the Transform decoder; the decoder is used for updating query and representative point coordinates; a sparse query set is introduced to represent a target, features of fine granularity of the target are mined through representative points, an end-to-end target detection model is obtained, the framework reaches the industry leading level on a public detection standard data set, and the framework has more accurate, faster and smaller effects; the trained model is directly loaded to an end reasoning chip, and target detection of an actual scene can be carried out, such as pedestrian detection, smoke and fire detection, smoke alarm, pedestrian recognition and other specific application fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image detection, and specifically relates to an end-to-end object detection framework based on representative points. Background Art

[0002] An object detector is a computer vision algorithm that can identify and locate specific objects in an image or video. These algorithms are commonly used in various scenarios, such as autonomous driving, security monitoring, medical image analysis, and other fields.

[0003] Currently, in the existing technology, the object detector accurately locates and detects the target through a bounding box. Specifically, in the operation, the bounding box is parameterized as the coordinates of the center point of the target and the width and height, or the coordinates of the upper vertex and the lower vertex. During the training process of the network, the deviation from the pre-set anchor box is predicted, and then this is compared with the true deviation of the actual target to calculate the training loss.

[0004] However, using a bounding box to represent the target has a relatively coarse granularity, and the bounding box often contains a large amount of background area, making the learned representation not fine enough, which in turn affects the final detection effect. Summary of the Invention

[0005] In order to solve the problem in the existing technology that using a bounding box to represent the target has a relatively coarse granularity, and the bounding box often contains a large amount of background area, making the learned representation not fine enough, which in turn affects the final detection effect, the present invention proposes an end-to-end object detection framework based on representative points.

[0006] The technical solution adopted by the present invention to solve its technical problems is: an end-to-end object detection framework based on representative points according to the present invention includes a detector, an encoder, and a decoder; the encoder is used to extract dense representative points and target categories predicted on the refined feature map; the detector is used to take out the top N targets among the representative points as the query of the Transformer decoder; the decoder is used to update the query and the coordinates of the representative points.

[0007] Preferably, the encoder processes and extracts features from the encoded data in cooperation with the backbone network.

[0008] Preferably, the decoder updates the query and the coordinates of the representative points using the self-attention mechanism and the cross-attention mechanism.

[0009] Preferably, the final output result uses the maximum-minimum method to obtain the minimum bounding box of the representative points.

[0010] Preferably, the decoder performs two samplings on the representative points using bilinear interpolation sampling to obtain the representative point features.

[0011] Preferably, the backbone network uses ResNet50 as the pre-trained neural network foundation.

[0012] Preferably, the detection framework uses the Transformer structure to introduce a sparse query set to represent image targets.

[0013] Preferably, the end-to-end object detection framework detection method based on representative points includes the following steps:

[0014] S1: Input an 800x1300 image into the visual Backbone to extract a feature map HxWxC with rich semantics.

[0015] S2: Input the extracted feature map into the stacked Transformer encoder, and use the self-attention mechanism in the Transformer to obtain a refined feature map with global perception.

[0016] S3: Use a representative point predictor composed of linear layers to slide a window on the HxWxC feature map to output dense representative point results and class estimates as a rough estimate of the target.

[0017] S4: According to the class estimate score, select the top 300 estimated targets with confidence and obtain K representative points corresponding to the targets.

[0018] S5: Use bilinear interpolation sampling to obtain representative point features and use average pooling to aggregate them to obtain a sparse query set.

[0019] S6: Input the sparse query set into the Transformer decoder based on the self-attention mechanism and cross-attention mechanism, and mine the target saliency region features through the K representative points corresponding to the target to update the query.

[0020] S7: Input the updated query into the localization and classification head, where the localization head outputs updated representative points and the classification head outputs the target class.

[0021] S8: Convert the refined representative points into bounding boxes as the final detection result.

[0022] The beneficial effects of the present invention are:

[0023] 1. The present invention provides an end-to-end target detection framework based on representative points. The detection framework mainly consists of two parts: a dense representative point detector based on a sliding window and a Transformer representative point detection decoder based on a sparse query. The dense representative point detector predicts dense representative points and target categories on the refined feature map extracted by the backbone network and the Transformer encoder, and then takes out the top N targets as the query of the Transformer decoder through the category confidence, and then uses the attention mechanism and cross attention mechanism in the Transformer decoder to update the query and representative point coordinates. Finally, the classification head is used to predict the categories of N targets, and the maximum and minimum transformation is used to obtain the optimal representative point. The small bounding box is used as the final output of the detection result. Different from the previous target detector that uses the bounding box representation, the target is represented as a more fine-grained representative point to mine the semantically salient area of ​​the target, thereby providing a better representation for target positioning and recognition. This framework also combines the advanced Transformer structure to introduce a sparse query set to represent the target, and uses representative points to mine the fine-grained features of the target to obtain an end-to-end target detection model. This framework has reached the industry-leading level on public detection standard datasets, with more accurate, faster, and smaller effects. The trained model can be directly loaded onto the on-end inference chip to perform target detection in actual scenarios, such as pedestrian detection, fireworks detection, smoke alarms, pedestrian recognition and other specific application areas.

[0024] 2. The present invention provides an end-to-end target detection framework based on representative points, which adopts dual-line difference sampling at the decoder layer. The sampling rate of the signal can be increased by sampling the signal twice and then performing linear difference between the two sampling points. This enables the decoder to obtain better results when restoring images or audio signals and reduce distortion and noise. In addition, dual-line difference sampling can also reduce the computational complexity of the decoder and improve the decoding speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0026] Figure 1 is a flow chart of the present invention;

[0027] Figure 2 This is a Transformer target detection framework diagram based on representative point modeling in the present invention;

[0028] Figure 3 A schematic diagram of modeling representative points in the present invention;

[0029] Figure 4 This is the design diagram of the Transformer decoder based on representative points in the present invention; Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] Please refer to Figures 1 - 4 , the present invention provides an end-to-end object detection framework based on representative points, including a detector, an encoder, and a decoder; the encoder is used to extract densely predicted representative points and object categories on the refined feature map; the detector is used to take out the top N objects in the representative points as the queries of the Transformer decoder; the decoder is used to update the queries and the coordinates of the representative points. The detection framework mainly consists of two parts, namely: a dense representative point detector based on a sliding window and a Transformer representative point detection decoder based on sparse queries. The dense representative point detector predicts dense representative points and object categories on the refined feature map after being extracted by the backbone network and the Transformer encoder, and then takes out the top N objects as the queries of the Transformer decoder through the class confidence. Then, the attention mechanism and the cross-attention mechanism in the Transformer decoder are used to update the queries and the coordinates of the representative points. Finally, the classification head is used to predict the categories of the N objects, and the minimum bounding box of the representative points is obtained through the maximum-minimum transformation as the final output detection result. Different from previous object detectors that use the representation of bounding boxes, the objects are represented as more fine-grained representative points to mine the semantic significant regions of the objects, so as to provide a better representation for the localization and recognition of the objects. This framework also combines an advanced Transformer structure, represents the objects by introducing a sparse set of queries, and mines the fine-grained features of the objects through the representative points, obtaining an end-to-end object detection model. This kind of framework has reached the leading level in the industry on the publicly available detection standard datasets, with the effects of being more accurate, faster, and smaller. The trained model can be directly loaded onto the on-chip inference chip for object detection in actual scenarios, such as pedestrian detection, fireworks detection, smoke alarm, pedestrian recognition and other specific application fields.

[0032] It should be noted here that in a Transformer, a query is a special input vector. In the detection framework, the target categories and representative point coordinates predicted by the dense representative point detector are used as the queries for the Transformer decoder.

[0033] Furthermore, as Figure 2 shown, the encoder processes and extracts features from the encoded data in cooperation with the backbone network.

[0034] Furthermore, as Figure 4 shown, the decoder updates the queries and representative point coordinates using the self-attention mechanism and the cross-attention mechanism. The decoder based on the sparse representative point set uses the cross-attention mechanism of the Transformer to mine the features of the target semantic significant regions, which can provide a more fine-grained perception of the target compared with the traditional bounding box representation.

[0035] Furthermore, as Figure 3 shown, the final output result uses the maximum-minimum method to obtain the minimum bounding box of the representative points. The maximum-minimum transformation is the process of calculating the minimum bounding box from the dense representative points. This method can ensure that even if multiple representative points coincide, only one minimum bounding box is needed to contain them, thus reducing the number of redundant boxes and improving the detection efficiency. The representative point modeling of the present invention can capture the features of the target semantic significant regions, and the minimum bounding box of the representative point set can be obtained through the maximum-minimum transformation as the final output result.

[0036] Furthermore, the decoder uses bilinear interpolation sampling to sample the representative points twice to obtain the representative point features. Using bilinear interpolation sampling in the decoder layer can improve the sampling rate of the signal by sampling the signal twice and then performing linear interpolation between the two sampling points. This can enable the decoder to obtain better results when restoring the image or audio signal, reducing distortion and noise. In addition, bilinear interpolation sampling can also reduce the computational complexity of the decoder and improve the decoding speed.

[0037] Furthermore, the backbone network uses ResNet50 as the basis of the pre-trained neural network. On the validation set of the COCO dataset, for the detection model with ResNet50 as the backbone network, only by training for 12 epochs, it can reach an accuracy of 48.0 mAP at 25 FPS.

[0038] It should be noted here that: An epoch refers to a complete learning process for the entire training set. Therefore, 12 epochs means that only 12 complete training iterations of the model are required to reach the expected performance level.

[0039] Furthermore, the detection framework uses a Transformer structure to introduce a sparse query set to represent image targets.

[0040] Furthermore, the end-to-end object detection framework detection method based on representative points includes the following steps:

[0041] S1: Input an 800x1300 image into the visual Backbone to extract a feature map HxWxC with rich semantics.

[0042] S2: Input the extracted feature map into a stacked Transformer encoder, and use the self-attention mechanism in the Transformer to obtain a refined feature map with global perception.

[0043] S3: Use a representative point predictor composed of linear layers to slide a window on the HxWxC feature map to output dense representative point results and class estimates as a rough estimate of the target.

[0044] S4: According to the class estimate scores, select the top 300 estimated targets with confidence and obtain K representative points for the corresponding targets.

[0045] S5: Use bilinear interpolation sampling to obtain representative point features and use average pooling to aggregate them into a sparse query set.

[0046] S6: Input the sparse query set into a Transformer decoder based on self-attention and cross-attention mechanisms, and mine the target saliency region features through the K representative points of the corresponding target to update the query.

[0047] S7: Input the updated query into the localization and classification head, where the localization head outputs updated representative points and the classification head outputs the target class.

[0048] S8: Convert the refined representative points into bounding boxes as the final detection result.

[0049] Working principle:

[0050] The detection framework mainly consists of two parts: a dense representative point detector based on sliding windows and a Transformer representative point detection decoder based on sparse queries. The dense representative point detector predicts dense representative points and target categories on the refined feature map extracted by the backbone network and the Transformer encoder, and then takes out the top N targets as the query of the Transformer decoder through the category confidence, and then uses the attention mechanism and cross attention mechanism in the Transformer decoder to update the query and representative point coordinates. Finally, the classification head is used to predict the categories of N targets, and the minimum bounding box of the representative point is obtained as the final output detection result through the maximum and minimum transformation. The test results show that unlike previous target detectors that use bounding boxes as representations, this framework represents the target as more fine-grained representative points to mine the semantically significant areas of the target, thereby providing better representation for target positioning and recognition. This framework also combines the advanced Transformer structure to introduce a sparse query set to represent the target, and uses representative points to mine the fine-grained features of the target, thus obtaining an end-to-end target detection model. This framework has reached the industry-leading level on public detection standard datasets, with more accurate, faster, and smaller effects. The trained model can be directly loaded onto the end-to-end inference chip to perform target detection in actual scenarios, such as pedestrian detection, fireworks detection, smoke alarm, pedestrian recognition and other specific application areas;

[0051] The two-line difference sampling is used at the decoder layer. The sampling rate of the signal can be increased by sampling the signal twice and then performing a linear difference between the two sampling points. This allows the decoder to achieve better results when restoring the image or audio signal and reduce distortion and noise. In addition, the two-line difference sampling can also reduce the computational complexity of the decoder and increase the decoding speed.

[0052] The decoder based on the sparse representative point set uses the cross-attention mechanism of Transformerd to mine the semantically salient area features of the target, which can provide a more fine-grained perception of the target compared to the traditional bounding box representation.

[0053] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected.

Claims

1. An end-to-end object detection framework based on representative points, characterized in that: It includes a detector, an encoder, and a decoder; the encoder is used to extract densely predicted representative points and object categories on the refined feature map; the detector is used to take the top N objects among the representative points as the queries of the Transformer decoder; the decoder is used to update the queries and the coordinates of the representative points.

2. The end-to-end object detection framework based on representative points according to claim 1, characterized in that: The encoder processes and extracts features from the encoded data in cooperation with the backbone network.

3. The end-to-end object detection framework based on representative points according to claim 2, characterized in that: The decoder uses the self-attention mechanism and the cross-attention mechanism to update the queries and the coordinates of the representative points.

4. The end-to-end object detection framework based on representative points according to claim 3, characterized in that: The final output result uses the maximum-minimum method to obtain the minimum bounding box of the representative points.

5. The end-to-end object detection framework based on representative points according to claim 4, characterized in that: The decoder performs two samplings on the representative points using bilinear interpolation sampling to obtain the representative point features.

6. The end-to-end object detection framework based on representative points according to claim 5, characterized in that: The backbone network uses ResNet50 as the basis of the pre-trained neural network.

7. The end-to-end object detection framework based on representative points according to claim 6, characterized in that: The detection framework introduces a sparse query set to represent image objects through the Transformer structure.

8. The end-to-end object detection framework based on representative points according to claim 7, characterized in that: The detection method of the end-to-end object detection framework based on representative points includes the following steps: S1: Input an 800x1300 image into the visual Backbone to extract a semantically rich feature map HxWxC; S2: Input the extracted feature map into the stacked Transformer encoder, and use the self-attention mechanism in the Transformer to obtain a refined feature map with global perception; S3: Use a representative point predictor composed of linear layers to slide the window on the HxWxC feature map to output dense representative point results and class estimates as a rough estimate of the object; S4: According to the class estimate score, take out the top 300 estimated objects with confidence and obtain K representative points of the corresponding objects; S5: Use bilinear interpolation sampling to obtain representative point features and use average pooling aggregation to obtain a sparse query set; S6: Input the sparse query set into the Transformer decoder based on the self-attention mechanism and the cross-attention mechanism, and mine the target saliency region features through the K representative points of the corresponding object to update the queries; S7: Input the updated query into the localization and classification head, where the localization head outputs an updated representative point and the classification head outputs the target category; S8: Convert the refined representative point into a bounding box as the final detection result.