Improved end-to-end target detection system and method
By improving the encoder structure and loss function training, the number of feature extraction and iterations of Sparse R-CNN is optimized, and the problems of weak feature extraction capabilities and high computational complexity are solved, achieving more efficient object detection.
Patent Information
- Application Number
- CN202510321232.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-08-01
AI Technical Summary
The existing Sparse R-CNN has weak feature extraction capabilities in object detection, and the long iterative structure leads to high parameter quantity and calculation complexity. The existing improved methods have failed to effectively solve this problem.
The improved encoder structure is adopted, including the ConvNeXt-Tiny backbone network, the bridge module and 3 cascaded dynamic heads, shallow and deep feature extraction are performed, combined with the mixed loss function HIoU for training, and the decoder long iterative structure is optimized.
It significantly improves feature extraction capabilities, reduces parameter quantity and calculation complexity, improves detection accuracy and robustness, reduces the number of iterations from 6 to 4 times, improves detection accuracy and reduces calculation costs.
Smart Images

Figure CN120411535A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and computer vision object detection, and particularly relates to an improved end-to-end object detection system and method. Background Art
[0002] In the field of object detection, redundant bounding box generation is common in traditional single-stage and two-stage detection architectures. Their detection processes rely on post-processing mechanisms such as non-maximum suppression (NMS) to complete the screening of repeated predictions. Such detection paradigms based on dense candidate boxes (dense methods) adopt a one-to-many label matching mechanism, and their detection performance is significantly sensitive to preset anchor box parameters (including quantity, size ratio, and generation strategy). In contrast, new object detection algorithms represented by DETR and Sparse R-CNN have pioneered a sparse detection paradigm (sparse methods), which achieves a one-to-one label matching mechanism by presetting a fixed number of queries. This innovative design not only eliminates complex post-processing processes but also constructs a highly concise end-to-end detection system. It is worth noting that Sparse R-CNN redefines the query mechanism, representing it as a bounding box and its corresponding embedding. Based on inheriting the design concept of Faster R-CNN, this architecture introduces a learnable set of candidate boxes and their corresponding feature matrices as the initial input, and gradually improves the detection accuracy through a multi-level iterative optimization mechanism. Different from the encoder-decoder architecture adopted by the DETR series, Sparse R-CNN follows the classic backbone network combined with the feature extraction scheme of FPN (Feature Pyramid Network) and constructs a deep query iterative optimization path. However, the structures of the backbone network and the feature pyramid network for feature extraction are already very old, resulting in poor feature extraction ability and general detection accuracy. Coupled with the fact that the multi-level iterative optimization mechanism of Sparse R-CNN acts on the feature maps output by the backbone network and the feature pyramid network, therefore, this iterative structure requires 6 times to obtain a better optimization result, and this long iterative structure will greatly increase the number of parameters and computational complexity of the algorithm.
[0003] Zheng et al. proposed a progressive component based on Sparse R-CNN in 2022. This component aims to improve the detection accuracy of Sparse R-CNN in dense scenarios. This component acts on the last stage of the long iterative structure of Sparse R-CNN. Specifically, first, it selects effective queries (accepted queries) that are easy to generate true positive predictions, and then optimizes the remaining noisy queries based on the previously accepted prediction results. Although this method focuses on improving the detection accuracy of Sparse R-CNN, it mainly acts on the last stage of the long iterative structure, not only does not improve the feature extraction ability of the algorithm itself, but also brings additional parameter quantities and computational amounts.
[0004] In the patent application "An Object Detection Method Based on Improved Sparse R-CNN" (Patent Application No.: CN202310081364.6) of Yihailante Technology Development (Changsha) Co., Ltd., the first step: image feature extraction step, using the backbone network to extract features from the input image and output the feature map through convolutional processing; the second step: regional feature extraction step, taking the initial candidate boxes and the output feature map obtained in the image feature extraction step as inputs, and using the RoiAlign method for bilinear interpolation processing to extract the regional features of the regions where the initial candidate boxes are located; the third step: regional feature mixing step, taking the initial candidate box features and the regional features obtained in the regional feature extraction step as inputs, and fusing the regional features according to the feature mixing function to obtain the mixed regional features corresponding to each candidate box. The mixed regional features have both high-dimensional abstract semantics and low-dimensional position information at the same time. This method uses the method of regional feature mixing to improve the feature processing flow of Sparse R-CNN in the long iterative stage and enhances the feature processing method of Sparse R-CNN in the long iterative stage, making it more efficient. However, the improvement of this method on Sparse R-CNN still stays at a certain step in the long iterative stage, and this feature mixing function will bring additional parameter quantities and computational amounts, and overall does not improve the problems of weak feature extraction ability and excessive number of iterations of Sparse R-CNN.
[0005] The patent application of Chongqing University of Technology, "A Cervical Cell Detection Method Based on Multi-Scale Spatial Information" (Patent Application No.: CN202411369428.3), the first step: obtaining a data set of cervical lesion cells, the data set including a test set and a training set; the second step: preprocessing the data set of cervical cell images in the training data set; the third step: constructing a multi-scale spatial information extraction branch module and constructing a channel attention module, and based on the multi-scale spatial information extraction branch module and the constructed channel attention module, constructing an improved Sparse R-CNN network; the fourth step: inputting the preprocessed training set into the improved Sparse R-CNN network and outputting the training result; the fifth step: performing learning training on the Sparse R-CNN detection network, inputting the cervical lesion cell images in the training data set into the Sparse R-CNN network for training, and saving the trained model; the sixth step: using the optimal improved Sparse R-CNN network model to detect cervical lesion cells and retaining the detected result pictures. This method uses Sparse R-CNN for the detection of cervical lesion cells, constructs a multi-scale spatial information extraction branch and constructs a channel attention module. These improvements can enhance the detection ability of Sparse R-CNN for small targets and be applied to the detection of cervical lesion cells. However, from a structural perspective, this method still does not improve the problems of poor feature extraction ability of Sparse R-CNN and long iteration times of the long iteration structure.
[0006] As can be seen from the above method, the existing improvement methods of Sparse R-CNN generally make small-scale improvements to the details of Sparse R-CNN, without touching on the problems of weak feature extraction ability of Sparse R-CNN and excessive parameter quantity and computational complexity caused by too many iteration times of the long iteration structure. Summary of the Invention
[0007] To solve the above-mentioned defects existing in the prior art, the purpose of the present invention is to provide an improved end-to-end object detection system and method for enhancing the feature extraction ability of Sparse R-CNN, optimizing its long iteration structure, and reducing the parameter quantity and computational complexity of Sparse R-CNN while improving the detection accuracy.
[0008] The present invention is realized through the following technical solutions.
[0009] One aspect of the present invention provides an improved end-to-end object detection method, including:
[0010] Construct an improved end-to-end object detection network, including an improved encoder structure and a decoder structure. The improved encoder structure includes a ConvNeXt-Tiny backbone network, a bridge module, and 3 cascaded dynamic heads, which are used for shallow feature extraction, simple feature enhancement, and deep feature extraction of the enhanced features respectively.
[0011] Use the CrowdHuman training set to train the constructed improved end-to-end object detection network to obtain an improved end-to-end object detection model.
[0012] Use the CrowdHuman validation set to evaluate the improved end-to-end object detection model, and adjust the model training parameters according to the training logs and evaluation metrics to obtain an optimized improved end-to-end object detection model.
[0013] Select the improved end-to-end object detection model with the best performance to perform object detection on dense pedestrian images to obtain an object detection result image.
[0014] Preferably, construct an improved end-to-end object detection network, including:
[0015] Construct an improved encoder structure for feature extraction.
[0016] Construct an improved decoder structure for output classification and regression information.
[0017] Preferably, construct an improved encoder structure for feature extraction, including:
[0018] Use the ConvNeXt-Tiny backbone network of the improved encoder for shallow feature extraction.
[0019] Use the bridge module to perform simple feature enhancement on the shallow features output by the ConvNeXt-Tiny backbone network.
[0020] Use 3 cascaded dynamic heads to perform deep feature extraction on the features enhanced by the bridge module.
[0021] Preferably, the ConvNeXt-Tiny backbone network uses 4 stages for feature extraction, each stage contains a fixed number of ConvNeXt Blocks, and a convolutional layer is used for downsampling after each stage.
[0022] Preferably, the bridge module of the improved encoder includes a channel part and a spatial part, and the bridge module calculates channel attention and spatial attention on the shallow features output by the ConvNeXt-Tiny backbone network to achieve simple feature enhancement.
[0023] Preferably, the three cascaded dynamic head modules of the improved encoder perform deep feature extraction on the enhanced features of the bridge module. A dynamic head module includes a scale-aware attention sequence, a spatial-aware attention sequence, and a task-aware attention sequence. When the features output by the bridge module enter a dynamic head, the scale-aware attention sequence fuses different-scale features based on their semantic importance, the spatial-aware attention sequence calculates the discriminative ability focusing on different spatial positions, and the task-aware attention sequence dynamically switches feature channels to assist in the calculation of different tasks.
[0024] Preferably, an improved decoder structure is constructed to output classification and regression information, including performing region of interest feature alignment on the output deep features according to the position information of the candidate boxes. The aligned features will perform attention calculation with the candidate features in the dynamic instance interaction head, and the classification information and regression information are obtained after processing.
[0025] Preferably, the constructed improved end-to-end object detection network is trained using the CrowdHuman training set, including:
[0026] Using the resize() function in the OpenCV library to adjust the size of the input image as the input for model training;
[0027] Determine all training parameters. The class loss function uses the Focal Loss function, and the bounding box loss function uses the hybrid loss function HIoU that combines SIoU and CIoU. The improved end-to-end object detection network is trained using the CrowdHuman training set;
[0028] Record the loss value and AP metric for each iteration during training; obtain the improved end-to-end object detection model.
[0029] In another aspect of the present invention, an improved end-to-end object detection system for the method is provided, including an improved encoder structure and an improved decoder structure;
[0030] The improved encoder structure includes:
[0031] The ConvNeXt-Tiny backbone network for performing shallow feature extraction;
[0032] The bridge module for performing simple feature enhancement on the shallow features output by the ConvNeXt-Tiny backbone network;
[0033] Three cascaded dynamic heads for performing deep feature extraction on the enhanced features of the bridge module;
[0034] The improved decoder structure includes:
[0035] The dynamic instance interaction head calculates the attention for the candidate features and the features aligned with the region of interest, outputs the target features, and decodes the target features to obtain regression information and classification information; among them, the regression information will be used as the candidate bounding box for the next iteration, and the target features will be used as the candidate features for the next iteration.
[0036] Due to the adoption of the above technical solutions, the present invention has the following beneficial effects:
[0037] 1. The present invention adopts the hybrid loss function HIoU, and trains the algorithm by combining two different loss functions, which can enhance the detection accuracy and robustness of the algorithm.
[0038] 2. The ConvNeXt-Tiny backbone network is adopted to extract features using large convolutional kernels, which significantly improves the algorithm's learning ability for long-range dependencies, and while improving the basic feature extraction ability, it also increases the ability to capture context relationships.
[0039] 3. The bridge module is adopted to focus on weighting the channels and spaces of the shallow features, improve the multi-scale detection ability of the algorithm, and use the third-generation deformable convolution in the space part to increase the receptive field.
[0040] 4. Three cascaded dynamic heads are adopted, and the internal shallow features are strongly extracted through a sequence of scale-aware attention, space-aware attention, and task-aware attention, focusing on improving the feature extraction ability of Sparse R-CNN.
[0041] Due to the adoption of the improved encoder structure, the feature extraction ability of Sparse R-CNN is greatly enhanced, and the number of iterations of the long iteration structure is also reduced from 6 times to 4 times. While improving the detection accuracy of the original algorithm, the overall number of parameters and computational complexity are reduced. Description of the Drawings
[0042] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute an improper limitation to the present invention. In the drawings:
[0043] Figure 1 It is the flowchart of the specific implementation manner provided by the present invention;
[0044] Figure 2 It is the structural diagram of an improved end-to-end target detection method provided by the present invention;
[0045] Figure 3 It is the structural diagram of the bridge module provided by the present invention;
[0046] Figure 4 It is the structural diagram of the deformable module provided by the present invention;
[0047] Figure 5 Structural diagram of the dynamic head provided by the present invention;
[0048] Figure 6 Subjective quality comparison chart of the method of the present invention and Sparse R-CNN. Specific implementation manners
[0049] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. Here, the illustrative embodiments of the present invention and the description are used to explain the present invention, but are not intended to limit the present invention.
[0050] Refer to Figure 1 , the embodiment of the present invention provides an improved end-to-end object detection method, and the specific implementation manners include the following steps:
[0051] Step S101, construct an improved end-to-end object detection network, including an improved encoder structure and an improved decoder structure.
[0052] Step S11, construct an improved encoder structure for feature extraction.
[0053] Refer to Figure 2 , the improved encoder uses the ConvNeXt-Tiny backbone network, the bridge module and 3 cascaded dynamic heads to replace the ResNet50 backbone network and the feature pyramid network in Sparse R-CNN.
[0054] Among them, the improved encoder uses the ConvNeXt-Tiny backbone network for shallow feature extraction.
[0055] The ConvNeXt-Tiny backbone network uses 4 stages for feature extraction. Each stage contains a fixed number of ConvNeXt Blocks. Each ConvNeXt Block includes a 7×7 large kernel convolutional layer, a LayerNorm layer, a 1×1 convolutional layer, a GeLU activation function layer, a 1×1 convolutional layer and a residual connection. The first stage contains 3 ConvNeXt Blocks, the second stage contains 3 ConvNeXt Blocks, the third stage contains 9 ConvNeXt Blocks, and the fourth stage contains 3 ConvNeXt Blocks. After each stage, a convolutional layer with a kernel size of 2×2 and a stride of 2 is used for downsampling to reduce the spatial size and increase the number of channels at the same time. ConvNeXt-Tiny finally outputs the shallow features of the image.
[0056] Among them, the improved encoder uses a bridge module to perform simple feature enhancement on the shallow features output by the ConvNeXt-Tiny backbone network to improve the multi-scale detection capability and receptive field.
[0057] like Figure 3 As shown in Figure 2, the bridge module can be divided into a channel part and a spatial part. The bridge module processes the shallow features output by ConvNeXt-Tiny. The processing process can be expressed as follows:
[0058]
[0059] Among them, F is the shallow feature of the input channel part. F1 is the output of the entire channel part, M c represents the channel attention operation, F c It is the feature map output by the channel attention operation. F2 is the output of the entire spatial part and is also the enhanced feature of the final output. s represents the spatial attention operation, F s is the feature map output by the spatial attention operation. Represents element-wise multiplication.
[0060] Reference Figure 3 ,The bridge module processes shallow features including:
[0061] Channel attention processing is performed on shallow features, which can be expressed as follows:
[0062] M c (F) = Conv(ReLU(Conv(Maxpool(F)))) (3) where Maxpool represents the maximum pooling layer, Conv represents the 1×1 convolutional layer, and ReLU represents the ReLU activation function. The essence of this step is to focus on "which channels are more important." That is, the features of different channels may contain different information. For example, one channel may represent edge features, while another may represent color features. The channel part calculates the importance of different channels and assigns higher weights to important channels.
[0063] Perform spatial attention processing on the features output by the channel part. The spatial part is the core of the entire bridge module and plays a vital role in improving the performance of the algorithm. It can be expressed as follows:
[0064] M s (F) = σ(DM(F1)) (4)
[0065] Where σ represents the Sigmoid activation function. DM represents the deformable module. The deformable module is a newly designed small module that integrates deformable convolution. Figure 4 As shown, the DM process can be expressed as:
[0066] DM(F1) = SiLU(DConv(SiLU(BN(Conv(F1))))) (5)
[0067] Where Conv represents a 1×1 convolutional layer, BN represents a batch normalization layer, SiLU represents the SiLU activation function, and DConv is the third-generation deformable convolution. The third-generation deformable convolution shares weights among convolutional neurons, introduces multiple groups of mechanisms, and changes the element-wise Sigmoid function to the softmax function. All these improvements aim to improve efficiency and obtain a larger receptive field. The essence of this step is to focus on "which positions" are more important. It is noted that pixel points at different positions in the feature map may contribute different importance. For example, the edges, contours, or key parts of the target may be more important than the background area. The spatial part assigns higher weights to important positions by calculating the importance at different positions.
[0068] After being processed by the spatial part, the final output of the bridge module is the enhanced feature.
[0069] Among them, the improved encoder uses 3 cascaded dynamic heads to perform deep feature extraction on the enhanced features output by the bridge module.
[0070] The dynamic head is an innovative detection head module proposed by Microsoft and can be used for enhanced feature extraction to obtain deep features. As Figure 5 shown, a dynamic head module includes a scale-aware attention sequence, a space-aware attention sequence, and a task-aware attention sequence. When the features output by the bridge module enter a dynamic head, they all go through the calculations of these three attention sequences. This process can be expressed as:
[0071] W(F) = π C (π S )(π L (F)·F)·F)·F (6)
[0072] Where π C 、π S and π L are the three attention sequences corresponding to the scale-aware attention, space-aware attention, and task-aware attention respectively, and F is the feature output by the bridge module.
[0073] The scale-aware attention can fuse different-scale features based on their semantic importance. The process is as follows:
[0074] π L (F)·F = σ(f(Avgpool(F)))·F (7)
[0075] Among them, F is the feature output by the bridge module. f(·) is a linear function and can be replaced by a 1×1 convolution. Avgpool is an average pooling operation, and σ(·) is a hard-sigmoid activation function.
[0076] Spatial-aware attention can focus on the discriminative ability of different spatial positions. The process is as follows:
[0077]
[0078] Among them, F is the feature output by scale-aware attention. K is the number of sparse sampling positions, and k is the enumerated sampling point. F(l; p k +Δp k ; c) is a deformable convolution operation, specifically representing the value of the feature F at the position p k +Δp k in the c-th channel of the l-th layer feature. p k is the k-th position predefined for network sampling in the regular convolution, and Δp k is the offset corresponding to the k-th network sampling position. p k +Δp k represents the offset position after self-learned spatial offset Δp k to focus on the discriminative region. Δm k is the modulation scalar at the position Δp k . Both are learned from the intermediate layer input features of F.
[0079] Task-aware attention can dynamically switch feature channels to assist different tasks. The process can be characterized as:
[0080] π C (F)·F = max(α 1 (F)·F c +β 1 (F), α 2 (F)·F c +β 2 (F)) (9) Among them, F is the feature output by spatial-aware attention. [α 1 , α 2 , β 1 , β 2 T = θ(·) are hyperparameters used to control the activation threshold. And θ(·) is similar to Dynamic ReLU. Dynamic ReLU first performs global average pooling in the L×S dimension for dimensionality reduction, then uses two fully connected layers and a normalization layer, and finally applies a shifted sigmoid function to normalize the output to [-1,1].
[0081] After three cascaded dynamic heads, the deep features of the image are obtained, and the improved encoder is now constructed.
[0082] Step S12: construct an improved decoder structure to output classification and regression information.
[0083] like Figure 2 As shown, the decoder part includes a dynamic instance interaction head, which is the long iterative structure of the method of the present invention, and the number of iterations is 4. The method of the present invention presets 500 groups of queries to represent the candidate boxes and candidate features respectively. In the decoder, the deep features output in step S11 are aligned with the region of interest features according to the position information of the candidate boxes. After that, this part of the features will be calculated with the candidate features in the dynamic instance interaction head, and the result of the calculation is the target feature. After processing the target feature, classification information and regression information will be obtained, among which the regression information will be used as a new candidate box and the target feature will be used as a new candidate feature. They will participate in the next iteration until the fourth iteration.
[0084] The method of the present invention significantly enhances feature extraction capabilities due to its improved encoder structure. Consequently, the number of iterations in the decoder's long iterative structure can be reduced from six in Sparse R-CNN to four. Optimizing the number of iterations in the decoder's long iterative structure significantly reduces the number of parameters and computational complexity, greatly improving computational efficiency and reducing redundant parameters. This results in an improved end-to-end object detection network.
[0085] Step S102: prepare the CrowdHuman dataset and use the CrowdHuman training set to train the constructed improved end-to-end object detection network to obtain an improved end-to-end object detection model.
[0086] The specific steps include:
[0087] Step S21: Prepare the CrowdHuman dataset.
[0088] The CrowdHuman dataset is a public, dense pedestrian detection dataset consisting of 15,000 images in the training set and 4,370 images in the validation set. Place the training and validation images in the same folder and organize the labels for each image.
[0089] Step S22: Use the CrowdHuman training set to train the improved end-to-end object detection network.
[0090] (a) The resize() function in the OpenCV library is used to resize the input image to 1400×800 as the input for model training.
[0091] (b) Determine all training parameters, load the pre-trained weights corresponding to the backbone network, select AdamW as the optimizer, adjust the learning rate to 0.0001, set the weight decay to 0.05, and set the Batch Size to 4.
[0092] (c) Use Focal Loss for the class loss function and the hybrid loss function HIoU for the bounding box loss function. HIoU combines the SIoU and CIoU loss functions to train an improved end-to-end object detection network. The formula for the HIoU loss function is:
[0093]
[0094] where, L HIoU represents the loss value of HIoU; t represents time, and T mid is the moment when training reaches half, and T total represents the complete time required for training; L SIoU means that when the training time is less than or equal to T mid , that is, in the first half of training, SIoU is used as the loss function; L CIoU means that when the training time is greater than T mid and less than or equal to T total , that is, in the second half of training, CIoU is used as the loss function.
[0095] In HIoU, the granularity of time t is the number of training epochs. A total of 36 epochs are required to train this method. Therefore, T mid is 18, and T total is 36.
[0096] Use the CrowdHuman training set to train the improved end-to-end object detection network.
[0097] Step S23, record the training log, and the log content is the loss value and AP metric for each iteration during training.
[0098] Step S24, complete the training to obtain the improved end-to-end object detection model.
[0099] Step S103, evaluate the improved end-to-end object detection model using the CrowdHuman validation set, adjust the model training parameters according to the training log and evaluation metrics to obtain an optimized model;
[0100] Step S104, select the improved end-to-end object detection model with the best performance to perform object detection on dense pedestrian images to obtain the object detection result image.
[0101] To better illustrate the effectiveness of the present invention, experiments are conducted on the CrowdHuman dataset and the MS COCO dataset, and compared with multiple object detection methods.
[0102] CrowdHuman is a large-scale pedestrian detection dataset designed specifically for pedestrian detection tasks in crowded scenes and is widely used in object detection and pedestrian re-identification research. The dataset consists of 15,000 high-resolution images (including 11,800 training images and 4,370 validation images), containing approximately 343,000 pedestrian instances in total. The annotation information includes full body bounding boxes, visible body parts, and head bounding boxes, providing more refined object information and being suitable for detection tasks at different granularities. The images in CrowdHuman come from a wide range of sources, including street surveillance, public places, stadiums, etc., with characteristics of high occlusion and dense distribution, making it an important benchmark for evaluating the robustness of object detection.
[0103] The comparison results of different methods on the CrowdHuman dataset are shown in Table 1.
[0104]
[0105]
[0106] As can be seen from the table, the AP value of the method of the present invention on the CrowdHuman data is 92.7%, with the best performance. The method of the present invention is 4.1% better than Sparse R-CNN and 0.7% better than Progressive Sparse R-CNN in terms of the AP value. In addition, the number of parameters of the method of the present invention decreases by 15.3% compared with Sparse R-CNN, and the computational amount decreases by 17%.
[0107] MS COCO (Microsoft Common Objects in Context) is a versatile dataset widely used in computer vision tasks, covering tasks such as object detection, instance segmentation, keypoint detection, panoramic segmentation, and caption generation. It was released by the Microsoft team and contains 330,000 images (118,000 training images and 5,000 validation images are used for object detection), and approximately 200,000 of them have detailed pixel-level annotations. The COCO dataset contains 80 classes of objects, covering common items in daily life, such as people, animals, vehicles, furniture, etc. Each image usually contains multiple objects, with complex backgrounds and various object interaction relationships, conforming to real-world visual scenes. In addition, COCO also provides 5 types of human keypoint annotations (such as hands, feet, nose, etc.), as well as text descriptions for the Captions task, supporting multimodal research. The COCO Challenge is an important competition in the field of computer vision, which has promoted the development of advanced algorithms such as Faster R-CNN, Mask R-CNN, YOLO, and DETR. This dataset is widely used in tasks such as object detection, instance segmentation, and pose estimation, and is one of the core benchmarks for deep learning research.
[0108] The comparison results of different methods on the MS COCO dataset are shown in Table 2.
[0109]
[0110]
[0111] As can be seen from the table, the method of the present invention ranks first in all AP metrics. With a 15.3% decrease in the number of parameters and a 17% decrease in the computational amount, AP 50:95 increases by 2.7%, AP 50 increases by 2.4%, AP 75 increases by 3.5%, AP S increases by 2.5%, AP M increases by 2.5%, and AP L increases by 1.7%.
[0112] The experimental results in Table 1 and Table 2 fully illustrate that the feature extraction ability of the method of the present invention is greatly enhanced. Because the number of iterations of the long iterative structure is optimized, the number of parameters and the computational amount are also reduced compared to Sparse R-CNN, and the detection accuracy is also more excellent.
[0113] Figure 6 is the subjective quality comparison chart of the method of the present invention and Sparse R-CNN, where Figure 6 (a) shows the detection results of Sparse R-CNN,Figure 6 (b) is the detection result of the method of the present invention. Figure 6 The first row of images in shows a blind basketball game. The sizes of the targets in this image are significantly different. In particular, the sizes of the three people on the right side of the image are larger, while the sizes of the two people far from the court in the middle of the image are extremely small. In addition, there is partial overlap between the two people on the left side of the image, and the overall detection difficulty is relatively high. It can be seen from the detection results that Sparse R-CNN has obvious misdetections. There are only two people on the left side of the image, but the annotation results of Sparse R-CNN show three bounding boxes, while the method of the present invention accurately identifies these two people. It should be noted that the method of the present invention not only correctly detects all the targets in the image, but also the confidence score of each target is higher than that of Sparse R-CNN. This not only verifies that the method of the present invention has stronger feature extraction ability, but also subjectively affirms the multi-scale detection ability of the method of the present invention; Figure 6 The second row of images in shows a close-up group photo of multiple people. There are a total of five people in the image, and there is more or less overlap between each person. It can be seen from the detection results that Sparse R-CNN detected 6 targets, and there are again misdetections. This can mutually testify with the argument about Sparse R-CNN proposed in the present invention: that is, the design concept of Sparse R-CNN is novel, but the structure is obsolete and the feature extraction ability is weak. For the same image, after being inferred by the method of the present invention, 5 people are accurately marked, and the overall confidence score is also higher; Figure 6 The third row of images in is a large-format image with high resolution. There are many small and dense targets in this image (the crowd in the image background), uneven illumination (the crowd far from the camera is in the shadow of the building), and there are targets with extremely large scale differences (the size difference between the lady closest to the camera and the crowd in the distance). It is a picture with the superposition of multiple complex scene features, and the detection difficulty is extremely high. It can be seen from the detection results that although Sparse R-CNN recognized the lady closest to the camera, it missed the man in white clothes on the left side of the image. The method of the present invention not only recognized many small targets in the distance of the image, but also accurately recognized the man in white clothes on the left side of the image and the lady in front of the camera, fully verifying the boosting effect of the bridge module on the multi-scale detection ability of the method of the present invention.
[0114] Referring to Figure 2 As shown, the embodiment of the present invention provides an improved end-to-end target detection system, including an improved encoder structure and an improved decoder structure:
[0115] The improved encoder structure includes:
[0116] The ConvNeXt-Tiny backbone network is used for shallow feature extraction;
[0117] A bridge module for simply enhancing the shallow features output by ConvNeXt-Tiny;
[0118] Three cascaded dynamic heads for extracting deep features from the features enhanced by the bridge module.
[0119] An improved decoder structure, including:
[0120] A dynamic instance interaction head that calculates attention on candidate features and features aligned with regions of interest, outputs target features, and decodes the target features to obtain regression information and classification information; among them, the regression information will be used as the candidate box for the next iteration, and the target features will be used as the candidate features for the next iteration.
[0121] The present invention is not limited to the above embodiments. Based on the technical solutions disclosed in the present invention, those skilled in the art can make some substitutions and deformations to some of the technical features without creative labor according to the disclosed technical content, and these substitutions and deformations are all within the protection scope of the present invention.
Claims
1. An improved end-to-end object detection method, characterized in that, Including: Construct an improved end-to-end object detection network, including an improved encoder structure and a decoder structure. The improved encoder structure includes a ConvNeXt-Tiny backbone network, a bridge module, and 3 cascaded dynamic heads, which are used for shallow feature extraction, simple feature enhancement, and deep feature extraction of the enhanced features respectively; Use the CrowdHuman training set to train the constructed improved end-to-end object detection network to obtain an improved end-to-end object detection model; Use the CrowdHuman validation set to evaluate the improved end-to-end object detection model, and adjust the model training parameters according to the training logs and evaluation metrics to obtain an optimized improved end-to-end object detection model; Select the improved end-to-end object detection model with the best performance to perform object detection on dense pedestrian images to obtain an object detection result image.
2. The improved end-to-end object detection method according to claim 1, characterized in that, Construct an improved end-to-end object detection network, including: Construct an improved encoder structure for feature extraction; Construct an improved decoder structure for output classification and regression information.
3. The improved end-to-end object detection method according to claim 2, wherein Construct an improved encoder structure for feature extraction, including: Adopt the ConvNeXt-Tiny backbone network of the improved encoder for shallow feature extraction; Adopt a bridge module to perform simple feature enhancement on the shallow features output by the ConvNeXt-Tiny backbone network; Adopt 3 cascaded dynamic heads to perform deep feature extraction on the features enhanced by the bridge module.
4. The improved end-to-end object detection method according to claim 3, characterized in that, The ConvNeXt-Tiny backbone network uses 4 stages for feature extraction. Each stage contains a fixed number of ConvNeXt Blocks, and a convolutional layer is used for downsampling after each stage.
5. The improved end-to-end object detection method according to claim 3, wherein, The bridge module of the improved encoder includes a channel part and a spatial part. The bridge module performs channel attention and spatial attention calculations on the shallow features output by the ConvNeXt-Tiny backbone network to achieve simple feature enhancement.
6. The improved end-to-end object detection method according to claim 3, wherein The 3 cascaded dynamic head modules of the improved encoder perform deep feature extraction on the features enhanced by the bridge module. A dynamic head module includes a scale-aware attention sequence, a spatial-aware attention sequence, and a task-aware attention sequence. When the features output by the bridge module enter a dynamic head, they are fused by the scale-aware attention sequence based on their semantic importance for different-scale features, the discriminative ability calculation of different spatial positions is focused by the spatial-aware attention sequence, and the feature channels are dynamically switched by the task-aware attention sequence to assist the calculation of different tasks.
7. The improved end-to-end object detection method according to claim 2, wherein Construct an improved decoder structure for output classification and regression information, including performing region of interest feature alignment on the output deep features according to the position information of the candidate boxes. The aligned features will perform attention calculation with the candidate features in the dynamic instance interaction head, and the classification information and regression information are obtained after processing.
8. The improved end-to-end object detection method according to claim 1, wherein, Use the CrowdHuman training set to train the constructed improved end-to-end object detection network, including: Use the resize() function in the OpenCV library to adjust the size of the input image as the input for model training; Determine all training parameters. Use the Focal Loss function as the class loss function and the hybrid loss function HIoU, which combines SIoU and CIoU, as the bounding box loss function. Train the improved end-to-end object detection network using the CrowdHuman training set; Record the loss value and AP metric for each iteration during training; obtain the improved end-to-end object detection model.
9. The improved end-to-end object detection method according to claim 8, wherein, The formula for the hybrid loss function HIoU is: Among them, L HIoU represents the loss value of HIoU; t represents time, and T mid is the moment when the training reaches the midpoint, and T total represents the complete time required for training; L SIoU represents that when the training time is less than or equal to T mid , that is, in the first half of the training, SIoU is used as the loss function; L CIoU represents that when the training time is greater than T mid and less than or equal to T total , that is, in the second half of the training, CIoU is used as the loss function.
10. An improved end-to-end object detection system for the method according to any one of claims 1-9, characterized in that, It includes an improved encoder structure and an improved decoder structure; The improved encoder structure includes: The ConvNeXt-Tiny backbone network for shallow feature extraction; The bridge module for simple feature enhancement of the shallow features output by the ConvNeXt-Tiny backbone network; Three cascaded dynamic heads for deep feature extraction of the features enhanced by the bridge module; The improved decoder structure includes: The dynamic instance interaction head, which calculates the attention on the candidate features and the features aligned with the regions of interest, outputs the target features, and decodes the target features to obtain the regression information and classification information; among them, the regression information will be used as the candidate bounding box for the next iteration, and the target features will be used as the candidate features for the next iteration.
Citation Information
Patent Citations
An object detection method based on improved Sparse R-CNN
CN116385732B
Cervical cell detection method based on multi-scale spatial information
CN119359638A