Target detection method and system based on sparse region extraction

The sparse area extraction module performs sparse area screening on high-resolution feature layers and combines with the object detection head module to detect it, solving the problem of large number of candidate boxes and large amount of calculations in high-resolution layer, and achieving efficient and real-time object detection.

CN119625288BActive Publication Date: 2025-05-16HANGZHOU DIANZI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510155496.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-16
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing object detection methods have a large number of candidate boxes and a large amount of calculations on high-resolution layers, which affects real-time performance and is difficult to take into account both accuracy and speed.

Method used

The object detection method based on sparse area extraction is adopted, and the high-resolution feature layer is filtered through the sparse area extraction module to reduce redundant calculations, and the detection of small targets and global targets is carried out in combination with the first and second target detection head modules.

Benefits of technology

It effectively reduces the number of candidate areas and redundant calculations on high-resolution feature layers, improves detection rate and accuracy, and improves overall detection efficiency and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625288B_ABST
    Figure CN119625288B_ABST
Patent Text Reader

Abstract

The present invention discloses a target detection method and system based on sparse region extraction. The present invention first extracts the features of the image to be detected and performs fusion enhancement to obtain enhanced features of different resolutions; then, for the high-resolution features, the feature window is screened out through the sparse region extraction module; then, a first priori frame is generated based on the feature window, and its category and offset are predicted through the first target detection head module; at the same time, a second priori frame is generated for the remaining enhanced features, and is predicted through the second target detection head module; finally, all the priori frames are added to the candidate anchor frame set after being offset according to the offset, and the final target detection frame is obtained through non-maximum suppression. The present invention performs sparse region extraction on the high-resolution features, selects the key parts of the area for small-sized target detection, and performs normal global target detection on the remaining low-resolution features, thereby improving the detection efficiency while improving the accuracy of small-sized target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to computer vision and image processing technology, and in particular to a target detection method based on sparse region extraction. Background Art

[0002] In the field of computer vision, target detection is a key technology and is widely used in scenarios such as autonomous driving and security monitoring. Traditional target detection methods rely on multi-level feature extraction to achieve target recognition and positioning through high- and low-resolution feature maps. However, although high-resolution feature maps can capture details and improve the ability to detect small targets, they lead to too many candidate areas and large amounts of calculation, which affects real-time performance. Low-resolution feature maps have high computational efficiency, but they perform poorly in detecting small targets and it is difficult to balance accuracy and speed. Therefore, how to optimize the computational burden of high-resolution layers while ensuring detection accuracy and improving overall detection efficiency has become the focus of current research.

[0003] RetinaNet is a single-stage object detection method proposed by Facebook AI Research. Its core innovation is the focal loss, which aims to solve the problem of class imbalance in traditional object detection methods. The traditional cross entropy loss function tends to ignore difficult-to-detect targets when facing class imbalance. The focal loss reduces the weight of easy-to-classify samples and strengthens the training of difficult-to-classify samples, thereby improving the model's recognition ability for small objects and difficult-to-detect targets. RetinaNet combines convolutional neural networks with an end-to-end fully convolutional structure to avoid the computational bottleneck of traditional two-stage detection methods and improve detection efficiency. However, RetinaNet still has some shortcomings. Although the focal loss has advantages in dealing with class imbalance, the detection accuracy may be affected when facing multi-scale, dense targets and complex scenes. In addition, RetinaNet has a large computational overhead, especially when processing high-resolution images, the number of generated prior boxes is exponentially increasing relative to low-resolution, which may still lead to inefficiency. Although it performs well in many tasks, it is still a challenge to further optimize computational efficiency and accuracy when processing large-scale datasets.

[0004] Chenhongyi Yang et al. proposed a cascaded sparse query target detection method in their paper (QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection), which aims to solve the computational efficiency and accuracy problems in high-resolution small object detection. Unlike traditional methods, QueryDet reduces redundant calculations through a sparse query mechanism, while using a cascade structure to gradually improve the positioning accuracy of small objects. This method significantly improves the detection efficiency by focusing on the most relevant areas in the image. However, the model only participates in the loss calculation during training, and does not participate in the selection calculation of the prior box. Therefore, the number of prior boxes does not decrease. Only during inference, the sparse query mechanism is used to select the prior box. Therefore, despite its excellent performance in small object detection, further optimization is still needed to cope with more complex scenarios. Summary of the invention

[0005] Based on the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a target detection method based on sparse region extraction, which mainly solves the problems of a large number of candidate boxes and a large amount of calculation in high-resolution layers by performing sparse region extraction on high-resolution layers, while also improving the detection rate.

[0006] The present invention adopts the following technical solutions to solve the above problems:

[0007] In a first aspect, the present invention provides a target detection method based on sparse region extraction, which comprises:

[0008] S1. For the input image to be detected, the feature extraction module is used to extract image features from it, and then the image features are fused and enhanced by the feature fusion module to obtain enhanced features with different resolutions, and the two layers of enhanced features with the highest resolution are used as the first features, and the remaining enhanced features are used as the second features;

[0009] S2. For each layer of the first feature, input it into the sparse region extraction module, and after passing through the patch embedding layer, the window multi-head self-attention layer, the moving window multi-head self-attention layer, and the layer normalization layer in the module, the number of channels is compressed to 1 through the convolution layer, and then the compressed single-channel feature is divided into a series of windows, and the weighted sum of the maximum value and the average value of the pixels in each window is calculated and used as the probability value of the window being selected, and a part of the windows with the largest probability value are screened out from all windows as feature windows;

[0010] S3, for each layer of the first feature, generate multiple first a priori boxes based on each feature point in all feature windows, splice the local feature blocks of all feature windows in the first feature and input them into the first object detection head module, and obtain the predicted category and offset of each first a priori box through the first category prediction branch and the first box prediction branch respectively;

[0011] S4. For each layer of the second feature, generate multiple second a priori boxes based on each feature point in the second feature, and directly input the second feature into the second object detection head module, and obtain the predicted category and offset of each second a priori box through the second category prediction branch and the second box prediction branch respectively;

[0012] S5. All first a priori boxes and all second a priori boxes are offset according to their respective offsets and added to the candidate anchor box set. Then, the redundant anchor boxes are removed through the non-maximum suppression module, and finally the detection boxes of all targets in the image to be detected are obtained.

[0013] As a preferred embodiment of the first aspect, the feature extraction module adopts a Resnet network.

[0014] As a preferred embodiment of the first aspect, the feature fusion module adopts a feature pyramid model (FPN), and the final output enhanced features have a total of 6 layers, the two layers of enhanced features with the highest resolution are used as the first features, and the remaining 4 layers of enhanced features are used as the second features.

[0015] As a preferred embodiment of the above-mentioned first aspect, in the sparse area extraction module, the first category prediction branch and the first box prediction branch have the same structure, both of which are cascaded by 2 layers of window multi-head self-attention layers, layer normalization layers and 3×3 convolution layers.

[0016] As a preferred embodiment of the first aspect above, in the second target detection head module, the second category prediction branch and the second box prediction branch have the same structure, both of which are formed by cascading multiple layers of 3×3 convolutional layers.

[0017] As a preferred embodiment of the above-mentioned first aspect, the target detection model composed of the feature extraction module, the feature fusion module, the sparse region extraction module, the first target detection head module, the second target detection head module and the non-maximum suppression module needs to be supervised trained on a labeled training data set in advance.

[0018] As a preferred embodiment of the first aspect above, the loss function used for the supervised training is a weighted sum of query loss, category prediction loss, and position regression loss; the query loss is a focal loss (Focal Loss) between the Sigmoid activation output of the single-channel feature and the true value labeled feature map, wherein in the true value labeled feature map, only the anchor frame area where the target whose size is smaller than a preset value is located is marked as 1, and the remaining positions are marked as 0; the category prediction loss is a focal loss (Focal Loss) between the predicted category and the true value category of all detection boxes; the position regression loss is an L1Loss loss between the predicted offset and the true offset of all detection boxes.

[0019] In a second aspect, the present invention provides a target detection system based on sparse region extraction, comprising:

[0020] An image input module is used to input an image to be detected that needs to be detected;

[0021] a detection module, configured to perform target detection on the image to be detected inputted in the image input module according to the target detection method based on sparse region extraction as described in any one of the schemes of the first aspect above, and obtain detection frames of all targets in the image to be detected and a predicted category corresponding to each detection frame;

[0022] The output module is used to output the detection box and the predicted category obtained in the detection module according to a preset output method.

[0023] In a third aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the target detection method based on sparse region extraction as described in any of the schemes of the first aspect.

[0024] In a fourth aspect, the present invention provides a computer electronic device comprising a memory and a processor;

[0025] The memory is used to store computer programs;

[0026] The processor is used to implement the target detection method based on sparse area extraction as described in any solution of the first aspect when executing the computer program.

[0027] Compared with the prior art, the present invention adopts the above technical solution and has the following benefits:

[0028] 1. The present invention realizes the sparse area extraction of the high-resolution feature layer, and can focus on the small target detection in the feature window, effectively reducing the number of candidate areas and redundant calculations on the high-resolution feature layer. Therefore, while ensuring that the high-resolution layer retains sufficient detail information to improve the small target detection accuracy, the present invention significantly reduces the calculation burden of the overall detection and improves the detection rate.

[0029] 2. The present invention uses a first target detection head module for detecting small targets to efficiently predict small targets on high-resolution feature blocks, and directly uses a second target detection head module in the form of a convolution detection head to quickly infer categories and target boxes on low-resolution feature layers. This allows the present invention to significantly improve real-time performance and efficiency while maintaining or improving detection accuracy, and has broad application prospects and obvious technical advantages in a variety of scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of the steps of a target detection method based on sparse region extraction;

[0031] Figure 2 This is a schematic diagram of the module composition of the target detection model;

[0032] Figure 3 A schematic diagram of the module composition of a target detection system based on sparse region extraction;

[0033] Figure 4 A schematic diagram of the components of computer electronic equipment;

[0034] Figure 5 A schematic diagram of a method flow in an embodiment of the present invention;

[0035] Figure 6 It is a schematic diagram of the internal process of the sparse region extraction module;

[0036] Figure 7 Schematic diagram of the internal process of the first target detection head module in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.

[0038] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features.

[0039] like Figure 1 As shown, in a preferred embodiment of the present invention, a target detection method based on sparse region extraction is provided, which includes 5 steps from S1 to S5. The key is to perform sparse region extraction on the high-resolution features extracted from the image, screen out key areas to focus on the detection of small-sized targets (of course large-sized and medium-sized targets will also be detected), and perform normal global target detection on the remaining low-resolution features, thereby improving the detection accuracy of small-sized targets, while solving the problems of a large number of high-resolution layer candidate boxes and a large amount of calculation, and improving the detection efficiency.

[0040] S1. For the input image to be detected, the feature extraction module is used to extract image features from it, and then the image features are fused and enhanced through the feature fusion module to obtain enhanced features of different resolutions, and the two layers of enhanced features with the highest resolution are used as the first features, and the remaining enhanced features are used as the second features.

[0041] It should be noted that the feature extraction module and feature fusion module in the present invention can theoretically be implemented by using corresponding network modules that can extract multi-scale features. In an embodiment of the present invention, the feature extraction module adopts a Resnet network, such as Resnet-50 or Resnet-100. In addition, the feature fusion module can adopt a feature pyramid model (FPN), and the features extracted by the feature extraction module can be further input into the FPN. Based on the top-down feature fusion method inside the FPN, high-level features (rich in semantic information but low resolution) are organically combined with low-level features (high resolution but less semantic information), thereby obtaining multi-scale features with rich semantic information at different levels. In an embodiment of the present invention, the feature extraction module and the feature fusion module preferably adopt a combination of Resnet-50+FPN, and the enhanced features finally output by the FPN have a total of 6 layers, and the two layers of enhanced features with the highest resolution are used as the first features, and the remaining 4 layers of enhanced features are used as the second features. These 2 layers of first features and 4 layers of second features will be used as inputs for subsequent modules.

[0042] S2. For each layer of the first feature, it is input into the sparse region extraction module. After passing through the patch embedding layer, the window multi-head self-attention layer (WindowMSA), the moving window multi-head self-attention layer (Shifted Window MSA), and the layer normalization (LayerNorm) layer in the module, the number of channels is compressed to 1 through the convolution layer, and then the compressed single-channel feature is divided into a series of windows. The weighted sum of the maximum value and the average value of the pixels in each window is calculated and used as the probability value of the window being selected, and a part of the windows with the largest probability value are screened out from all windows as feature windows.

[0043] It should be noted that the Patch Embedding layer, Window MSA layer, Shifted Window MSA layer, and LayerNorm layer are all functional modules in the prior art and can be implemented according to their respective prior art practices. The convolution layer connected after the LayerNorm layer is used to compress the number of channels to 1, which can be implemented by 3×3 convolution in the embodiment of the present invention.

[0044] In addition, for the compressed single-channel features output by the 3×3 convolution, it is necessary to divide it into a series of windows, and the size of each window can be reasonably adjusted according to the actual situation. In the embodiment of the present invention, the window size is set to 8×8. After the window division is completed, the probability value of each window being selected as the feature window needs to be calculated. The probability value It can be calculated as follows:

[0045]

[0046] Where: represents the maximum eigenvalue within the calculation window, represents the average eigenvalue within the calculation window, and Represents the corresponding weight hyperparameter.

[0047] After calculating the probability value of each window being selected as a feature window, all windows can be sorted from high to low according to the probability value, and then the windows with the highest probability value are selected as feature windows according to the preset ratio or number of windows. These feature windows will be used as key areas for detecting small-sized objects (of course, large-sized and medium-sized objects will also be detected, but the main purpose of this part is to detect more small-sized objects), so that some key areas can be sparsely extracted from the entire global area.

[0048] S3. For each layer of the first feature, multiple first prior frames are generated based on each feature point in all feature windows, and the local feature blocks of all feature windows in the first feature are spliced ​​and input into the first target detection head module, and the predicted category and offset of each first prior frame are respectively obtained through the first category prediction branch and the first frame prediction branch.

[0049] It should be noted that each layer of the first feature needs to perform the above step S3 to obtain the predicted categories and offsets of all first a priori boxes in the first feature of this layer. The first a priori box can be offset according to the offset of the first a priori box to obtain the final detection anchor box.

[0050] It should also be noted that when generating the first priori box based on each feature point in all feature windows, the specific number of first priori boxes generated is a hyperparameter that can be optimized and adjusted. In an embodiment of the present invention, the generation of the first priori box can be achieved by directly calling AnchorGenerator (anchor generator), and 9 first priori boxes are preferably generated for each feature point, so that for an 8*8 feature window, 8*8*9 first priori boxes can be generated. If a total of M feature windows are extracted in the first feature of each layer, the number of first priori boxes generated in the first feature is M*8*8*9.

[0051] In an embodiment of the present invention, the above-mentioned sparse region extraction module includes two parallel branches, namely the first category prediction branch and the first box prediction branch. The first category prediction branch can predict the target category contained in each first priori box, and the first box prediction branch can predict the offset of each first priori box. In an embodiment of the present invention, the first category prediction branch and the first box prediction branch have the same structure, both of which are cascaded by 2 layers of window multi-head self-attention (WindowMSA) layers, layer normalization (LayerNorm) layers and 3×3 convolution layers. The input of each branch is the concatenation result of all local feature blocks in a single-layer first feature. The number of output channels of the last 3×3 convolution layer in the two branches is different. The number of output channels of the 3×3 convolution layer in the first category prediction branch is the number of first priori boxes for each feature point multiplied by the number of target category labels, and the number of output channels of the 3×3 convolution layer in the first box prediction branch is 4 times the number of first priori boxes for each feature point in the first feature. Here, 4 times represents the 4 dimensional values ​​of the offset.

[0052] S4. For each layer of the second feature, multiple second prior boxes are generated based on each feature point in the second feature, and the second feature is directly input into the second target detection head module, and the predicted category and offset of each second prior box are respectively obtained through the second category prediction branch and the second box prediction branch.

[0053] In the embodiment of the present invention, when the second prior frame is generated based on each feature point in the second feature, the specific number of second prior frames generated is a hyperparameter that can be optimized and adjusted. In the embodiment of the present invention, the generation of the second prior frame can also be achieved by directly calling AnchorGenerator (anchor generator), and each feature point preferably generates 9 second prior frames, so that for one The feature window can be generated A second prior box.

[0054] In an embodiment of the present invention, the second target detection head module includes two parallel branches, namely a second category prediction branch and a second box prediction branch. The second category prediction branch can predict the target category contained in each second prior box, and the second box prediction branch can predict the offset of each second prior box. In an embodiment of the present invention, the second category prediction branch and the second box prediction branch have the same structure, both of which are cascaded by multiple layers of 3×3 convolutional layers, and the specific number of 3×3 convolutional layers is preferably 5 layers. Similarly, the number of output channels of the last 3×3 convolutional layer in the second category prediction branch and the second box prediction branch is different. The number of output channels of the 3×3 convolutional layer in the second category prediction branch is the number of second prior boxes for each feature point in the second feature multiplied by the number of target category labels, and the number of output channels of the 3×3 convolutional layer in the second box prediction branch is 4 times the number of second prior boxes for each feature point in the second feature, where 4 times also represents the 4-dimensional values ​​of the offset.

[0055] S5. All first a priori boxes and all second a priori boxes are offset according to their respective offsets and added to the candidate anchor box set. Then, the redundant anchor boxes are removed through the non-maximum suppression module, and finally the detection boxes of all targets in the image to be detected are obtained.

[0056] It should be noted that the non-maximum suppression (NMS) operation performed in the non-maximum suppression module is an existing technology in the field of target detection, which aims to solve the problem of overlapping bounding boxes that occur during the detection process, thereby obtaining more accurate target detection results. The core concept of NMS is that when there are multiple overlapping bounding boxes, only the bounding box with the highest confidence is retained, and other bounding boxes with lower confidence are suppressed. Since NMS belongs to the existing technology, its specific principle is not repeated here, and the existing encapsulation module can be directly called to perform this function. After all the candidate anchor boxes in the candidate anchor box set pass through the NMS module, the retained candidate anchor boxes are the target detection boxes in the final image to be detected, and since the prediction categories of these detection boxes have been determined in the aforementioned S3 and S4 steps, the coordinate position of the detection box and the target category can be output at the same time.

[0057] In addition, it should be noted that the target detection method based on sparse region extraction described in S1 to S5 above uses a feature extraction module, a feature fusion module, a sparse region extraction module, a first target detection head module, a second target detection head module and a non-maximum suppression module. Figure 2 As shown, these modules constitute the framework of the target detection model of the present invention. Based on the conventional model training process, it can be seen that the target detection model needs to be supervised in advance on a labeled training data set before being used for actual reasoning and prediction. The loss function used in supervised training can be designed according to the loss of the conventional target detection algorithm, as long as the detection accuracy of the model can be guaranteed.

[0058] In an embodiment of the present invention, a loss function for supervised training of the above target detection model is provided, and its form is optimized as a weighted sum of three parts: query loss, category prediction loss, and position regression loss. The specific forms of these three parts are described below:

[0059] The first part of the loss is the query loss, which is the focal loss between the Sigmoid activation output of the above single-channel feature (output by the convolution layer in the sparse region extraction module) and the true value annotation feature map, where only the anchor box area where the target with a size smaller than the preset value is located in the true value annotation feature map is marked as 1, and the rest of the positions are marked as 0. It should be noted that the size of the target marked as 1 in the original image to be detected needs to be smaller than the preset value. This preset value is a hyperparameter used to control the size of small-sized targets for key detection.

[0060] The second part of the loss is the category prediction loss, which is the focal loss between the predicted category and the true value category of all detection boxes. The specific form of Focal Loss belongs to the existing technology and will not be repeated here.

[0061] The third part of the loss is the position regression loss, which is the L1Loss loss between the predicted offset and the actual offset of all detection boxes. The specific form of the L1Loss loss belongs to the prior art and will not be repeated here.

[0062] The weighted sum of the query loss, category prediction loss, and position regression loss is the total loss function used by the target detection model. The weight of each part can be an optimization parameter. After the total loss function is set, the labeled training data set can be used to extract samples in batches to iteratively train the target detection model. When the accuracy on the validation set meets the requirements, the iteration can be stopped and the final model parameters can be output for reasoning and actual detection.

[0063] It should be noted that the target detection method steps based on sparse region extraction shown in the above S1 to S5 can essentially be implemented in the form of a computer program.

[0064] In addition, based on the same inventive concept, Figure 3 As shown, the present invention provides a target detection system based on sparse region extraction, which includes:

[0065] An image input module is used to input an image to be detected that needs to be detected;

[0066] A detection module, configured to perform target detection on the image to be detected inputted from the image input module according to the target detection method based on sparse region extraction shown in S1 to S5 above, and obtain detection frames of all targets in the image to be detected and a predicted category corresponding to each detection frame;

[0067] The output module is used to output the detection box and the predicted category obtained in the detection module according to a preset output method.

[0068] It should be noted that the above-mentioned image input module and output module can be implemented through a GUI interface or code statements. It is preferred to design a GUI interface to implement an image input function that can be intuitively operated by the user, and at the same time intuitively display the detection box and predicted category on the image to be detected through the GUI interface.

[0069] Similarly, based on the same inventive concept, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the target detection method based on sparse region extraction as described above.

[0070] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.

[0071] Therefore, based on the same inventive concept, the present invention provides a computer-readable storage medium corresponding to a target detection method based on sparse area extraction, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it can implement the target detection method based on sparse area extraction as described above.

[0072] Therefore, based on the same inventive concept, Figure 4 As shown, the present invention also provides a computer electronic device corresponding to the target detection method based on sparse region extraction provided in the above embodiment, which includes a memory and a processor;

[0073] The memory is used to store computer programs;

[0074] The processor is used to implement the target detection method based on sparse region extraction as described above when executing the computer program;

[0075] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by the processor to perform the above steps S1 to S5.

[0076] It is understandable that the above storage medium may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. The storage medium may also be a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc., which can store program codes.

[0077] It is understandable that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0078] It should also be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0079] In order to better demonstrate the specific implementation method and technical effect of the above-mentioned target detection method based on sparse region extraction of the present invention, a specific example is given below to demonstrate the construction, training and test performance of its target detection model on a data set.

[0080] Example

[0081] In this embodiment, the specific process of the target detection method based on sparse region extraction is as follows: Figure 5 As shown, it specifically includes step 1 to step 4. The specific implementation of each step is described in detail below.

[0082] Step 1: Build a target detection model based on sparse region extraction

[0083] The target detection model consists of a feature extraction module, a feature fusion enhancement module, a sparse region extraction module, a first target detection head module, a second target detection head module and a non-maximum suppression module. The feature extraction module is used to extract feature information of the image to be detected. The feature fusion module is used to fuse and enhance the features processed by the feature extraction module. The sparse region extraction module is used to select key areas of the high-resolution feature layer. The first target detection head module is a small target detection head module, which is used to perform small-size target detection on the features selected based on the sparse region extraction module. The second target detection head module is a conventional convolution detection head module, which is used to perform global target detection on the low-resolution feature layer. The non-maximum suppression module is used to perform non-maximum suppression on all candidate boxes detected by the first target detection head module and the second target detection head module, and retain the final detection box.

[0084] In this embodiment, Resnet-50 is used as a feature extraction module to extract image features. For an image to be detected with a size of [H, W], a multi-level feature map is generated through the Resnet-50 network. The size of the output feature map of each layer is In the embodiment, 4 groups of residual modules of Resnet-50 are used. The four layers of output features (i.e., the outputs of Stage2, Stage3, Stage4, and Stage5), the feature size corresponding to the i-th layer feature is , In order to make better use of features at different levels, the feature pyramid model FPN is used as a feature fusion enhancement module to fuse and extract multi-level features. FPN further fuses the previous 4-layer output features of Resnet-50, and finally outputs 6 layers of enhanced feature maps of different scales. The final multi-channel enhanced feature map is }.

[0085] For target detection based on anchor frames, multiple pre-specified anchor frames of different sizes are generated for each feature point position in each layer of feature map of feature representation to enhance adaptability to various target shapes. However, this will result in a large number of pre-selected frames, and the amount of calculation for high-resolution feature layers will be particularly large. In order to enhance the detection effect of small target detection, it is necessary to process it at the high-resolution feature layer. Therefore, the present invention converts the above multi-channel enhanced feature map into } is divided into two categories, Enhance features accordingly The two layers of enhanced features with the highest resolution, enhanced features and Enhanced Features It can be used as two layers as the first feature to screen key areas and detect small-sized targets from them, and the rest Enhance features accordingly The four layers of enhanced features with lower resolution are enhanced features. , Enhanced Features , Enhanced Features , Enhanced Features The fourth layer can be used as the second feature to detect targets in the global range.

[0086] when When the first features of these two layers of high resolution are used, they need to first pass through the sparse region extraction module (Sparse Region Extractor, SRE) to select multiple local regions, and then only these local regions and the anchors generated by the corresponding regions are used for subsequent detection. The sparse region extraction module consists of a Patch Embedding layer, a WindowMSA layer, a Shifted Window MSA layer, a layer normalization (LayerNorm) layer, and a 3×3 convolution layer. The process within the module is as follows: Patch Embedding first divides the input image to be detected into multiple fixed-size image patches (patches), and converts each image patch into a low-dimensional vector representation, and then sends it to two layers of attention layers based on the window mechanism. The first WindowMSA layer uses a fixed window, and the second Shifted Window MSA layer uses a moving window for calculation. After that, it is normalized by the LayerNorm layer, and finally a 3×3 convolution layer is used for final prediction. The overall process is as follows Figure 6 shown.

[0087] In this embodiment, the parameters of each layer inside the sparse region extraction module are set as follows: the first layer, the PatchEmbedding layer, is calculated using Conv2d, the convolution kernel and the step size are set to the same size of 1, and the output channel is 256; the window size used by the WindowMSA layer and the Shifted Window MSA layer is set to 4, and the activation layer used is GELU; the embed_dims in the LayerNorm layer used is set to 256; the last 3×3 convolution layer uses the convolution kernel size set to 3, the step size is set to 1, and the number of output channels is set to 1.

[0088] Therefore, it is assumed that the first feature size of the input sparse region extraction module is , when passing through the PatchEmbedding module, the module uses convolution operations to obtain The data in each window is processed through the WindowMSA layer and the Shifted Window MSA layer. These two layers of modules will not change the shape of the feature layer. Then the LN layer is used for normalization to prevent gradient explosion. Finally, it passes through a convolutional layer with a kernel of 3×3 and changes the number of channels to 1, which is the output single-channel feature.

[0089] For the single-channel features output from the sparse region extraction module, this embodiment uses an 8*8 window to perform sliding segmentation and extraction, with a step size of 8. Each extracted window contains 8*8 feature values, and the probability value of the window being selected as the feature window is calculated based on the local features in the window. , the calculation formula is as follows:

[0090]

[0091] Where: represents the maximum eigenvalue within the calculation window, represents the average eigenvalue within the calculation window, and Represents the corresponding weight hyperparameter, where Set to 0.75, Set to 0.25.

[0092] In this embodiment, for each first feature (enhanced feature , Enhanced Features ), we need to take the M=100 windows with the largest probability value as the feature windows most likely to contain small-sized targets. For each of the 100 8*8 feature windows selected from the first feature, map each feature window to the corresponding first feature, and extract a local feature block of size [C,8,8] within the feature window range. Then, splice the 100 feature blocks along the height H direction of the feature to form a feature window of size The concatenated local feature map , the concatenated local feature map It is calculated as input to subsequent modules.

[0093] For the first feature of each layer, a priori boxes need to be generated before target detection. The feature layer, For each feature point in all feature windows, n=9 anchor boxes need to be formed as prior boxes and flattened to form The vector of size represents the anchor. The 4-dimensional information here represents the coordinates in the format of xyxy. The anchor box remains unchanged throughout the process. According to the transformed coordinates, we can extract 100*8*8*9 anchors. Note that the order here is first based on the feature block order, and then sorted by the H and W order in each feature block.

[0094] After generating 9 first prior frames based on each feature point in all feature windows for the first feature of each layer, the local feature blocks of all feature windows in the first feature can be spliced ​​to form a spliced ​​local feature map. The first target detection head module is input, and the predicted category and offset position of each first priori box are obtained through the first category prediction branch and the first box prediction branch respectively.

[0095] The first target detection head module includes a first category prediction branch and a first box prediction branch, such as Figure 7 As shown in Figure 2, each branch consists of 2 WindowMSA layers (WindowSize is 4), an LN layer, and a 3×3 convolutional layer. The input is The concatenated local feature map , the output of the category prediction branch , the output channel K is the number of target categories, n is the number of prior boxes generated by a single feature point; the output of the first box prediction branch In this embodiment, n=9.

[0096] In addition, when These four layers of low-resolution second features need to perform global target detection, rather than referring to the first features for small-size target detection in local key areas. In an embodiment of the present invention, for each layer of second features, multiple second priori boxes are generated based on each feature point in the second features, and the second features are directly input into the second target detection head module, and the predicted category and offset position of each second priori box are obtained through the second category prediction branch and the second box prediction branch.

[0097] It should be noted that for For the low-resolution feature layer part of the layer, the present invention does not perform special processing through the sparse region extraction module, directly uses the features output by FPN, and subsequently connects the convolution detection head module for prediction, that is, it is necessary to generate a second priori box for each feature point in the entire feature map. In this embodiment, the convolution detection head adopts a standard Retinanet detection head module, and the second category prediction branch and the second box prediction branch are each cascaded twice by 5 layers of 3×3 convolutions, and can output the box category and box position respectively.

[0098] It should be noted that the output formats of the first target detection head module and the second target detection head module of the two detection heads used in the present invention are similar, and each detection head module contains two prediction layers, a category prediction layer and a box prediction layer. The input feature size of each detection head is , the number of channels will change after passing through the detection head module. For each feature point, the category prediction layer outputs the result of multiplying the number of anchor boxes for each point by the number of detection categories, which indicates the confidence of each category corresponding to each box. The higher the confidence, the greater the probability that the target belongs to that category. The output of the box prediction layer for each feature point is 4 times the number of prior boxes. Each anchor box corresponds to 4 values. These 4 numbers do not directly represent the coordinates of the box for the image, but represent the offset of the prior box. The coordinates of the prior box can be corrected based on the offset. For a pre-generated prior box, assume that its coordinates are , and the corresponding offset prediction , the coordinates of the new detection frame after offset can be obtained by the following formula :

[0099]

[0100]

[0101]

[0102] The last module is the non-maximum suppression module. All candidate boxes detected by the first target detection head module and the second target detection head module can be merged and input into the non-maximum suppression module for non-maximum suppression. The final detection box is retained, which is the final output detection result.

[0103] Therefore, a target detection model based on sparse region extraction can be formed by the feature extraction module, feature fusion module, sparse region extraction module, first target detection head module, second target detection head module and non-maximum suppression module. The model needs to be supervised trained on an annotated training data set through subsequent steps.

[0104] Step 2: Prepare training dataset

[0105] The training data set consists of images to be detected containing targets. Each image has corresponding target annotation information, which includes the vertex coordinates of the box of each target and its category.

[0106] Step 3: Train the neural network based on sparse region extraction

[0107] The images used for training are preprocessed to form the same format and size, and then input into the target detection model based on sparse region extraction. After being output by the target detection model, the total loss is calculated, and the gradient is updated using the SGD algorithm based on the calculated loss value. The training data is continuously extracted batch by batch for iterative training to finally achieve the training effect.

[0108] In this embodiment, the total loss used is mainly composed of three loss losses, namely query loss query_loss, category prediction loss cls_loss and location regression loss bbox_loss. Among them, query_loss is for the high-resolution layer, that is, the feature layer using the sparse region extraction module, and its formula is as follows

[0109]

[0110] SRE(feather) represents the output of the first feature map feather after passing through the sparse region extraction module (SRE module), and GT is the true value labeled feature map with a 0-1 label, where 1 indicates that there is an object at that position in the image, and 0 indicates that there is no object.

[0111] When generating a true value annotation feature map, anchor boxes can be pre-annotated on the image to be detected after the initial size is unified, and then all the annotated anchor boxes are scaled and mapped to the size of the corresponding first feature map. Then, according to the size threshold in the current first feature map, the anchor boxes whose sizes are smaller than the size threshold after mapping are screened out, and the areas within these anchor boxes are marked as 1, and the rest are marked as 0.

[0112] Since the sizes of the first feature maps of the two layers are different, the corresponding size thresholds are also different. The feature layer of the size threshold is set to 32*32, so only the inner area of ​​the annotation anchor box with an area smaller than 32*32 is selected and marked as 1, and the rest of the feature points are marked as 0; For the feature layer, the size threshold is set to 64*64, so only the inner area of ​​the annotation anchor box with an area smaller than 64*64 is selected and marked as 1, and the rest of the feature points are marked as 0.

[0113] Cls_loss represents the category prediction loss, and also uses the Focal Loss loss function. Bbox_loss represents the box regression loss, and uses the L1Loss loss loss function. It should be noted that when calculating the above loss, the loss can be calculated based on the candidate anchor frame set before the NMS layer. However, not all prediction results of each feature point and all corresponding anchor frames are involved in the loss calculation. The anchor frames generated in advance determine which frames are used for loss calculation. This embodiment calculates the iou value of the predicted frame composed of the above coordinate-filtered frames and all frames of the low-resolution layer and the real frame of the image. There is a value for each predicted frame and each real frame. Only when the maximum value is greater than a certain threshold value, it will be used for subsequent loss calculation. For each real frame, this embodiment will assign a frame with the largest iou value as the selected frame to ensure that each real frame has a predicted frame to predict it, so as to calculate the loss corresponding to the true value.

[0114] The final total loss function Loss is calculated as follows:

[0115]

[0116] in , , Represents three weight hyperparameters.

[0117] Step 4: Use the image to be detected for detection

[0118] The image to be detected is placed into the trained target detection model, and then the detection boxes of all targets in the image to be detected and the corresponding target categories are finally obtained.

[0119] In order to demonstrate the technical effect of this embodiment, this embodiment verifies the trained target detection model (denoted as the present invention) on the visdrone2019 dataset, using AP as the evaluation index. Table 1 shows the overall detection results of the present invention and the traditional Retinanet model on the dataset, and its evaluation indicators include mAP, mAP_50, mAP_75, mAP_s, mAP_m and mAP_l. Table 2 shows the mAP detection results on 10 different types of targets on the dataset, and the target categories are: pedestrian, people, bicycle, car, van, truck, tricycle, awning-tricycle, bus, and motorcycle. It can be seen that the present invention is superior to the Retinanet model in overall results.

[0120] Table 1

[0121] mAP mAP_50 mAP_75 mAP_s mAP_m mAP_l Retinanet 0.208 0.378 0.205 0.110 0.329 0.376 The present invention 0.231 0.425 0.220 0.151 0.327 0.336

[0122] Table 2

[0123] pedestrian figure bike car truck truck Three-wheeler Covered tricycle the bus motorcycle Retinanet 0.164 0.104 0.079 0.502 0.28 0.224 0.132 0.09 0.354 0.161 The present invention 0.218 0.147 0.09 0.536 0.296 0.205 0.146 0.082 0.392 0.201

[0124] It can be seen that the present invention selects feature areas in high-resolution layers and can achieve accurate detection of small targets without predicting the entire feature layer.

[0125] The above-described embodiments are only some preferred implementations of the present invention, but are not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.

Claims

1. A target detection method based on sparse region extraction, characterized in that: include: S1. For the input image to be detected, the feature extraction module is used to extract image features from it, and then the image features are fused and enhanced by the feature fusion module to obtain enhanced features with different resolutions, and the two layers of enhanced features with the highest resolution are used as the first features, and the remaining enhanced features are used as the second features; S2. For each layer of the first feature, input it into the sparse region extraction module, and after passing through the patch embedding layer, the window multi-head self-attention layer, the moving window multi-head self-attention layer, and the layer normalization layer in the module, the number of channels is compressed to 1 through the convolution layer, and then the compressed single-channel feature is divided into a series of windows, and the weighted sum of the maximum value and the average value of the pixels in each window is calculated and used as the probability value of the window being selected, and a part of the windows with the largest probability value are screened out from all windows as feature windows; S3, for each layer of the first feature, generate multiple first a priori boxes based on each feature point in all feature windows, splice the local feature blocks of all feature windows in the first feature and input them into the first object detection head module, and obtain the predicted category and offset of each first a priori box through the first category prediction branch and the first box prediction branch respectively; S4. For each layer of the second feature, generate multiple second a priori boxes based on each feature point in the second feature, and directly input the second feature into the second object detection head module, and obtain the predicted category and offset of each second a priori box through the second category prediction branch and the second box prediction branch respectively; S5. All first a priori boxes and all second a priori boxes are offset according to their respective offsets and added to the candidate anchor box set. Then, the redundant anchor boxes are removed through the non-maximum suppression module, and finally the detection boxes of all targets in the image to be detected are obtained.

2. The target detection method based on sparse region extraction according to claim 1, characterized in that: The feature extraction module adopts the Resnet network.

3. The target detection method based on sparse region extraction according to claim 1, characterized in that: The feature fusion module adopts a feature pyramid model (FPN), and the final output enhanced features have a total of 6 layers, the two layers of enhanced features with the highest resolution are used as the first features, and the remaining 4 layers of enhanced features are used as the second features.

4. The target detection method based on sparse region extraction according to claim 1, characterized in that: In the first target detection head module, the first category prediction branch and the first box prediction branch have the same structure, both of which are cascaded by 2 layers of window multi-head self-attention layers, layer normalization layers and 3×3 convolution layers.

5. The target detection method based on sparse region extraction according to claim 1, characterized in that: In the second object detection head module, the second category prediction branch and the second box prediction branch have the same structure, both of which are formed by cascading multiple layers of 3×3 convolutional layers.

6. The target detection method based on sparse region extraction according to claim 1, characterized in that: The target detection model composed of the feature extraction module, the feature fusion module, the sparse area extraction module, the first target detection head module, the second target detection head module and the non-maximum suppression module needs to be supervised trained on a labeled training data set in advance.

7. The target detection method based on sparse region extraction according to claim 6, characterized in that: The loss function used in the supervised training is the weighted sum of query loss, category prediction loss, and position regression loss; the query loss is the focal loss (FocalLoss) between the Sigmoid activation output of the single-channel feature and the true value labeled feature map, in which only the anchor frame area where the target with a size smaller than a preset value is located in the true value labeled feature map is marked as 1, and the rest of the positions are marked as 0; the category prediction loss is the focal loss (Focal Loss) between the predicted category and the true value category of all detection boxes; the position regression loss is the L1Loss loss between the predicted offset and the true offset of all detection boxes.

8. A target detection system based on sparse region extraction, characterized in that: include: An image input module is used to input an image to be detected that needs to be detected; A detection module, configured to perform target detection on the image to be detected inputted in the image input module according to the target detection method based on sparse region extraction according to any one of claims 1 to 7, and obtain detection frames of all targets in the image to be detected and a predicted category corresponding to each detection frame; The output module is used to output the detection box and the predicted category obtained in the detection module according to a preset output method.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the target detection method based on sparse region extraction as described in any one of claims 1 to 7 can be implemented.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the target detection method based on sparse region extraction as described in any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Video target detection method and device based on sparse foreground prior, and storage medium

    CN112434618A

  • Arbitrary-shape scene text detection method based on contour feature enhancement

    CN119296094A