Object Detection Method, Device, Electronic Device, and Storage Medium
By extracting and fusing feature maps of different sizes and enhancing feature representations using discrete number estimation models, the problem of inaccurate detection in dense object detection is solved, and higher object detection accuracy is achieved.
Patent Information
- Application Number
- CN202510345995.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art has the problem of inaccurate detection in dense object detection, especially when the target objects are blocked or overlapped in dense scenes, it is difficult to accurately detect and locate the target position.
By extracting the feature maps of different sizes of the image to be detected, the discrete number estimation model is used to estimate the original feature map in discrete number, and fuse it with the original feature map to generate a mixed feature map, and input the trained object detection model for target object detection.
The accuracy of object detection is improved, especially in dense target scenarios, which can more accurately locate and identify target objects.
Smart Images

Figure CN119851219B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of object detection, and particularly to an object detection method, device, electronic device, and storage medium. Background Art
[0002] Dense object detection is a detection task performed when there are a large number of closely arranged object targets (such as pedestrians, vehicles, etc.) in an image. These targets may be occluded or overlapped with each other, increasing the difficulty of detection. In practical application scenarios such as urban traffic monitoring and shopping mall passenger flow statistics, accurately detecting and analyzing dense targets is crucial for traffic management and abnormal behavior recognition, and has important practical significance.
[0003] Currently, object detection for dense scenes mainly estimates the number of objects by generating a density map of the object distribution and indirectly locates the object positions. For example, models such as multi-column convolutional neural networks are used to extract multi-scale features to generate a density map, or by the anchor method, the image is divided into grids, and a fixed number of anchors are predicted for each grid, and the objects are detected by sliding on the feature map. However, due to the large number of objects in dense scenes, and the close distance and small size between objects, the model often has difficulty focusing on a single object, resulting in limited detection performance. The overlap between anchors may also lead to inaccurate detection results.
[0004] Regarding the problem of inaccurate object detection in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] In this embodiment, an object detection method, device, electronic device, and storage medium are provided to solve the problem of inaccurate object detection in the related art.
[0006] In a first aspect, in this embodiment, an object detection method is provided, including:
[0007] Extract the image features of the image to be detected to obtain the original feature map of the image to be detected, where the original feature map includes feature maps of different sizes;
[0008] Based on the trained discrete quantity estimation model, perform discrete quantity estimation on the original feature map and output a discrete quantity estimation feature map;
[0009] Fuse the original feature map and the discrete quantity estimation feature map to obtain a mixed feature map;
[0010] Input the mixed feature map into the trained object detection model for object detection to obtain an object detection result.
[0011] In some of these embodiments, extracting the image features of the image to be detected to obtain the original feature map of the image to be detected includes:
[0012] Extracting an initial feature map from the image to be detected through a preset convolutional network model;
[0013] Performing downsampling and then upsampling on the initial feature map;
[0014] Adding the feature maps with the same size and the same number of channels during the downsampling process and the upsampling process to obtain the original feature map of the image to be detected.
[0015] In some of these embodiments, during the downsampling process of the initial feature map, the method further includes:
[0016] Inputting the initial feature map into the first convolutional module and the bottleneck module of the preset convolutional network model, and outputting the first feature map of the image to be detected;
[0017] Inputting the initial feature map into the second convolutional module of the preset convolutional network model, and outputting the second feature map of the image to be detected;
[0018] Concatenating the first feature map and the second feature map to obtain the initial feature map after downsampling.
[0019] In some of these embodiments, the training process of the discrete quantity estimation model includes:
[0020] Training an initial discrete quantity estimation model according to depthwise separable convolution and a classifier to obtain a trained discrete quantity estimation model;
[0021] Among them, in the classifier, training the initial discrete quantity estimation model includes:
[0022] After partitioning the discrete quantity estimation feature map, estimating the quantity of the target objects in the discrete quantity estimation feature map to obtain the quantity estimation results of each target object, and training the initial discrete quantity estimation model according to the quantity estimation results of each target object.
[0023] In some of these embodiments, fusing the original feature map and the discrete quantity estimation feature map to obtain a hybrid feature map includes:
[0024] Performing depthwise separable convolution processing on the original feature map to obtain two third feature maps and fourth feature maps with the same size;
[0025] Reshape the discrete quantity estimation feature map, the third feature map, and the fourth feature map to feature maps of the same size, obtaining the reshaped discrete quantity estimation feature map, the reshaped third feature map, and the reshaped fourth feature map;
[0026] Fuse the reshaped third feature map, the reshaped fourth feature map, and the reshaped discrete quantity estimation feature map to obtain a mixed feature map.
[0027] In some embodiments, the fusing the reshaped third feature map, the reshaped fourth feature map, and the reshaped discrete quantity estimation feature map to obtain a mixed feature map includes:
[0028] Transpose the reshaped discrete quantity estimation feature map to obtain the transposed discrete quantity estimation feature map,
[0029] Multiply the transposed discrete quantity estimation feature map by the reshaped third feature map and input the result into a classifier, and calculate the distribution probability of the target object through the classifier to obtain a fifth feature map;
[0030] Perform matrix operation on the fifth feature map and the reshaped fourth feature map to obtain a sixth feature map;
[0031] Reshape the size of the sixth feature map to the same size as the original feature map to obtain the mixed feature map.
[0032] In some embodiments, the training process of the target detection model includes:
[0033] Obtain the mixed feature map of the training image;
[0034] Pre-annotate the target object in the training image to obtain an initial target object annotation result;
[0035] Based on the initial target object annotation result and the mixed feature map of the training image, train the initial target detection model for the target object detection task to obtain the trained target detection model.
[0036] In a second aspect, a target detection device is provided in this embodiment, including: an original feature extraction module, a discrete quantity estimation model, a mixed feature module, and a target detection module, where
[0037] The original feature extraction module is configured to extract the image features of the image to be detected to obtain the original feature map of the image to be detected, where the original feature map includes feature maps of different sizes;
[0038] The discrete quantity estimation model is used to perform discrete quantity estimation on the original feature map and output a discrete quantity estimation feature map;
[0039] The hybrid feature module is used to fuse the original feature map and the discrete quantity estimation feature map to obtain a hybrid feature map;
[0040] The target detection module is used to input the hybrid feature map into a trained target detection model for target object detection to obtain a target object detection result.
[0041] In a third aspect, an electronic device is provided in this embodiment, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the target detection method described in the first aspect above is implemented.
[0042] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored. When the program is executed by a processor, the target detection method described in the first aspect above is implemented.
[0043] Compared with the related art, in the target detection method provided in this embodiment, by extracting the image features of the image to be detected, the original feature map of the image to be detected is obtained, where the original feature map includes feature maps of different sizes; based on a trained discrete quantity estimation model, discrete quantity estimation is performed on the original feature map to output a discrete quantity estimation feature map; the original feature map and the discrete quantity estimation feature map are fused to obtain a hybrid feature map; the hybrid feature map is input into a trained target detection model for target object detection to obtain a target object detection result. It performs discrete quantity estimation on the original feature map through a discrete quantity estimation model to obtain a discrete quantity estimation feature map, then fuses the discrete quantity estimation feature map with the original feature map to enhance the image features and obtain a hybrid feature map, and performs target detection based on the hybrid feature map with enhanced image features, thereby improving the accuracy of target detection.
[0044] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0046] Figure 1 It is a hardware structure block diagram of the terminal of the target detection method in this embodiment.
[0047] Figure 2 It is a flowchart of the object detection method of this embodiment.
[0048] Figure 3 It is a schematic diagram of the original feature map extraction network structure of the object detection method of this embodiment.
[0049] Figure 4 It is a processing structure diagram of the C3 convolution module of the object detection method of this embodiment.
[0050] Figure 5 It is a schematic diagram of the discrete quantity estimation model structure of the object detection method of this embodiment.
[0051] Figure 6 It is a partition diagram of the discrete quantity estimation model of the object detection method of this embodiment.
[0052] Figure 7 It is a schematic diagram of the feature mixing module structure of the object detection method of this embodiment.
[0053] Figure 8 It is a flowchart of another object detection method of this embodiment.
[0054] Figure 9 It is a structural block diagram of the object detection device of this embodiment. Detailed implementation manners
[0055] To understand the purpose, technical solution and advantages of this application more clearly, the following describes and explains this application in combination with the accompanying drawings and embodiments.
[0056] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings as understood by those of ordinary skill in the technical field to which this application pertains. In this application, words such as "a", "an", "one kind", "the", "these", etc. do not indicate a limitation in quantity and can be singular or plural. The terms "include", "comprise", "have" and any variations thereof used in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connect", "be connected", "couple" and the like used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" used in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Generally, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. used in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0057] The method embodiments provided in this embodiment can be executed on a terminal, a computer, or a similar computing device. For example, running on a terminal, Figure 1 is a hardware structure block diagram of the terminal of the object detection method in this embodiment. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 and a memory 104 for storing data. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those shown in Figure 1 the figure, or have a different configuration from that shown in Figure 1 the figure.
[0058] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the object detection method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0059] The transmission device 106 is used to receive or send data via a network. The above-mentioned network includes a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0060] In this embodiment, an object detection method is provided. Figure 2 is a flowchart of the object detection method of this embodiment, as Figure 2 shown, this process includes the following steps:
[0061] Step S201, extract the image features of the image to be detected to obtain the original feature map of the image to be detected, where the original feature map includes feature maps of different sizes.
[0062] Specifically, image features of the image to be detected are extracted through a convolutional neural network (CNN) to obtain the original feature map of the image to be detected. Specifically, the input image to be detected first undergoes processing by multiple convolutional layers and pooling layers, gradually extracting local features and global features in the image. These local features and global features constitute the original feature map of the image to be detected. Each convolutional operation generates a set of feature maps, which reflect the abstract representation of the image at different levels. As the number of network layers deepens, the size of the feature map gradually decreases, but the number of channels increases, which means that the feature map contains more abstract and high-level semantic information. For example, shallow feature maps may contain low-level features such as edges and textures, while deep feature maps may contain semantic information about object parts or the overall object. Smaller feature maps can capture global information in the image, such as the overall structure, layout, and large-scale texture features of the image. These global information are crucial for understanding the overall semantics and context relationships of the image. Larger feature maps, on the other hand, pay more attention to local details in the image and can extract microscopic features such as edges, textures, and shapes. These detail features are important for identifying specific objects or texture patterns in the image. By constructing an original feature representation that includes feature maps of different sizes, multi-scale information of the image can be more comprehensively retained, thereby providing a richer and more accurate feature basis for subsequent image analysis tasks. For example, in object detection tasks, such multi-scale feature maps can simultaneously help the algorithm better locate the position of the object and identify the detail features of the object; in image classification tasks, it can more effectively extract features related to the category, thereby improving the accuracy and robustness of classification.
[0063] Step S202: Based on the trained discrete quantity estimation model, perform discrete quantity estimation on the original feature map and output a discrete quantity estimation feature map.
[0064] Specifically, in this embodiment, to enhance the accuracy of object detection, discrete quantity estimation is performed on the original image features through a discrete quantity estimation model, dividing continuous feature data into several discrete intervals, converting the classification task of continuous quantities into a discrete classification task, evaluating each feature value in the original feature map, and mapping the evaluation result of each feature value to a preset discrete interval or category to generate a discrete quantity estimation feature map, learning quantity-related feature representations from the original feature map, and predicting the discrete quantity of the object in the image to obtain a discrete quantity estimation feature map. Among them, the discrete quantity estimation model is pre-trained and does not need to be trained for the quantity estimation task in the object detection task. The discrete quantity estimation model can reduce the complexity of the model and improve the accuracy of classification. Improve the accuracy of object recognition.
[0065] Step S203: Fuse the original feature map and the discrete quantity estimation feature map to obtain a hybrid feature map.
[0066] Specifically, in order to enhance the target detection model's ability to understand image features, the original feature map and the discrete quantity estimation feature map are fused to obtain a more effective feature representation. By combining the quantity information, semantic information with the original features, the model can better understand the distribution and quantity of the targets. By incorporating the quantity information, it can also help the target detection model to more accurately locate the targets. Among them, the fusion method can be through concatenation, directly concatenating the original feature map and the discrete quantity estimation feature map together to form a higher-dimensional feature map. The direct concatenation method will increase the number of channels of the features, but will not change the information within each feature map. For example, if the number of channels of the original feature map is C1 and the number of channels of the discrete feature map is C2, then the number of channels of the concatenated hybrid feature map is C1 + C2. Or add the two feature maps pixel by pixel to enhance the information content of each feature map. This method will not increase the number of channels of the features, but will enhance the information content of each feature map. Or other element-wise operations (such as multiplication, maximum value, etc.) can also be used to fuse the feature maps. For example, element-wise multiplication can highlight the regions that are commonly significant in the two feature maps. By fusing the original feature map and the discrete quantity estimation feature map, the detailed information of the original feature map and the semantic information of the discrete quantity estimation feature map can be integrated, and the target quantity information in the discrete quantity estimation feature map is also incorporated, improving the detection accuracy.
[0067] Step S204, input the hybrid feature map into the trained target detection model for target object detection to obtain the target object detection result.
[0068] The target detection model is a specially designed neural network for extracting the location and category information of target objects from the hybrid feature map. Among them, the target detection model can be a two-stage detection model such as Faster R-CNN, Mask R-CNN, or a single-stage detection model such as YOLO (You Only Look Once), SSD (Single Shot MultiBoxDetector). This embodiment does not make specific restrictions on the target detection model, and it can be selected according to the actual situation. The target detection model in this embodiment is obtained by pre-training with a large number of hybrid feature maps. Input the hybrid feature map into the trained target detection model for target object detection to obtain the target object detection result, where the target object detection result includes bounding boxes, label categories, confidence scores, etc.
[0069] Through the above steps S201 to S204, the image features of the image to be detected are extracted to obtain the original feature map of the image to be detected. Among them, the original feature map includes feature maps of different sizes; based on the trained discrete quantity estimation model, the original feature map is subjected to discrete quantity estimation, and a discrete quantity estimation feature map is output; the original feature map and the discrete quantity estimation feature map are fused to obtain a mixed feature map; the mixed feature map is input into the trained target detection model for target object detection to obtain the target object detection result. Compared with the density map or anchor point method predicted by the model in the prior art, in this embodiment, by extracting the original feature maps of different sizes, the feature representation of the image is enhanced through the original feature maps of different sizes, and the original feature map is input into the trained discrete quantity estimation model for discrete quantity estimation to obtain the discrete quantity estimation feature map. The discrete quantity estimation model performs discrete quantity estimation on the original feature map, converts the classification task of continuous quantity into a discrete classification task, evaluates each feature value in the original feature map, and maps the evaluation result of each feature value to a preset discrete interval or category to generate a discrete quantity estimation feature map, further enhancing the feature representation of the image to be detected. And the joint discrete quantity estimation model-assisted training can also further improve the ability and accuracy of the target detection model to identify targets. Then, the original feature map and the discrete quantity estimation feature map are fused in terms of features, further enhancing the feature representation of the image to be detected to obtain a mixed feature map. Finally, target detection is performed based on the processed mixed feature map to obtain the target detection result. The accuracy of target detection is improved. In addition, the target detection method in this embodiment is not only applicable to the detection of a single target object, but even when applied to a scenario with a large number of targets and closely arranged targets, it can also achieve a high target detection accuracy through the auxiliary training and feature fusion of the discrete quantity estimation model.
[0070] In some of these embodiments, extracting the image features of the image to be detected to obtain the original feature map of the image to be detected includes: extracting a preset number of image features of different sizes from the image to be detected through a preset convolutional network model to obtain each initial feature map; performing downsampling and then upsampling processing on each initial feature map; adding the feature maps with the same size and the same number of channels in the downsampling process and the upsampling process to obtain the original feature map of the image to be detected.
[0071] Specifically, in the above step S201, in order to obtain an original feature map with better feature performance, this embodiment is extracted by a combination of upsampling and downsampling. Figure 3 It is a schematic diagram of the network structure for extracting the original feature map of the target detection method in this embodiment. As Figure 3 shown, first, initial feature maps are extracted from the image to be detected through an existing convolutional network model. The size of the initial feature map is H×W, and the number of channels is 3.
[0072] The initial feature map is first downsampled using the convolutional module Conv(3, 64, 3, 2, 2). The dimensions of the convolutional module Conv(3, 64, 3, 2, 2) are as follows:
[0073] Number of input channels: 3;
[0074] Number of output channels: 64;
[0075] Kernel size: 3x3;
[0076] Stride: 2;
[0077] Number of groups: 2;
[0078] The feature map output by the convolutional module Conv(3, 64, 3, 2, 2) is input into the convolutional module C3(64) for further processing, and the output is: a feature map with a size of H / 2 × W / 2 and 64 channels.
[0079] Based on the image output by the convolutional module C3(64), a second downsampling is performed using the convolutional module Conv(64, 128, 3, 2). The dimensions of the convolutional module Conv(64, 128, 3, 2) are as follows:
[0080] Number of input channels: 64;
[0081] Number of output channels: 128;
[0082] Kernel size: 3x3;
[0083] Stride: 2;
[0084] The feature map output by the convolutional module Conv(64, 128, 3, 2) is input into the convolutional module C3(128) for further processing, and the output is: a feature map with a size of H / 4 × W / 4 and 128 channels.
[0085] Based on the image output by the convolutional module C3(128), a third downsampling is continued using the convolutional module Conv(128, 256, 3, 2). The dimensions of the convolutional module Conv(128, 256, 3, 2) are as follows:
[0086] Number of input channels: 128;
[0087] Number of output channels: 256;
[0088] Kernel size: 3x3;
[0089] Stride: 2;
[0090] The feature map output by the convolutional module Conv(128, 256, 3, 2) is input into the convolutional module C3(256) for further processing. Output: The size of the feature map is H / 8×W / 8; the number of channels is 256.
[0091] Based on the image output by the convolutional module C3(256), the fourth downsampling is continued using the convolutional module Conv(256, 512, 3, 2). The dimensions of this convolutional module Conv(256, 512, 3, 2) are as follows:
[0092] Number of input channels: 256;
[0093] Number of output channels: 512;
[0094] Kernel size: 3x3;
[0095] Stride: 2;
[0096] The feature map output by the convolutional module Conv(256, 512, 3, 2) is input into the convolutional module convolutional module C3(512) for further processing. Output: The size of the feature map is H / 16×W / 16, and the number of channels is 512.
[0097] Based on the image output by the convolutional module C3(512), the fifth downsampling is continued using the convolutional module Conv(512, 1024, 3, 2). The dimensions of this convolutional module Conv(512, 1024, 3, 2) are as follows:
[0098] Number of input channels: 512;
[0099] Number of output channels: 1024;
[0100] Kernel size: 3x3;
[0101] Stride: 2;
[0102] The feature map output by the convolutional module Conv(512, 1024, 3, 2) is input into the convolutional module C3(1024) for further processing. Output: The size of the feature map is H / 32×W / 32, and the number of channels is 1024.
[0103] Based on the image output by the convolutional module C3(1024), the spatial resolution of the feature map is increased and the number of channels is reduced through transposed convolution or interpolation methods (such as bilinear interpolation) for upsampling processing. The first upsampling is performed using the convolutional module Conv(1024, 512, 1, 1) to reduce the number of channels. The dimensions of this convolutional module Conv(1024, 512, 1, 1) are as follows:
[0104] Number of input channels: 1024;
[0105] Number of output channels: 512;
[0106] Convolution kernel size: 1x1;
[0107] Stride: 1.
[0108] Use the upsampling operation to restore the feature map size output by convolution module C3(1024) from H / 32×W / 32 to H / 16×W / 16.
[0109] Output: Feature map size is H / 16×W / 16, and the number of channels is 512.
[0110] Based on the image output by convolution module Conv(1024, 512, 1, 1), use convolution module Conv(512, 256, 1, 1) for the second upsampling to reduce the number of channels. The dimensions of this convolution module Conv(512, 256, 1, 1) are as follows:
[0111] Number of input channels: 512;
[0112] Number of output channels: 256;
[0113] Convolution kernel size: 1x1;
[0114] Stride: 1;
[0115] Use the upsampling operation (such as bilinear interpolation or transposed convolution) to restore the feature map size output by convolution module Conv(1024, 512, 1, 1) from H / 16×W / 16 to H / 8×W / 8. Output: Feature map size is H / 8×W / 8, and the number of channels is 256.
[0116] Among them, the Conv module is a convolutional module, and its parameters are the number of input channels, output dimension, kernel size, stride size, and number of groups of a standard convolution respectively. A convolutional module contains a standard convolutional layer, a BN layer, and an activation function SiLU. Upsample is responsible for upsampling, with a factor of 2, and the sampling method is bilinear sampling. The specific number of upsampling and downsampling operations is not limited and can be set according to the actual situation. The C3 module is another convolutional module in the feature extraction network structure. The C3 convolutional module (CSP Bottleneck with 3 convolutions) is an efficient feature extraction module introduced in YOLOv5 and YOLOv8. As an improved version of CSP (Cross Stage Partial), C3 further enhances the feature expression ability by introducing multiple Bottleneck layers while maintaining lightweight computational complexity. The main function of the C3 module is to combine multiple Bottleneck modules based on a branch structure to perform deeper feature extraction and fusion. It plays an important role in the backbone network and detection head parts of YOLO series models. The original feature map of the image to be detected is jointly extracted by the convolutional module Conv and the convolutional module C3.
[0117] After upsampling and downsampling the initial feature map, the feature maps with the same size and the same number of channels in the upsampling process and the downsampling process are added together. For example, the feature map with a size of H / 8×W / 8 and 256 channels output by the convolutional module Conv(512, 256, 1, 1) in the upsampling process is added to the feature map with a size of H / 8×W / 8 and 256 channels output by the convolutional module C3(256) in the downsampling process to obtain the original feature map P1; the feature map with a size of H / 16×W16 and 512 channels output by the convolutional module Conv(1024, 512, 1, 1) in the upsampling process is added to the feature map with a size of H / 16×W16 and 512 channels output by the convolutional module C3(512) in the downsampling process to obtain the original feature map P2; the feature map with a size of H / 32×W / 32 and 1024 channels output by the convolutional module C3(1024) in the upsampling process is added to the feature map with a size of H / 32×W / 32 and 1024 channels output by the convolutional module C3(1024) in the downsampling process to obtain the original feature map P3.
[0118] Through Figure 3 multiple downsampling and upsampling operations of the convolutional network, multi-scale features are gradually extracted and fused. The downsampling process reduces the size of the feature map through convolution and operations with a stride greater than 1, while the upsampling process restores the size of the feature map through transposed convolution or interpolation methods, which can fuse shallow and deep feature maps and improve the feature representation of the original feature map.
[0119] In another embodiment, during the downsampling process of the initial feature map, it further includes:
[0120] Input the initial feature map into the first convolution module and the bottleneck module of a preset convolutional network model to output the first feature map of the image to be detected; input the initial feature map into the second convolution module of the preset convolutional network model to output the second feature map of the image to be detected; splice the first feature map and the second feature map to obtain the downsampled initial feature map.
[0121] Specifically, during the above-mentioned downsampling process of the initial feature map, this embodiment also uses the C3 convolution module to process the initial feature map. The C3 convolution module (CSP Bottleneck with 3 convolutions) is an efficient feature extraction module introduced in YOLOv5 and YOLOv8. As an improved version of CSP (Cross Stage Partial), C3 further enhances the feature expression ability by introducing multiple Bottleneck layers while maintaining a lightweight computational load. The main function of the C3 module is to perform deeper feature extraction and fusion on the basis of a branch structure by combining multiple Bottleneck modules. It plays an important role in the backbone network and the detection head part of the YOLO series models. Figure 4 It is the processing structure diagram of the C3 convolution module of the object detection method in this embodiment, as Figure 4 shown. The basic structure of the C3 convolution module contains two branches: one branch includes the first convolution module and the bottleneck module (residual Bottleneck module). The Bottleneck module usually contains multiple convolutional layers and residual connections for enhancing feature representation. The other branch includes the second convolution module. Input the feature map with the size of C×H×W, where C is the number of channels, H is the height, and W is the width. Perform a convolution operation on the initial feature map through the first convolution module Conv(C, 0.5C, 1, 1):
[0122] Number of input channels: C;
[0123] Number of output channels: 0.5C;
[0124] Convolution kernel size: 1x1;
[0125] Stride: 1;
[0126] After the convolution operation, the size of the output feature map is 0.5C×H×W.
[0127] Based on the image output by the first convolution module Conv(C, 0.5C, 1, 1), perform a Bottleneck operation to further extract features.
[0128] The size of the output feature map remains 0.5C×H×W, and the first feature map is obtained.
[0129] Perform a convolution operation on the initial feature map through the second convolution module Conv(C, 0.5C, 1, 1):
[0130] Number of input channels: C;
[0131] Number of output channels: 0.5C;
[0132] Convolution kernel size: 1x1;
[0133] Stride: 1;
[0134] After the convolution operation through the second convolution module Conv(C, 0.5C, 1, 1), the size of the output feature map is 0.5C×H×W, and the second feature map is obtained.
[0135] Concatenate (Concat) the output feature maps of the two branches in the channel dimension:
[0136] Size of the first feature map output by branch one: 0.5C×H×W;
[0137] Size of the second feature map output by branch two: 0.5C×H×W;
[0138] After concatenating the first feature map and the second feature map, the size of the obtained initial feature map is C×H×W.
[0139] Perform a final convolution operation Conv(C, C, 1, 1) on the concatenated feature map:
[0140] Number of input channels: C;
[0141] Number of output channels: C;
[0142] Convolution kernel size: 1x1;
[0143] Stride: 1;
[0144] After the convolution operation, the size of the output downsampled initial feature map remains C×H×W.
[0145] The size of the output downsampled initial feature map is the same as that of the input feature map, but it has undergone feature enhancement and multi-branch fusion processing. This structure can enhance the feature representation ability through multi-branch processing and feature fusion, and further improve the accuracy of object detection.
[0146] In some embodiments, the training process of the discrete quantity estimation model includes:
[0147] The initial discrete quantity estimation model is trained according to depthwise separable convolution and a classifier to obtain a trained discrete quantity estimation model; wherein, in the classifier, training the initial discrete quantity estimation model includes: after partitioning the discrete quantity estimation feature map, estimating the quantity of target objects in the discrete quantity estimation feature map to obtain the quantity estimation results of each target object, and training the initial discrete quantity estimation model according to the quantity estimation results of each target object.
[0148] Specifically, the discrete quantity estimation model in this embodiment is pre-trained through a sample feature map. Figure 5 It is a schematic diagram of the structure of the discrete quantity estimation model of the object detection method in this embodiment; the discrete quantity estimation model in this embodiment includes depthwise separable convolution and a classifier. As Figure 5 shown, its training process includes:
[0149] a. Process the original feature maps P1, P2, P3 through the DWConv convolution module to obtain new feature maps D1, D2, D3; wherein, the DWConv convolution module is depthwise separable convolution, which replaces the standard convolution in the Conv convolution module with depthwise separable convolution. The DWConv convolution module includes depth convolution and pointwise convolution. The depth convolution performs convolution operations on each channel of the input feature map respectively to extract spatial features; then, through pointwise convolution, a 1x1 convolution kernel is used to map the output of the depth convolution to the target number of channels to fuse the information between channels, obtaining new feature maps D1, D2, D3; then, global average pooling (Global AvgPooling) is performed on each new feature map D1, D2, D3 to compress the spatial dimensions (height and width) to 1x1, obtaining G1, G2, G3;
[0150] b. Map the feature vectors G1, G2, G3 after global average pooling to a fixed dimension (such as 64 dimensions, linear(512, 64)) through a linear layer to obtain F1, F2, F3, mapping features of different scales to the same dimensional space for subsequent fusion;
[0151] c. Concatenate (Concat) F1, F2, F3 in the channel dimension to obtain a joint feature F. By fusing multi-scale features, the expression ability of the model is enhanced.
[0152] Finally, the classifier partitions the feature map and estimates the number of target objects in each partition. The loss is calculated based on the difference between the estimated result of the number in each partition and the true label, and the initial discrete number estimation model is optimized to obtain the trained discrete number estimation model. The classifier in this embodiment includes a linear layer (linear(192, 102)) and a Softmax layer. During the training of the classifier, the discrete number estimation feature map is first divided into multiple regions, and each partition corresponds to a local region in the image. The number of partitions can be adjusted according to the task requirements. Figure 6 is a schematic diagram of the partition of the discrete number estimation model of the object detection method in this embodiment, as Figure 6 shown. For example, the rule for dividing 102 interval categories is to divide by an interval of 10, and a regression method is used to predict the number of target objects in each partition. Or a classification method is used to predict the number category of target objects in each partition (such as 0, 1, 2, etc.), where 0 means no target, and 1 - 10 represent that the number of target objects is between 1 and 10. The loss is calculated based on the difference between the estimated result of the number in each partition and the true label: if the regression method is used, the mean squared error (MSE) loss is usually adopted. If the classification method is used, the cross-entropy loss is usually adopted, and the model parameters are updated through the backpropagation algorithm to minimize the loss function. Through feature map partitioning and number estimation, the model can learn the distribution and number information of target objects, thereby improving the accuracy of object detection.
[0153] In another embodiment, the original feature map and the discrete number estimation feature map are fused to obtain a hybrid feature map, including:
[0154] The original feature map is processed by depthwise separable convolution to obtain two third feature maps and fourth feature maps of the same size; the discrete number estimation feature map, the third feature map, and the fourth feature map are reshaped to feature maps of the same size to obtain the reshaped discrete number estimation feature map, the reshaped third feature map, and the reshaped fourth feature map; the reshaped third feature map, the reshaped fourth feature map, and the reshaped discrete number estimation feature map are fused to obtain a hybrid feature map.
[0155] Specifically, when fusing the original feature map and the discrete number estimation feature map, the original feature map is first passed through two depthwise separable convolution modules to generate a third feature map and a fourth feature map respectively, where the third feature map and the fourth feature map are of the same size. Then, the size of the discrete number estimation feature map is reshaped to be the same as that of the third feature map and the fourth feature Figure 1Perform the following. Fuse the reshaped third feature map, the reshaped fourth feature map, and the reshaped discrete quantity estimation feature map. The fusion process can be achieved through matrix multiplication, attention mechanism, or other operations to combine the information of these feature maps, resulting in a hybrid feature map that contains the information of the original feature map and the discrete quantity estimation feature map.
[0156] In some embodiments, fusing the reshaped third feature map, the reshaped fourth feature map, and the reshaped discrete quantity estimation feature map to obtain a hybrid feature map includes: transposing the reshaped discrete quantity estimation feature map to obtain a transposed discrete quantity estimation feature map, multiplying the transposed discrete quantity estimation feature map by the reshaped third feature map and inputting the result into a classifier, calculating the distribution probability of the target object through the classifier to obtain a fifth feature map; performing matrix operation on the fifth feature map and the reshaped fourth feature map to obtain a sixth feature map; reshaping the size of the sixth feature map to the same size as the original feature map to obtain the hybrid feature map.
[0157] Specifically, Figure 7 is a schematic diagram of the feature mixing module structure of the object detection method in this embodiment. As Figure 7 shown, taking the original feature map P1 and the discrete quantity estimation feature map D1 as examples, the size of the original feature map P1 is 256×H / 8×W / 8, where H is the height of the feature map and W is the width of the feature map. The feature map P1 passes through two depthwise separable convolution (DWConv(256, 1, 1)) modules to obtain two feature maps with the same size. Then, perform a reshaping operation on the two feature maps to merge the width and height into one dimension, obtaining the reshaped third feature map K and the reshaped fourth feature map V, with a size of 256×H / 8×W / 8. The discrete quantity estimation feature map D1 undergoes a reshaping operation to generate the reshaped discrete quantity estimation feature map Q, with a size of 256×H / 8×W / 8. Transpose the reshaped discrete quantity estimation feature map Q to obtain a transposed discrete quantity estimation feature map, multiply the transposed discrete quantity estimation feature map by the reshaped third feature map K to obtain an intermediate result, apply the Softmax operation to this intermediate result to obtain the fifth feature map A, with a size of (H / 8×W / 8)×(H / 8×W / 8). Perform matrix multiplication operation on the fifth feature map A and the reshaped fourth feature map V to obtain a weighted feature map, the sixth feature map E. Reshape the sixth feature map through a reshape operation to restore the sixth feature map to the size of the original feature map P1, that is, 256×(H / 8)×(W / 8), to obtain the hybrid feature map E1, with a size of 256×H / 8×W / 8.
[0158] In another embodiment, the target detection model training process includes: obtaining a mixed feature map of the training images; pre-labeling the target objects in the training images to obtain an initial target object labeling result; and training the initial target detection model for the target object detection task based on the initial target object labeling result and the mixed feature map of the training images to obtain a trained target detection model.
[0159] Specifically, according to the method for extracting the mixed feature map in the above embodiment, extract the mixed feature map from the training images, and pre-label the target objects in the training images first. The labeling includes the category and bounding box of the target object, and the pre-labeling can be done manually or using a pre-trained model or an automated tool. Based on the requirements of the target detection task, construct an initial target detection model, and train the initial target detection model based on the initial target object labeling result and the mixed feature map of the training images. During the training process, the model learns how to identify and locate the target objects from the mixed feature map. Each point in each mixed feature map predicts the position and category of a target box. The position is a regression task, and the category is a classification task. Calculate the loss function (such as cross-entropy loss, IoU loss, etc.) based on the pre-labeling result and the prediction result of the initial target detection model to measure the difference between the prediction result of the initial target detection model and the true labeling, and optimize the parameters of the initial target detection model to minimize the loss. Evaluate the performance of the initial target detection model, and adjust and optimize the initial target detection model according to the evaluation result to obtain a trained target detection model.
[0160] In this embodiment, a target detection method is also provided. Figure 8 It is a flowchart of another target detection method in this embodiment, as Figure 8 shown. The process includes the following steps:
[0161] Step S801, extract an initial feature map from the image to be detected through a preset convolutional network model; perform downsampling and then upsampling on the initial feature map; add the feature maps with the same size and the same number of channels during the downsampling process and the upsampling process to obtain the original feature map of the image to be detected;
[0162] Step S802, based on the trained discrete quantity estimation model, perform discrete quantity estimation on the original feature map and output a discrete quantity estimation feature map;
[0163] Step S803, perform depthwise separable convolution processing on the original feature map to obtain two third feature maps and fourth feature maps with the same size;
[0164] Step S804: Reshape the discrete quantity estimation feature map, the third feature map, and the fourth feature map to feature maps of the same size, obtaining the reshaped discrete quantity estimation feature map, the reshaped third feature map, and the reshaped fourth feature map.
[0165] Step S805: Transpose the reshaped discrete quantity estimation feature map to obtain the transposed discrete quantity estimation feature map. Multiply the transposed discrete quantity estimation feature map by the reshaped third feature map and input the result into a classifier. Calculate the distribution probability of the target object through the classifier to obtain a fifth feature map.
[0166] Step S806: Perform matrix operations on the fifth feature map and the reshaped fourth feature map to obtain a sixth feature map.
[0167] Step S807: Reshape the size of the sixth feature map to the same size as the original feature map to obtain a mixed feature map.
[0168] Step S808: Input the mixed feature map into the trained target detection model for target object detection to obtain the target object detection result.
[0169] Through the above steps S801 to S808, compared with the density map or anchor point method predicted by the model in the prior art, in this embodiment, different-sized original image features of the target object in the image to be detected are extracted; different-sized original image features are combined to train a discrete quantity estimation model, and the original image features are input into the trained discrete quantity estimation model to generate a discrete quantity estimation feature map; through a mixed feature module, the information of the discrete quantity estimation feature map and the original feature map is fused to generate a mixed feature map; a target detection task is performed based on the mixed feature map; among them, by adding a discrete quantity estimation model, a mixed feature map is obtained by combining the discrete quantity estimation feature map and the original feature map, and the detection of the target object is performed based on the mixed feature map, improving the accuracy of target detection.
[0170] In this embodiment, a target detection device is further provided. This device is used to implement the above embodiment and preferred implementation manners, and those that have been described will not be repeated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the device described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0171] Figure 9 is the structural block diagram of the target detection device in this embodiment. As Figure 9 shown, the device 90 includes: an original feature extraction module 91, a discrete quantity estimation model 92, a mixed feature module 93, and a target detection module 94, where
[0172] The original feature extraction module 91 is used to extract the image features of the image to be detected, and obtain the original feature map of the image to be detected, where the original feature map includes feature maps of different sizes;
[0173] The discrete quantity estimation model 92 is used to perform discrete quantity estimation on the original feature map and output a discrete quantity estimation feature map;
[0174] The hybrid feature module 93 is used to fuse the original feature map and the discrete quantity estimation feature map to obtain a hybrid feature map;
[0175] The target detection module 94 is used to input the hybrid feature map into the trained target detection model for target object detection to obtain the target object detection result.
[0176] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.
[0177] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0178] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0179] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:
[0180] S1. Extract the image features of the image to be detected to obtain the original feature map of the image to be detected, where the original feature map includes feature maps of different sizes;
[0181] S2. Based on the trained discrete quantity estimation model, perform discrete quantity estimation on the original feature map and output a discrete quantity estimation feature map;
[0182] S3. Fuse the original feature map and the discrete quantity estimation feature map to obtain a hybrid feature map;
[0183] S4. Input the hybrid feature map into the trained target detection model for target object detection to obtain the target object detection result.
[0184] It should be noted that specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0185] In addition, in combination with the object detection method provided in the above embodiments, a storage medium can also be provided in this embodiment to implement it. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the object detection methods in the above embodiments is implemented.
[0186] It should be understood that the specific embodiments described herein are only used to explain this application, rather than to limit it. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0187] Obviously, the accompanying drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative efforts. Additionally, it can be understood that although the work done during the development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application.
[0188] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0189] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0190] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A target detection method, characterized in that, Including: Extracting image features of the image to be detected to obtain an original feature map of the image to be detected, where the original feature map includes feature maps of different sizes; Based on the trained discrete quantity estimation model, performing discrete quantity estimation on the original feature map, dividing continuous feature data into discrete intervals, converting the classification task of continuous quantity into a discrete classification task, evaluating each feature value in the original feature map to obtain an evaluation result of each feature value in the original feature map, and mapping the evaluation result to a preset discrete interval or category to output a discrete quantity estimation feature map; Performing depthwise separable convolution processing on the original feature map to obtain two third feature maps and fourth feature maps with the same size; Reshaping the sizes of the discrete quantity estimation feature map, the third feature map, and the fourth feature map into feature maps of the same size to obtain a reshaped discrete quantity estimation feature map, a reshaped third feature map, and a reshaped fourth feature map; Transposing the reshaped discrete quantity estimation feature map to obtain a transposed discrete quantity estimation feature map, multiplying the transposed discrete quantity estimation feature map by the reshaped third feature map and inputting the result into a classifier, calculating the distribution probability of the target object through the classifier to obtain a fifth feature map; performing matrix operation on the fifth feature map and the reshaped fourth feature map to obtain a sixth feature map; reshaping the size of the sixth feature map to be the same as that of the original feature map to obtain a mixed feature map; Inputting the mixed feature map into the trained target detection model for target object detection to obtain a target object detection result.
2. The object detection method according to claim 1, wherein The extracting image features of the image to be detected to obtain an original feature map of the image to be detected includes: Extracting an initial feature map from the image to be detected through a preset convolutional network model; Performing downsampling and then upsampling on the initial feature map; Adding the feature maps with the same size and the same number of channels during the downsampling process and the upsampling process to obtain the original feature map of the image to be detected.
3. The object detection method according to claim 2, wherein During the downsampling process of the initial feature map, the method further includes: Inputting the initial feature map into the first convolutional module and the bottleneck module of the preset convolutional network model to output a first feature map of the image to be detected; Inputting the initial feature map into the second convolutional module of the preset convolutional network model to output a second feature map of the image to be detected; Concatenating the first feature map and the second feature map to obtain a downsampled initial feature map.
4. The object detection method according to claim 1, wherein The training process of the discrete quantity estimation model includes: Training an initial discrete quantity estimation model according to depthwise separable convolution and a classifier to obtain a trained discrete quantity estimation model; Wherein, in the classifier, training the initial discrete quantity estimation model includes: After partitioning the discrete quantity estimation feature map, perform quantity estimation on the target objects in the discrete quantity estimation feature map to obtain the quantity estimation results of each target object, and train the initial discrete quantity estimation model according to the quantity estimation results of each target object.
5. The object detection method according to claim 1, wherein The training process of the target detection model includes: Obtain the mixed feature map of the training image; Perform pre-annotation on the target objects in the training image to obtain the initial target object annotation results; Based on the initial target object annotation results and the mixed feature map of the training image, train the initial target detection model for the target object detection task to obtain the trained target detection model.
6. A target detection device, characterized in that, It includes: An original feature extraction module, a discrete quantity estimation model, a mixed feature module, and a target detection module, where The original feature extraction module is used to extract the image features of the image to be detected to obtain the original feature map of the image to be detected, where the original feature map includes feature maps of different sizes; The discrete quantity estimation model is used to perform discrete quantity estimation on the original feature map, divide the continuous feature data into discrete intervals, convert the classification task of continuous quantities into a discrete classification task, evaluate each feature value in the original feature map to obtain the evaluation result of each feature value in the original feature map, and map the evaluation result to a preset discrete interval or category, and output a discrete quantity estimation feature map; The mixed feature module is used to perform depthwise separable convolution processing on the original feature map to obtain two third feature maps and fourth feature maps of the same size; Reshape the sizes of the discrete quantity estimation feature map, the third feature map, and the fourth feature map into feature maps of the same size to obtain the reshaped discrete quantity estimation feature map, the reshaped third feature map, and the reshaped fourth feature map; Transpose the reshaped discrete quantity estimation feature map to obtain the transposed discrete quantity estimation feature map, multiply the transposed discrete quantity estimation feature map by the reshaped third feature map and input it into a classifier, calculate the distribution probability of the target object through the classifier to obtain a fifth feature map; perform matrix operation on the fifth feature map and the reshaped fourth feature map to obtain a sixth feature map; reshape the size of the sixth feature map to be the same as that of the original feature map to obtain a mixed feature map; The target detection module is used to input the mixed feature map into the trained target detection model for target object detection to obtain the target object detection result.
7. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is set to run the computer program to execute the target detection method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the target detection method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for estimating number of image recognition objects and storage medium
CN114973115A