Small target detection method and device, electronic equipment, storage medium and program product
By introducing asymmetric convolution modules and multi-scale feature fusion networks into convolutional neural networks, the feature extraction capability for small target detection is enhanced, solving the problem of insufficient accuracy in small target detection in existing technologies and achieving higher detection accuracy.
Patent Information
- Application Number
- CN202511446903.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Existing convolutional neural networks have insufficient feature extraction capabilities when detecting small objects, resulting in poor accuracy of detection results.
Feature extraction is performed through the front-end layer of the backbone network. Asymmetric convolution modules are used to enhance the directional details of the initial shallow feature maps. A multi-scale feature fusion network is then used to weight and fuse the shallow and back-end feature maps to generate a multi-scale fused feature map. Finally, a detection head is used to detect small targets.
It significantly improves the accuracy of small target detection. Through the synergistic mechanism of directional enhancement and multi-scale fusion, it enhances the expressive power and discriminative power of small target features.
Smart Images

Figure CN120912875A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a small target detection method and device, electronic equipment, storage medium and program product. BACKGROUND
[0002] Target detection is a core task in the field of computer vision, which aims to identify specific class of object instances in images or videos and determine their location and size. In recent years, with the development of deep learning technology, target detection algorithms based on convolutional neural networks (CNN) have become the mainstream solution.
[0003] In the existing technical solution, target detection usually includes two main stages: feature extraction and target positioning / classification, and the performance of the feature extraction stage has a decisive influence on the accuracy of the entire detection system. The current mainstream structural framework, such as ResNet (Residual Network) series, ShuffleNet (Shuffle Network) series, etc., performs well in processing medium or large scale targets, but its extraction ability significantly decreases when facing small target detection, often failing to capture effective representation information, resulting in poor accuracy of small target detection results.
[0004] Therefore, how to effectively improve the accuracy of small target detection is a problem that needs to be solved at present. SUMMARY
[0005] The present application provides a small target detection method, device, electronic equipment, storage medium and program product to improve the accuracy of small target detection results.
[0006] The present application provides a small target detection method, comprising: extracting features of a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map; processing the initial shallow feature map through an asymmetric convolution module to obtain a shallow enhanced feature map; inputting the shallow enhanced feature map into a back-end layer of the backbone network for further propagation to obtain back-end layer feature maps output by multiple network stages; performing weighted fusion on the shallow enhanced feature map and the back-end layer feature maps through a multi-scale feature fusion network to obtain a multi-scale fusion feature map; detecting the multi-scale fusion feature map through a detection head to obtain a small target detection result.
[0007] According to the small target detection method provided by the application, the initial shallow feature map is obtained by performing feature extraction on the to-be-detected image through the front-end layer of the backbone network, and the method comprises the following steps: The initial feature map is obtained by performing feature extraction on the to-be-detected image through the initial convolution layer of the backbone network. The initial shallow feature map is obtained by performing feature extraction on the initial feature map through the first network stage of the backbone network; wherein the first network stage is the first network stage after the first spatial down-sampling is completed in the backbone network.
[0008] According to the small target detection method provided by the application, the multi-scale fusion feature map is obtained by performing weighted fusion on the shallow enhanced feature map and the back-end layer feature map through the multi-scale feature fusion network, and the method comprises the following steps: The back-end layer feature map is up-sampled to have the same spatial size as the shallow enhanced feature map. The channel of the shallow enhanced feature map and the back-end layer feature map after up-sampling is adjusted through a 1*1 convolution layer to obtain a to-be-fused shallow enhanced feature map and a to-be-fused back-end layer feature map; the to-be-fused shallow enhanced feature map and the to-be-fused back-end layer feature map have the same number of channels. The adaptive weight is generated by the attention module according to the to-be-fused shallow enhanced feature map and the to-be-fused back-end layer feature map. The to-be-fused shallow enhanced feature map and the to-be-fused back-end layer feature map are weighted and summed according to the adaptive weight to obtain the multi-scale fusion feature map.
[0009] According to the small target detection method provided by the application, the multi-scale fusion feature map is obtained by performing weighted fusion on the shallow enhanced feature map and the back-end layer feature map through the multi-scale feature fusion network, and the method further comprises the following steps: The feature map output by the last two network stages in the back-end layer is selected from the back-end layer feature map as a target scale feature map. The multi-scale fusion feature map is obtained by performing weighted fusion on the shallow enhanced feature map and the target scale feature map through the multi-scale feature fusion network based on the attention mechanism.
[0010] According to the small target detection method provided by the application, the shallow enhanced feature map is obtained by processing the initial shallow feature map through the asymmetric convolution module, and the method comprises the following steps: A horizontal convolution kernel is applied to the initial shallow feature map to perform convolution operation and generate a horizontal direction feature map. A vertical convolution kernel is applied to the initial shallow feature map to perform convolution operation and generate a vertical direction feature map. The horizontal direction feature map and the vertical direction feature map are added and fused to obtain the shallow layer enhanced feature map.
[0011] According to the small target detection method provided by the application, the horizontal convolution kernel is a 1*m convolution kernel, and the vertical convolution kernel is an m*1 convolution kernel, wherein m is an integer greater than 1.
[0012] The application further provides a small target detection device, comprising: The feature extraction module is configured to extract features of the to-be-detected image through a front-end layer of the backbone network to obtain an initial shallow layer feature map. The first processing module is configured to process the initial shallow layer feature map through the asymmetric convolution module to obtain a shallow layer enhanced feature map. The second processing module is configured to input the shallow layer enhanced feature map into a back-end layer of the backbone network and continue to propagate forward to obtain back-end layer feature maps output by multiple network stages. The weighted fusion module is configured to perform weighted fusion on the shallow layer enhanced feature map and the back-end layer feature maps through a multi-scale feature fusion network to obtain multi-scale fusion feature maps. The small target detection module is configured to detect the multi-scale fusion feature maps through a detection head to obtain a small target detection result.
[0013] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the small target detection method according to any one of the above when executing the computer program.
[0014] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the small target detection method according to any one of the above.
[0015] The application further provides a computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the small target detection method according to any one of the above.
[0016] The small target detection method, device, electronic equipment, storage medium and program product provided by the application first perform feature extraction on a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map. The initial shallow feature map contains rich detail information while retaining a high spatial resolution. Then, the initial shallow feature map is processed through an asymmetric convolution module to enhance the directional details therein, and a shallow enhanced feature map is obtained, thereby significantly strengthening the detail features that are crucial to small target detection. Next, the shallow enhanced feature map is input into a back-end layer of the backbone network for continued propagation to obtain back-end layer feature maps output by multiple network stages. The shallow enhanced feature map and the back-end layer feature maps are weighted and fused through a multi-scale feature fusion network to obtain a multi-scale fused feature map. The fusion process effectively retains key semantics and detail information of different scales, enabling the final feature map to have both high-resolution details and high-level semantics. Finally, the multi-scale fused feature map is detected through a detection head to obtain a small target detection result. Through the cooperative mechanism of directional enhancement and multi-scale fusion, the asymmetric convolution and multi-layer feature fusion technologies are organically combined to improve the expression ability and discriminativeness of small target features, thereby significantly improving the accuracy of the small target detection result. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0018] Figure 1 is one of the flowcharts of the small target detection method provided by the application; Figure 2 is another flowchart of the small target detection method provided by the application; Figure 3 is a third flowchart of the small target detection method provided by the application; Figure 4 is a fourth flowchart of the small target detection method provided by the application; Figure 5 is a fifth flowchart of the small target detection method provided by the application; Figure 6 is a structural schematic diagram of the small target detection device provided by the application; Figure 7 is a structural schematic diagram of the electronic equipment provided by the application. DETAILED DESCRIPTION
[0019] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0020] Target detection is a core task in the field of computer vision, which aims to identify specific class object instances in images or videos and determine their location and size. In recent years, with the development of deep learning technology, target detection algorithms based on convolutional neural networks have become the mainstream solution.
[0021] In the existing technical solutions, target detection usually includes two main stages: feature extraction and target positioning / classification, and the performance of the feature extraction stage has a decisive influence on the accuracy of the entire detection system. The current mainstream structural framework, such as the ResNet series and the ShuffleNet series, performs excellently in processing medium or large scale targets, but its extraction ability significantly decreases when facing small target detection, often failing to capture effective feature information, resulting in poor accuracy of small target detection results.
[0022] Therefore, how to effectively improve the accuracy of small target detection is a problem that needs to be solved at present.
[0023] Through analysis, these mainstream structural frameworks generally adopt a single and symmetric deep network structure, which has two key limitations: First, the directional feature extraction ability is insufficient. Traditional convolution operations use square convolution kernels (such as 3x3 or 5x5), which use symmetric weight distribution in horizontal and vertical directions. This uniform attention to global features makes it difficult for the network to extract features with directional characteristics, such as horizontally oriented road markings or vertically distributed building edges. For small targets with limited pixel quantity, this symmetric design weakens their key discriminative features.
[0024] Second, the limitation of single-scale feature extraction. The existing single network architecture adopts a serial feature extraction mode, and as the network depth increases, the receptive field gradually expands. This design causes the effective features of small targets to be diluted in the deep network, for example, a 10x10 pixel target only occupies a very small proportion of the receptive field in the deep network.
[0025] Based on the above analysis, the present application proposes a small target detection method, device, electronic equipment, storage medium and program product, which will be described below in combination with Figures 1-7 .
[0026] Figure 1is one of the flowcharts of the small target detection method provided by the present application, as shown in Figure 1 The small target detection method includes steps S110, S120, S130, S140 and S150.
[0027] In step S110, feature extraction is performed on the to-be-detected image through the front-end layer of the backbone network to obtain an initial shallow feature map.
[0028] The backbone network can adopt a classical convolutional neural network structure, for example, ResNet series (such as ResNet-18, ResNet-50, ResNet-101), ShuffleNet series (such as ShuffleNet V1, ShuffleNet V2), GhostNet series, etc.
[0029] The front-end layer refers to the network layer close to the input image and located in the front part of the backbone network, which is the first layer passed by the to-be-detected image after entering the network. Specifically, the front-end layer can include an initial convolutional layer and a first network stage, or an initial convolutional layer, a first network stage and a second network stage. The first network stage is the first network stage after the first spatial down-sampling in the backbone network, and the second network stage is the next network stage of the first network stage.
[0030] Taking ResNet-18 as the backbone network as an example, the ResNet-18 architecture includes a stem stage (initial convolutional layer), stage1 (stage 1, also referred to as the first residual stage), stage2, stage3 and stage4. The initial convolutional layer corresponds to the stem stage, the first spatial down-sampling is completed by the stem stage, the first network stage corresponds to stage1, and the second network stage corresponds to stage2.
[0031] Taking ShuffleNet V2 as the backbone network as an example, the ShuffleNet V2 architecture includes Head (head, also referred to as initial convolutional layer), stage2, stage3 and stage4. The initial convolutional layer corresponds to Head, the first spatial down-sampling is completed by Head, the first network stage corresponds to stage2, and the second network stage corresponds to stage3.
[0032] The initial shallow feature map has a significantly reduced size compared to the original to-be-detected image, but still retains rich spatial position information, edges, textures and other low-level but crucial details for small target positioning.
[0033] In an embodiment, the initial convolutional layer of the backbone network is used to extract features of the image to be detected to obtain an initial feature map; and the first network stage of the backbone network is used to extract features of the initial feature map to obtain an initial shallow feature map; wherein the first network stage is a network stage in which the backbone network completes the first spatial down-sampling.
[0034] In another embodiment, the initial convolutional layer of the backbone network is used to extract features of the image to be detected to obtain an initial feature map; and the first network stage and the second network stage of the backbone network are used to extract features of the initial feature map in sequence to obtain an initial shallow feature map; wherein the first network stage is a network stage in which the backbone network completes the first spatial down-sampling, and the second network stage is a network stage subsequent to the first network stage.
[0035] It should be noted that the initial shallow feature map extracted by the first embodiment has a higher spatial resolution and richer bottom-level detail information than the initial shallow feature map extracted by the second embodiment, which can ensure that more rich detail sources can be obtained during subsequent direction enhancement and feature fusion, and which is more helpful to improve the accuracy of small target detection results.
[0036] In step S120, the initial shallow feature map is processed by the asymmetric convolution module to obtain a shallow enhanced feature map.
[0037] In order to enhance the expression ability of the shallow feature without significantly increasing the calculation cost, an asymmetric convolution module is introduced in the embodiment. The initial shallow feature map is processed by the asymmetric convolution module to obtain an enhanced shallow feature map, which is denoted as a shallow enhanced feature map.
[0038] The asymmetric convolution module includes two parallel branches: a horizontal convolution branch and a vertical convolution branch. In the horizontal convolution branch, a horizontal convolution kernel is applied to the shallow feature map to perform convolution operation to generate a horizontal direction feature map; in the vertical convolution branch, a vertical convolution kernel is applied to the shallow feature map to perform convolution operation to generate a vertical direction feature map; and the horizontal direction feature map and the vertical direction feature map are added and fused to obtain the shallow enhanced feature map. The specific execution process can be referred to the following embodiments, which will not be repeated here.
[0039] Further, the asymmetric convolution module adopts a double-channel output structure, the first half channel of the output shallow enhanced feature map is dedicated to storing horizontal direction features, and the second half channel is dedicated to storing vertical direction features, so as to maintain the independence of the direction features and provide a structured input basis for subsequent fusion.
[0040] Step S130, input the shallow layer enhanced feature map to the back-end layer of the backbone network, continue to forward propagate, and obtain the back-end layer feature maps output by multiple network stages.
[0041] The back-end layer refers to the network layer in the backbone network far away from the input image and located in the latter part of the backbone network. Specifically, it refers to the network layer in the backbone network other than the front-end layer.
[0042] The shallow layer enhanced feature map is input to the back-end layer of the backbone network, and continues to forward propagate to obtain the back-end layer feature maps output by multiple network stages.
[0043] Taking ResNet-18 as the backbone network as an example, the ResNet-18 architecture includes a stem stage, a stage1, a stage2, a stage3 and a stage4. When the front-end layer includes the stem stage and the stage1, the back-end layer includes the stage2, the stage3 and the stage4. Correspondingly, the back-end layer feature maps include the feature map processed by the stage2, the feature map processed by the stage3 and the feature map processed by the stage4.
[0044] Taking ShuffleNet V2 as the backbone network as an example, the ShuffleNet V2 architecture includes a Head, a stage2, a stage3 and a stage4. When the front-end layer includes the Head and the stage2, the back-end layer includes the stage3 and the stage4. Correspondingly, the back-end layer feature maps include the feature map processed by the stage3 and the feature map processed by the stage4.
[0045] Step S140, through the multi-scale feature fusion network, the shallow layer enhanced feature map and the back-end layer feature map are weighted fused to obtain a multi-scale fusion feature map.
[0046] The multi-scale feature fusion network includes multiple branches, each branch can independently process the features on the feature map of a specific scale. In this way, multiple scale feature maps can be input in parallel, the protection of features of each scale, especially the shallow layer features, can be realized, thereby effectively preventing the dilution of small target information.
[0047] In an embodiment, the back-end layer feature map is up-sampled to have the same spatial size as the shallow layer enhanced feature map; a 1x1 convolution layer is used to adjust the channel of the shallow layer enhanced feature map and the up-sampled back-end layer feature map respectively to obtain a to-be-fused shallow layer enhanced feature map and a to-be-fused back-end layer feature map; the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature map have the same number of channels; and the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature map are weighted and summed according to a preset weight to obtain a multi-scale fusion feature map.
[0048] In another embodiment, the feature maps output by the last two network stages in the back-end layer are selected from the back-end layer feature map as target scale feature maps; the target scale feature maps are up-sampled to have the same spatial size as the shallow layer enhanced feature map; a 1x1 convolution layer is used to adjust the channel of the shallow layer enhanced feature map and the up-sampled target scale feature map respectively to obtain a to-be-fused shallow layer enhanced feature map and a to-be-fused target scale feature map; the to-be-fused shallow layer enhanced feature map and the to-be-fused target scale feature map have the same number of channels; and the to-be-fused shallow layer enhanced feature map and the to-be-fused target scale feature map are weighted and summed according to a preset weight to obtain a multi-scale fusion feature map.
[0049] Further, in the weighted fusion, the weighted fusion is performed based on an attention mechanism, i.e., the shallow layer enhanced feature map and the back-end layer feature map are weighted and fused based on the attention mechanism by the multi-scale feature fusion network to obtain a multi-scale fusion feature map.
[0050] Specifically, in an embodiment, the back-end layer feature map is up-sampled to have the same spatial size as the shallow layer enhanced feature map; a 1x1 convolution layer is used to adjust the channel of the shallow layer enhanced feature map and the up-sampled back-end layer feature map respectively to obtain a to-be-fused shallow layer enhanced feature map and a to-be-fused back-end layer feature map; the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature map have the same number of channels; an adaptive weight is generated by an attention module according to the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature map; and the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature map are weighted and summed according to the adaptive weight to obtain a multi-scale fusion feature map. The specific execution process can refer to the following embodiments, which will not be described here.
[0051] In another implementation, feature maps output from the last two network stages in the backend layer are selected from the backend layer feature maps as target-scale feature maps. The target-scale feature maps are upsampled to have the same spatial size as the shallow enhancement feature maps. Channel adjustments are applied to both the shallow enhancement feature maps and the upsampled target-scale feature maps using 1×1 convolutional layers to obtain the shallow enhancement feature maps to be fused and the target-scale feature maps to be fused. The shallow enhancement feature maps to be fused and the target-scale feature maps to be fused have the same number of channels. An attention module generates adaptive weights based on the shallow enhancement feature maps to be fused and the target-scale feature maps to be fused. Based on the adaptive weights, a weighted sum is performed on the shallow enhancement feature maps to be fused and the target-scale feature maps to be fused to obtain a multi-scale fused feature map. The specific execution process can be referred to in the following embodiments, and will not be elaborated here.
[0052] Step S150: The multi-scale fused feature map is detected by the detection head to obtain the small target detection result.
[0053] The multi-scale fused feature map obtained above is input into the detection head to detect it and obtain the small target detection result.
[0054] The small target detection method provided in this invention first extracts features from the image to be detected through the front layer of the backbone network to obtain an initial shallow feature map. This initial shallow feature map retains high spatial resolution while containing rich detail information. Then, an asymmetric convolution module processes the initial shallow feature map to enhance its directional details, resulting in a shallow enhanced feature map, which significantly strengthens the detail features crucial for small target detection. Next, the shallow enhanced feature map is input into the back layer of the backbone network and continues to propagate forward, obtaining back layer feature maps output from multiple network stages. A multi-scale feature fusion network then performs weighted fusion of the shallow enhanced feature map and the back layer feature map to obtain a multi-scale fused feature map. This fusion process effectively preserves key semantic and detail information at different scales, enabling the final feature map to possess both high-resolution details and high-level semantics. Finally, a detection head performs detection on the multi-scale fused feature map to obtain the small target detection result. The embodiments of the present invention utilize a synergistic mechanism of directional enhancement and multi-scale fusion, organically combining asymmetric convolution and multi-layer feature fusion techniques to jointly enhance the expressive power and discriminative power of small target features, thereby significantly improving the accuracy of small target detection results.
[0055] Based on any of the above embodiments Figure 2 This is the second flowchart of the small target detection method provided by the present invention, as shown below. Figure 2 As shown, step S110 includes step S111 and step S112.
[0056] Step S111, performing feature extraction on the to-be-detected image through an initial convolutional layer of the backbone network to obtain an initial feature map.
[0057] The initial convolutional layer is used to perform preliminary feature extraction and substantial spatial down-sampling on the to-be-detected image, so as to reduce subsequent calculation amount and expand the receptive field.
[0058] Taking ResNet-18 as the backbone network as an example, in the ResNet-18 architecture, the initial convolutional layer is composed of a 7x7 convolutional layer (with a stride of 2) and a 3x3 max-pooling layer (with a stride of 2) in series. After processing an input to-be-detected image with a size of 224x224 through the initial convolutional layer, the spatial size is down-sampled to 56x56, and an initial feature map with a channel number of 64 is output.
[0059] Taking ShuffleNet V2 as the backbone network as an example, in the ShuffleNet V2 architecture, the initial convolutional layer is composed of a 3x3 convolutional layer (with a stride of 2) and a 3x3 max-pooling layer (with a stride of 2) in series. After processing an input to-be-detected image with a size of 224x224 through the initial convolutional layer, the spatial size is down-sampled to 56x56, and an initial feature map with a channel number of 24 is output.
[0060] Step S112, performing feature extraction on the initial feature map through a first network stage of the backbone network to obtain the initial shallow feature map; wherein the first network stage is the first network stage after the first spatial down-sampling in the backbone network.
[0061] The initial feature map is input into the first network stage of the backbone network, and the initial feature map is subjected to feature extraction through the first network stage to obtain the initial shallow feature map.
[0062] The first network stage is the first network stage after the first spatial down-sampling in the backbone network. For example, in the ResNet-18 architecture, the first network stage corresponds to the first group of residual modules (stage1), which specifically includes two basic residual blocks, each of which has two layers of 3x3 convolution, and the output channel number is 64. For another example, in the ShuffleNet V2 architecture, the first network stage corresponds to stage2, and specifically, the stage2 includes four sequentially connected ShuffleNet basic blocks (Block). Among them, the first basic block uses a depth convolution with a stride of 2 to realize spatial down-sampling, and the subsequent three basic blocks have a stride of 1 to keep the feature map size unchanged. Each basic block performs channel partitioning, grouped convolution, channel shuffling, and other operations, and finally outputs a feature map with a channel number of 244.
[0063] Taking ResNet-18 as the backbone network as an example, after obtaining the initial feature map of 56x56, the initial feature map is input into the stage1 of ResNet-18, the spatial size of the input and output is kept unchanged (still 56x56), and feature extraction is performed, and finally a feature map with a channel number of 64 is output, which is the initial shallow feature map.
[0064] Taking ShuffleNet V2 as the backbone network as an example, after obtaining the initial feature map of 56x56, the initial feature map is first input into the first basic block (stride=2) in the stage2 of ShuffleNet V2, and the spatial size is down-sampled to 28x28. Then, the subsequent three basic blocks are processed, and the size of 28x28 is kept unchanged, and finally a feature map with a channel number of 244 is output, which is the initial shallow feature map.
[0065] It should be noted that the output feature map of the first network stage is selected as the initial shallow feature map because it retains a very high spatial resolution while having undergone preliminary semantic information extraction, and contains rich low-level visual information such as edges, textures, and corner points. As the source of subsequent direction enhancement and feature fusion, it can ensure that these important information for small target detection can be effectively protected and utilized, thereby significantly overcoming the problem of detail loss caused by deepening of the network, and improving the recall rate and positioning accuracy of small targets.
[0066] The small target detection method provided by the embodiment of the application takes the feature map obtained through the initial convolution layer and the first network stage of the backbone network as the initial shallow feature map, which retains the maximum spatial resolution and the most original detail information, and takes it as the source of subsequent direction enhancement and feature fusion, which can fundamentally solve the problem of dilution or loss of small target information caused by deepening of the network.
[0067] Based on any of the above embodiments, Figure 3 is a flowchart of the third small target detection method provided by the present application, as shown in the figure, step S140 includes step S141, step S142, step S143 and step S144. Figure 3
[0068] Step S141, up-sampling the back-end layer feature map to the same spatial size as the shallow layer enhanced feature map.
[0069] Step S142, adjusting the channels of the shallow layer enhanced feature map and the up-sampled back-end layer feature map respectively through a 1x1 convolution layer to obtain a to-be-fused shallow layer enhanced feature map and a to-be-fused back-end layer feature map; the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature map have the same number of channels.
[0070] Considering that although the existing standard feature pyramid attempts to fuse multi-scale features, the one-way fusion method still causes deep semantic information to be submerged in fine spatial details of the shallow layer. Therefore, in this embodiment, a progressive feature fusion architecture is provided, multiple branches are set, each branch can independently process features on a specific scale feature map, and the weights of each branch are automatically learned based on an attention mechanism to realize dynamic weighted fusion, so as to protect small target information from being lost while realizing multi-scale feature complementary enhancement, thereby further improving the accuracy of small target detection results.
[0071] First, a common target size and channel number are selected to process the feature maps of each branch. Specifically, the size of the highest resolution shallow layer enhanced feature map can be selected as the reference to maximize the preservation of detail information. The channel number can be selected as 256, which can balance the performance and expression ability and the computational feasibility.
[0072] Since the spatial size of the back-end layer feature map is small, the back-end layer feature map is up-sampled to make the up-sampled back-end layer feature map have the same spatial size as the shallow layer enhanced feature map. Then, the channels of the shallow layer enhanced feature map and the up-sampled back-end layer feature map are adjusted through a 1x1 convolution layer to make the channel numbers of the feature maps the same, so as to facilitate subsequent fusion. Adjusting the feature dimension through a 1x1 convolution avoids distortion in cross-layer fusion.
[0073] Taking ResNet-18 as the backbone network as an example, the spatial size of a shallow layer enhanced feature map (denoted as C1) is 56x56, and the number of channels is 64. After processing by the back-end layers stage2, stage3 and stage4, a plurality of back-end layer feature maps are obtained, which are denoted as C2, C3 and C4 respectively, wherein the spatial size of C2 is 28x28, and the number of channels is 128; the spatial size of C3 is 14x14, and the number of channels is 256; the spatial size of C4 is 7x7, and the number of channels is 512.
[0074] Correspondingly, at this time, the multi-scale feature fusion network includes four branches, denoted as a first branch, a second branch, a third branch and a fourth branch.
[0075] The C1 is adjusted in the number of channels through the first branch to obtain a to-be-fused shallow layer enhanced feature map (denoted as F1), and when the number of channels is adjusted, the number of channels is adjusted to 256 using a 1x1 convolution, so that the spatial size of F1 is 56x56, and the number of channels is 256.
[0076] The C2 is up-sampled and adjusted in the number of channels through the second branch to obtain a to-be-fused feature map (denoted as F2), and when up-sampling, 2 times up-sampling is used, and then the number of channels is adjusted to 256 using a 1x1 convolution, so that the spatial size of F2 is 56x56, and the number of channels is 256.
[0077] The C3 is up-sampled through the third branch to obtain a to-be-fused feature map (denoted as F3), and when up-sampling, 4 times up-sampling is used, so that the spatial size of F3 is 56x56, and the number of channels is 256.
[0078] The C4 is up-sampled and adjusted in the number of channels through the fourth branch to obtain a to-be-fused feature map (denoted as F4), and when up-sampling, 8 times up-sampling is used, and then the number of channels is adjusted to 256 using a 1x1 convolution, so that the spatial size of F4 is 56x56, and the number of channels is 256.
[0079] At this time, the to-be-fused shallow layer enhanced feature map is F1, and the to-be-fused back-end layer feature maps include the above F2, F3 and F4.
[0080] In step S143, an attention module is used to generate adaptive weights according to the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature maps.
[0081] The to-be-fused shallow layer enhanced feature map and the plurality of to-be-fused back-end layer feature maps are input into the attention module to generate adaptive weights corresponding to each feature map.
[0082] Specifically, the element addition is first performed on the to-be-fused shallow layer enhanced feature map and the plurality of to-be-fused back-end layer feature maps to obtain an aggregated feature map, denoted as F', in the above example, F' = F1+F2+F3+F4, the spatial size of F' is 56x56, and the channel number is 256.
[0083] Global average pooling is applied to the aggregated feature map to compress the global information, which is then input into a fully connected neural network and normalized by a Softmax function to obtain adaptive weights corresponding to each feature map, denoted as α1, α2, α3 and α4, respectively.
[0084] In step S144, the adaptive weights are used to perform weighted summation on the to-be-fused shallow layer enhanced feature map and the to-be-fused back-end layer feature maps to obtain the multi-scale fusion feature map.
[0085] The adaptive weights corresponding to each feature map are used to perform weighted summation on each feature map to obtain a multi-scale fusion feature map.
[0086] In the above example, the multi-scale fusion feature map F = (α1xF1) + (α2xF2) + (α3xF3) + (α4xF4).
[0087] The small target detection method provided by the embodiments of the present application performs spatial size and channel number alignment processing on the shallow layer enhanced feature map and the plurality of back-end layer feature maps through each branch of the multi-scale feature fusion network. In the processing process, the spatial size of the shallow layer enhanced feature map with the highest resolution is selected as the reference, and spatial scale and channel matching are realized through upsampling and 1x1 convolution to maximize the retention of the detailed information which is crucial for small target detection. Then, based on the attention mechanism, the weights of each branch are automatically learned to realize dynamic weighted fusion, which can automatically increase the weight of the feature map with strong discrimination for small targets. Based on the asymmetric convolution processing of the shallow layer feature to realize small target information enhancement extraction, the embodiments of the present application combine the above multi-scale alignment and attention weighted fusion mechanism to form a closed-loop optimization system of "extraction-protection-enhancement": directional sensitive feature extraction is realized through asymmetric convolution, detailed information is protected without loss through scale alignment and feature fusion, and finally adaptive enhancement is realized in the fusion process through the attention mechanism. While effectively maintaining the integrity of the small target structure, the system fully fuses deep semantic and shallow details, significantly improves the discrimination ability of multi-scale feature expression, and further improves the accuracy of small target detection results.
[0088] Based on any of the above embodiments, Figure 4 is a fourth flowchart of the small target detection method provided by the present application, as shown in Figure 4 step S140 further includes steps S145 and S146.
[0089] Step S145, filtering the feature maps output by the last two network stages in the backend layer from the backend layer feature map as target scale feature maps.
[0090] Considering different network architectures, the backend layer includes different numbers of network stages, and considering the accuracy and efficiency, in the embodiment of the application, the feature maps output by the last two network stages in the backend layer are selected as target scale feature maps. The two target scale feature maps have large receptive fields and can represent global and high-level information of the target, and then are dynamically weighted and fused with the shallow layer enhanced feature maps, which can provide a feature representation with high resolution details and high-level semantics for small target detection, effectively solve the core problems of small target missing detection and inaccurate positioning, and improve the accuracy of small target detection results.
[0091] Taking the ResNet-18 architecture as an example, the last two network stages of the backend layer are Stage3 and Stage4, and the feature maps of the two network stages are selected as target scale feature maps, that is, C3 and C4.
[0092] Step S146, performing attention mechanism-based weighted fusion on the shallow layer enhanced feature map and the target scale feature map through a multi-scale feature fusion network to obtain a multi-scale fusion feature map.
[0093] Then, the shallow layer enhanced feature map and the target scale feature map are weighted and fused based on the attention mechanism through the multi-scale feature fusion network to obtain a multi-scale fusion feature map.
[0094] Specifically, the above step S146 includes steps S1461, S1462, S1463 and S1464.
[0095] Step S1461, upsampling the target scale feature map to have the same spatial size as the shallow layer enhanced feature map.
[0096] Step S1462, adjusting the channels of the shallow layer enhanced feature map and the target scale feature map after upsampling through a 1x1 convolution layer to obtain a to-be-fused shallow layer enhanced feature map and a to-be-fused target scale feature map; the to-be-fused shallow layer enhanced feature map and the to-be-fused target scale feature map have the same number of channels.
[0097] Step S1463, generating adaptive weights according to the to-be-fused shallow layer enhanced feature map and the to-be-fused target scale feature map through an attention module.
[0098] Step S1464, according to the adaptive weight, weighted sum of the to-be-fused shallow layer enhanced feature map and the to-be-fused target scale feature map is performed to obtain the multi-scale fusion feature map.
[0099] At this time, the target scale feature map includes 2, taking ResNet-18 as the backbone network as an example, the target scale feature map includes the feature map C3 output by stage3 and the C4 output by stage4. Correspondingly, the multi-scale feature fusion network includes 3 branches, denoted as the first branch, the second branch and the third branch.
[0100] The C1 is adjusted in channel number through the first branch to obtain the to-be-fused shallow layer enhanced feature map (denoted as F1). When adjusting the channel number, the channel number is adjusted to 256 using 1x1 convolution, and the spatial size of F1 is 56x56 and the channel number is 256.
[0101] The C3 is up-sampled through the second branch to obtain the to-be-fused feature map (denoted as F3). When up-sampling, 4 times up-sampling is used, and the spatial size of F3 is 56x56 and the channel number is 256.
[0102] The C4 is up-sampled and adjusted in channel number through the third branch to obtain the to-be-fused feature map (denoted as F4). When up-sampling, 8 times up-sampling is used, and then the channel number is adjusted to 256 using 1x1 convolution, and the spatial size of F4 is 56x56 and the channel number is 256.
[0103] At this time, the to-be-fused shallow layer enhanced feature map is F1, and the to-be-fused back-end layer feature map includes the above F3 and F4.
[0104] Then, the adaptive weights corresponding to each to-be-fused feature map are generated through the attention module, and are denoted as β1, β2 and β3 respectively. At this time, the multi-scale fusion feature map F is: F=(β1xF1)+(β2xF3)+(β3xF4).
[0105] The small target detection method provided in the embodiments of the present application filters the feature maps output by the last two network stages in the backend layer as target scale feature maps, and performs weighted fusion based on the attention mechanism with the shallow layer enhanced feature maps. Compared with selecting all the backend layer feature maps and the shallow layer enhanced feature maps for fusion, the fusion object in the embodiments of the present application is more targeted, can focus on the most representative semantic level (i.e., the target scale feature map) and the best quality detail source (i.e., the shallow layer enhanced feature map), avoids the interference of redundant and noisy features, enables the attention mechanism to more effectively perform fusion, thereby improving the key indicators (especially the small target detection accuracy), and at the same time, can significantly reduce the calculation amount and memory occupation in the multi-scale fusion process, improve the detection efficiency, and make it more suitable for real-time applications and resource-limited platforms. In addition, compared with selecting only the feature map output by the last network stage and the shallow layer enhanced feature map for fusion, the embodiments of the present application select the feature maps of the last two layers, which can ensure the diversity and hierarchy of high-level semantic information, and further improve the accuracy of the small target detection result.
[0106] Based on any of the above embodiments, Figure 5 is a fifth flowchart of the small target detection method provided by the present application, as shown in Figure 5 , step S120 includes step S121, step S122 and step S123.
[0107] It should be noted that step S121 and step S122 are executed in parallel.
[0108] Step S121 applies a horizontal convolution kernel to the initial shallow layer feature map to perform convolution operation and generate a horizontal direction feature map.
[0109] Step S122 applies a vertical convolution kernel to the initial shallow layer feature map to perform convolution operation and generate a vertical direction feature map.
[0110] The traditional convolution operation adopts a square convolution kernel (such as 3x3 or 5x5), which adopts a symmetric weight distribution in the horizontal and vertical directions. This uniform attention to global features makes it difficult for the network to extract features with directional characteristics, such as horizontally oriented road markings or vertically distributed building edges. And for small targets with limited pixel amount, this symmetric design weakens their key discriminative features. Therefore, in the present embodiment, an asymmetric convolution kernel (including a horizontal convolution kernel and a vertical convolution kernel) is designed in the asymmetric convolution module, which can break through the symmetry limitation of the traditional convolution and enhance the discriminative features such as the edge direction of the small target.
[0111] The asymmetric convolution module includes two parallel convolution branches: a horizontal convolution branch and a vertical convolution branch.
[0112] wherein, the horizontal convolution branch, applying a horizontal convolution kernel (denoted as k1) to the input initial shallow feature map (denoted as input) to generate the horizontal direction feature map (k1 input), the convolution operation slides in the horizontal direction of the feature map, which is specifically used to capture and enhance the horizontal structural features. The vertical convolution branch, applying a vertical convolution kernel (denoted as k2) to the same input initial shallow feature map, to obtain the vertical direction feature map (k2 input), the convolution operation slides in the vertical direction of the feature map, which is specifically used to capture and enhance the vertical structural features.
[0113] Further, the horizontal convolution kernel is a 1xm convolution kernel, and the vertical convolution kernel is a mx1 convolution kernel, where m is an integer greater than 1.
[0114] Correspondingly, the horizontal direction feature map and the vertical direction feature map are calculated as follows: ; ; wherein, a represents the input initial shallow feature map; i and j represent the position coordinates of the output feature map, i is the row index, and j is the column index; v represents the offset of the convolution kernel in the horizontal direction, which takes the value range of -m / 2~m / 2; u represents the offset of the convolution kernel in the vertical direction, which takes the value range of -m / 2~m / 2; represents the horizontal convolution kernel weight parameter; represents the vertical convolution kernel weight parameter.
[0115] It should be noted that by setting the convolution kernel of the same length in the horizontal direction and the vertical direction, balanced performance can be obtained on small targets of various shapes, avoiding missing detection due to favoring features in a specific direction, thereby further improving the overall accuracy of small target detection results.
[0116] In step S123, the horizontal direction feature map and the vertical direction feature map are added and fused to obtain the shallow enhanced feature map.
[0117] The two feature maps output by the horizontal convolution branch and the vertical convolution branch are added element by element, and the added result is added to a learnable bias term to obtain the shallow enhanced feature map. Specifically as follows: output= + +b; wherein, output is the shallow enhanced feature map, b is the bias, b is a scalar (single numerical value) which will be added to + Each element in the result matrix is a standard component in the neural network, which is used to increase the expression ability of the model.
[0118] By setting a 1xm convolution kernel and an mx1 convolution kernel, respectively used for generating a horizontal direction feature map and a vertical direction feature map, the dimensions of the horizontal direction feature map and the vertical direction feature map are the same, and then the shallow layer enhanced feature map can be directly added and fused. Through this fusion mode, the enhanced feature information in the horizontal and vertical directions can be superimposed without increasing the size and channel number of the feature map, forming a comprehensive feature representation more sensitive to linear and structural features. In addition, compared with using a traditional symmetric convolution kernel, the parameter quantity and the calculation complexity can be greatly reduced.
[0119] The small target detection method provided by the embodiment of the application uses an asymmetric convolution module, replaces the traditional symmetric convolution kernel with a horizontal direction convolution kernel and a vertical direction convolution kernel, greatly reduces the parameter quantity and the calculation complexity, and at the same time, through the special extraction and addition fusion of the horizontal and vertical features, a stronger representation ability for small target structures can be obtained, thereby improving the accuracy of the small target detection result.
[0120] Figure 6 is a structural schematic diagram of a small target detection device provided by the application, as shown in Figure 6 The device comprises a feature extraction module 610, a first processing module 620, a second processing module 630, a weighted fusion module 640 and a small target detection module 650; wherein: The feature extraction module 610 is used for performing feature extraction on a to-be-detected image through a front end layer of a backbone network to obtain an initial shallow layer feature map; The first processing module 620 is used for processing the initial shallow layer feature map through an asymmetric convolution module to obtain a shallow layer enhanced feature map; The second processing module 630 is used for inputting the shallow layer enhanced feature map to a back end layer of the backbone network, continuing to propagate forward, and obtaining back end layer feature maps output by multiple network stages; The weighted fusion module 640 is used for performing weighted fusion on the shallow layer enhanced feature map and the back end layer feature maps through a multi-scale feature fusion network to obtain a multi-scale fusion feature map; The small target detection module 650 is used for detecting the multi-scale fusion feature map through a detection head to obtain a small target detection result.
[0121] The small target detection device provided by the embodiment of the present application first extracts features of a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map. The initial shallow feature map contains rich detail information while retaining a high spatial resolution. Then, the initial shallow feature map is processed through an asymmetric convolution module to enhance directional details in the initial shallow feature map, so that a shallow enhanced feature map is obtained, thereby significantly strengthening the detail features that are crucial to small target detection. Next, the shallow enhanced feature map is input into a back-end layer of the backbone network for further propagation to obtain back-end layer feature maps output by multiple network stages. The shallow enhanced feature map and the back-end layer feature maps are weighted and fused through a multi-scale feature fusion network to obtain a multi-scale fused feature map. The fusion process effectively retains key semantics and detail information of different scales, so that the final feature map has both high-resolution details and high-level semantics. Finally, the multi-scale fused feature map is detected through a detection head to obtain a small target detection result. Through the cooperative mechanism of directional enhancement and multi-scale fusion, the embodiment of the present application organically combines the asymmetric convolution and multi-layer feature fusion technologies to improve the expression ability and discriminativeness of small target features, thereby significantly improving the accuracy of the small target detection result.
[0122] It should be noted that the small target detection device provided by the embodiment of the present application can implement all method steps achieved by the small target detection method embodiment and achieve the same technical effects. Therefore, the same parts and beneficial effects of the embodiment and the method embodiment will not be described in detail.
[0123] Figure 7 An example of a schematic diagram of a physical structure of an electronic device is shown in FIG. 7. Figure 7 As shown in FIG. 7, the electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740. The processor 710, the communications interface 720, and the memory 730 can communicate with each other through the communications bus 740. The processor 710 can invoke a logical instruction in the memory 730 to execute a small target detection method. The method includes extracting features of a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map; processing the initial shallow feature map through an asymmetric convolution module to obtain a shallow enhanced feature map; inputting the shallow enhanced feature map into a back-end layer of the backbone network for further propagation to obtain back-end layer feature maps output by multiple network stages; weighting and fusing the shallow enhanced feature map and the back-end layer feature maps through a multi-scale feature fusion network to obtain a multi-scale fused feature map; and detecting the multi-scale fused feature map through a detection head to obtain a small target detection result.
[0124] Further, the logic instructions in the memory 730 described above can be implemented by a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or partially contribute to the prior art, or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0125] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the small target detection method provided by the above-mentioned methods. The method comprises: performing feature extraction on a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map; processing the initial shallow feature map through an asymmetric convolution module to obtain a shallow enhanced feature map; inputting the shallow enhanced feature map into a back-end layer of the backbone network, and continuing to propagate forward to obtain a back-end layer feature map output by a plurality of network stages; performing weighted fusion on the shallow enhanced feature map and the back-end layer feature map through a multi-scale feature fusion network to obtain a multi-scale fusion feature map; and detecting the multi-scale fusion feature map through a detection head to obtain a small target detection result.
[0126] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the small target detection method provided by the above-mentioned methods. The method comprises: performing feature extraction on a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map; processing the initial shallow feature map through an asymmetric convolution module to obtain a shallow enhanced feature map; inputting the shallow enhanced feature map into a back-end layer of the backbone network, and continuing to propagate forward to obtain a back-end layer feature map output by a plurality of network stages; performing weighted fusion on the shallow enhanced feature map and the back-end layer feature map through a multi-scale feature fusion network to obtain a multi-scale fusion feature map; and detecting the multi-scale fusion feature map through a detection head to obtain a small target detection result.
[0127] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary hardware platform resources, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software product, and the computer software product can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0129] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A small object detection method characterized by, The method comprises the steps of: extracting features of a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map; processing the initial shallow feature map through an asymmetric convolution module to obtain a shallow enhanced feature map; inputting the shallow enhanced feature map into a back-end layer of the backbone network for further propagation to obtain back-end layer feature maps output by multiple network stages; performing weighted fusion on the shallow enhanced feature map and the back-end layer feature maps through a multi-scale feature fusion network to obtain a multi-scale fusion feature map; detecting the multi-scale fusion feature map through a detection head to obtain a small target detection result.
2. The small object detection method of claim 1, wherein The method of extracting features of a to-be-detected image through a front-end layer of a backbone network to obtain an initial shallow feature map comprises the steps of: extracting features of the to-be-detected image through an initial convolution layer of the backbone network to obtain an initial feature map; extracting features of the initial feature map through a first network stage of the backbone network to obtain the initial shallow feature map; wherein the first network stage is the first network stage after the first spatial down-sampling in the backbone network.
3. The small object detection method of claim 1, wherein The method of performing weighted fusion on the shallow enhanced feature map and the back-end layer feature maps through a multi-scale feature fusion network to obtain a multi-scale fusion feature map comprises the steps of: up-sampling the back-end layer feature maps to have the same spatial size as the shallow enhanced feature map; performing channel adjustment on the shallow enhanced feature map and the up-sampled back-end layer feature map through a 1x1 convolution layer to obtain a to-be-fused shallow enhanced feature map and a to-be-fused back-end layer feature map; the to-be-fused shallow enhanced feature map and the to-be-fused back-end layer feature map have the same number of channels; generating adaptive weights according to the to-be-fused shallow enhanced feature map and the to-be-fused back-end layer feature map through an attention module; performing weighted summation on the to-be-fused shallow enhanced feature map and the to-be-fused back-end layer feature map according to the adaptive weights to obtain the multi-scale fusion feature map.
4. The small object detection method according to any one of claims 1 to 3, characterized in that, The method of performing weighted fusion on the shallow enhanced feature map and the back-end layer feature maps through a multi-scale feature fusion network to obtain a multi-scale fusion feature map further comprises the steps of: selecting feature maps output by the last two network stages in the back-end layer from the back-end layer feature maps as target scale feature maps; performing attention mechanism-based weighted fusion on the shallow enhanced feature map and the target scale feature maps through a multi-scale feature fusion network to obtain a multi-scale fusion feature map.
5. The small object detection method according to any one of claims 1 to 3, characterized in that, The method of processing the initial shallow feature map through an asymmetric convolution module to obtain a shallow enhanced feature map comprises the steps of: applying a horizontal convolution kernel to the initial shallow feature map to perform convolution operation and generate a horizontal direction feature map; applying a vertical convolution kernel to the initial shallow feature map to perform convolution operation and generate a vertical direction feature map; performing addition fusion on the horizontal direction feature map and the vertical direction feature map to obtain the shallow enhanced feature map.
6. The small object detection method of claim 5, wherein The horizontal convolution kernel is a 1xm convolution kernel, and the vertical convolution kernel is a mx1 convolution kernel, wherein m is an integer greater than 1.
7. A small object detection apparatus, characterized by comprising: The method comprises the steps of: The feature extraction module is configured to perform feature extraction on the to-be-detected image through a front-end layer of the backbone network to obtain an initial shallow feature map; The first processing module is configured to perform processing on the initial shallow feature map through an asymmetric convolution module to obtain a shallow enhanced feature map; The second processing module is configured to input the shallow enhanced feature map into a back-end layer of the backbone network to continue forward propagation to obtain back-end layer feature maps output by multiple network stages; The weighting fusion module is configured to perform weighting fusion on the shallow enhanced feature map and the back-end layer feature maps through a multi-scale feature fusion network to obtain a multi-scale fusion feature map; The small target detection module is configured to perform detection on the multi-scale fusion feature map through a detection head to obtain a small target detection result.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to implement the small target detection method in any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the small target detection method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the small target detection method in any one of claims 1 to 6.
Citation Information
Patent Citations
Small target detection method based on multi-scale images and weighted fusion loss
CN111461110A
Improved YOLOv4 network model and small target detection method
CN114663654A
Remote sensing image target detection method based on multi-scale feature fusion and feature enhancement
CN114708511A
Feature enhancement and fusion small target detection method, device and equipment
CN116091942A
Infrared target detection method and device, computer equipment and storage medium
CN118229961A