Lightweight and efficient target detection method and device for small object aerial image
By enhancing the feature representation of small objects in aerial images through selective fusion and multi-scale pooling, the problem of difficult small object detection in complex backgrounds is solved, and efficient and accurate target detection is achieved.
Patent Information
- Application Number
- CN202510588380.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies have low target detection accuracy for small objects in aerial images, especially in complex backgrounds, where it is difficult to effectively extract and identify the features of small objects, resulting in low recognition rates.
By selectively fusing aerial image features at different levels, the fine-grained texture and edge information representation of small objects is enhanced, and contextual information is captured through multi-scale pooling and weighted fusion. The detection head is used to predict multi-scale feature maps to improve detection accuracy.
It achieves efficient and accurate detection of small objects in complex backgrounds to avoid omissions, and is suitable for scenarios such as drone monitoring, remote sensing, and disaster response.
Smart Images

Figure CN120612465A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a lightweight and efficient target detection method and device for small object aerial images. Background Art
[0002] There are many small, densely clustered objects (hereinafter referred to as "small objects") in high-altitude aerial images collected by aviation or drones (hereinafter referred to as "aerial images"). The existing technology has low accuracy in target detection of these small objects in aerial images, which seriously affects the recognition rate of small objects in scenarios such as drone monitoring, remote sensing, traffic monitoring and disaster response.
[0003] Small objects in aerial imagery typically occupy only a few pixels, yet they are embedded in complex backgrounds. Existing techniques often extract features through repeated downsampling and pooling, but these operations can cause the features of small objects to disappear or blend into the background, making them difficult to extract. Furthermore, when small objects are close to or even overlap with other objects in a complex background, their feature representations can overlap in the feature map, further increasing the difficulty of feature extraction. Consequently, detecting small objects from complex backgrounds is challenging.
[0004] In view of this, overcoming the defects of the prior art is an urgent problem to be solved in this technical field. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a lightweight and efficient target detection method and device for small object aerial images. Its purpose is to selectively fuse features in aerial images at different levels to enhance the representation of fine-grained texture and edge information corresponding to small objects, facilitate the effective extraction of features of small objects, and solve the problem that it is difficult to detect small objects from the complex background of aerial images.
[0006] The present invention adopts the following technical solutions: In a first aspect, the present invention provides a lightweight and efficient target detection method for small object aerial images, comprising: Performing a convolution operation on the aerial image to extract basic features; selectively fusing the basic features to obtain a first joint feature; performing a convolution operation on the first joint feature, selectively fusing the convolved features to obtain a next joint feature, and sequentially performing convolution and selective fusion on each joint feature until a final joint feature is obtained; Performing multi-scale pooling on the last joint feature, and weighted fusion of the pooled features based on learnable weights to obtain contextual features; Fusing all joint features and the context features to obtain a multi-scale feature map; The multi-scale feature map is input into the detection head, which predicts the bounding box and class probability at each spatial position and scale to obtain the object detection result.
[0007] Furthermore, the selective fusion of the basic features to obtain the first joint feature includes: Convolution and preprocessing are performed on the basic features to obtain features to be enhanced; Performing identity mapping on the features to be enhanced to obtain original features; generating an additional feature channel according to the feature to be enhanced, and obtaining a detail feature based on the additional feature channel and the feature to be enhanced; Performing feature enhancement on the feature to be enhanced based on the spatial attention weight to obtain an attention feature; A first joint feature is generated using the original feature, the detail feature, and the attention feature.
[0008] Furthermore, generating an additional feature channel according to the feature to be enhanced, and obtaining a detail feature based on the additional feature channel and the feature to be enhanced includes: Using a unit-size convolution kernel to perform a convolution operation on the feature to be enhanced to obtain an intrinsic feature map; Using multiple separable convolution kernels to convolve each intrinsic feature channel in the intrinsic feature map, respectively, to obtain a new feature channel corresponding to each intrinsic feature channel; Each of the intrinsic feature channels is fused with the corresponding new feature channel to obtain detail features.
[0009] Furthermore, the feature enhancement to be enhanced is performed based on the spatial attention weight to obtain the attention feature, which includes: Performing a standard convolution operation on the features to be enhanced to obtain high-level features; Generate a single-channel attention map of the high-level features; Processing the single-channel attention map using an activation function to obtain a spatial attention mask; The product of the spatial attention weight of each channel in the spatial attention mask and the pixel value of the corresponding feature channel in the high-level feature is calculated to obtain the attention feature.
[0010] Furthermore, generating a first joint feature using the original feature, the detail feature, and the attention feature includes: In the channel dimension, the original features, the detail features, and the attention features are concatenated to obtain multiple channel features; The multiple channel features are fused into a first joint feature.
[0011] Furthermore, the fusing of all joint features and the context features to obtain a multi-scale feature map includes: Assigning a preset number of joint features to at least one alignment group; performing multi-scale alignment on the joint features in each alignment group to obtain at least one first alignment feature map; Obtaining remaining joint features from all the joint features, and performing multi-scale alignment on the remaining joint features and the context features to obtain a second aligned feature map; The at least one first aligned feature map and the second aligned feature map are fused to obtain a multi-scale feature map.
[0012] Furthermore, performing multi-scale alignment on the joint features in each alignment group to obtain at least one first alignment feature map includes: Adjusting each joint feature in the alignment group to a predefined spatial size to obtain a plurality of features to be merged; Using a preset convolution kernel to perform a convolution operation on each of the features to be merged to adjust the feature channels of the features to be merged to obtain a plurality of adjusted features; In the channel dimension, performing feature splicing on the multiple adjusted features to obtain spliced features; A convolution operation is performed on the concatenated features using a unit-size convolution kernel to obtain a first aligned feature map.
[0013] Furthermore, the last joint feature is subjected to multi-scale pooling, and the pooled features are weightedly fused based on learnable weights to obtain contextual features including: Performing a convolution operation on the last joint feature using a lightweight convolution kernel to obtain a feature to be processed; Using pooling kernels of different sizes to pool the features to be processed respectively to obtain multiple pooled features; wherein a pooling kernel of one size corresponds to one pooled feature; Taking the number of the plurality of pooled features as a total weight; determining a plurality of different learnable weights according to the total weight; wherein the sum of the plurality of different learnable weights is one; Using different learnable weights, weighting each of the pooled features to obtain a plurality of weighted features; wherein one learnable weight corresponds to one weighted feature; Feature concatenation is performed on the multiple weighted features to obtain multi-scale features; and a convolution operation is performed on the multi-scale features to obtain context features.
[0014] In a second aspect, the present invention further provides a lightweight and efficient target detection device for small object aerial images, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the lightweight and efficient target detection method for small object aerial images described in the first aspect.
[0015] In a third aspect, the present invention further provides a non-volatile computer storage medium storing computer executable instructions, which are executed by one or more processors to complete the lightweight and efficient target detection method for small object aerial images described in the first aspect.
[0016] In a fourth aspect, a computer program product comprising instructions is provided, which, when executed on a computer or a processor, enables the computer or processor to execute the lightweight and efficient target detection method for small object aerial images as described in the first aspect.
[0017] In the fifth aspect, the present invention also provides a lightweight and efficient target detection system for small object aerial images, including the lightweight and efficient target detection device for small object aerial images as in the second aspect, and using the lightweight and efficient target detection method for small object aerial images as described in the first aspect to complete the interaction of the lightweight and efficient target detection device for small object aerial images of the second aspect.
[0018] Different from the prior art, the present invention has at least the following beneficial effects: The present invention selectively fuses features in aerial images at different levels to specifically enhance the representation of fine-grained texture and edge information corresponding to small objects. Multi-scale pooling and weighted fusion are performed on the last joint feature to capture contextual information of different receptive fields in the deepest joint features, thereby enabling the neural network model to focus on scales related to small object detection. The joint features and contextual features processed at different levels are aligned and fused, and then the detection head is used to predict the fused multi-scale feature map; since the features of small objects have been enhanced in the multi-scale feature map, the detection head can more accurately locate small objects, avoid missing small objects in the aerial image, and solve the problem of difficulty in detecting small objects from the complex background of the aerial image. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0020] Figure 1 This is a schematic diagram of the overall process of a lightweight and efficient target detection method for small object aerial images provided by an embodiment of the present invention; Figure 2 This is an overall schematic diagram of a neural network structure according to an embodiment of the present invention; Figure 3 is a schematic diagram of a specific example of a selective convolution module provided by an embodiment of the present invention; Figure 4 is a flow chart of step 10 provided in an embodiment of the present invention; Figure 5 is a flow chart of step 103 provided by an embodiment of the present invention; Figure 6 is a schematic diagram of a specific example of a filament module provided by an embodiment of the present invention; Figure 7 is a flow chart of step 104 provided by an embodiment of the present invention; Figure 8 is a flow chart of step 105 provided by an embodiment of the present invention; Figure 9 is a flow chart of step 20 provided in an embodiment of the present invention; Figure 10 1 is a schematic diagram of a specific example of a fast spatial pyramid pooling module provided by an embodiment of the present invention; Figure 11 is a flow chart of step 30 provided in an embodiment of the present invention; Figure 12 301 is a flowchart of a step provided by an embodiment of the present invention; Figure 13 is a schematic diagram of a specific example of a multi-scale alignment module provided by an embodiment of the present invention; Figure 14 is a schematic diagram of a specific example of a RepVGG module provided by an embodiment of the present invention; Figure 15 This is a schematic diagram of the architecture of a lightweight and efficient target detection device for small object aerial images provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0022] Unless the context requires otherwise, throughout the specification and claims, the term "including" is to be interpreted as meaning open inclusion, that is, "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "example", "specific example" or "some examples" and the like are intended to indicate that the specific features, structures, materials or characteristics associated with the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representation of the above terms does not necessarily refer to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner, that is, although they may be carried in the embodiments or examples of the above terms due to reasons such as the order and position of appearance, it is not limited to that they can be carried in combination by one embodiment or example.
[0023] In the description of the present invention, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present disclosure.
[0024] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "multiple" means two or more. In addition, for example, the description may also use the method of adding "A" and "B" at the end to describe the same type of nouns as two independent individuals. In this case, the corresponding features defined as "A" and "B" are only used to distinguish the description purposes of the same type of individuals, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated.
[0025] When describing some embodiments, the expressions “coupled”, “coupled” and “connected” and their derivatives may be used. For example, when describing some embodiments, the term “connected” may be used to indicate that two or more components are in direct physical or electrical contact with each other. For another example, when describing some embodiments, the term “coupled” may be used to indicate that two or more components are in direct physical or electrical contact. However, the term “connected” or “coupled” may also mean that two or more components are not in direct contact with each other, but still cooperate or interact with each other, such as “optical coupling”, “wireless connection”, etc. The embodiments disclosed herein are not necessarily limited to the contents of the present invention.
[0026] In the description of the present invention, the expression "A and / or B" (where A and B are used to formally represent specific characteristic contents) is involved, and the corresponding expressions include the following three combinations: only A, only B, and a combination of A and B.
[0027] As used herein, "about," "substantially," or "approximately" includes the stated value and an average value that is within an acceptable range of deviation from the particular value as determined by one of ordinary skill in the art taking into account the measurements in question and errors associated with measurement of the particular quantity (i.e., limitations of the measurement system).
[0028] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0029] Embodiment 1: Detecting small objects in aerial imagery is a challenge for standard convolutional neural network-based object detectors. Frequently seen objects in aerial imagery, such as vehicles, pedestrians, and infrastructure, occupy only a few pixels but are embedded in complex, cluttered scenes. Traditional one-stage and two-stage detectors, including You Only Look Once (YOLO) networks, Single Shot Multibox Detector (SSD) networks, and Faster Regions with Convolutional Neural Network (Faster R-CNN), perform well on common benchmarks but often suffer from reduced fidelity for aerial objects in aerial imagery.
[0030] Because repeated downsampling and pooling layers in the backbone network of the existing deep convolutional neural network will cause the features of small objects to disappear or merge with the background, the recall rate and positioning accuracy of small objects are low, and the practicality is poor.
[0031] Dense object clusters further complicate detection: when many objects are close together in aerial images, their feature representations overlap in corresponding feature maps, causing standard object detectors to confuse or merge adjacent objects. Noise introduced by complex backgrounds (e.g., cities, forests, or textured terrain) can drown out the signal of small objects because ordinary neural networks lack specialized mechanisms to suppress irrelevant background information.
[0032] Furthermore, existing object detectors often trade speed for accuracy when optimizing. Large models utilizing multi-scale fusion mechanisms achieve high accuracy, but at the expense of larger memory footprints and longer inference latency. This makes them less practical for edge computing platforms or drones with limited hardware resources.
[0033] In order to solve the above problems, Figure 1 As shown, an embodiment of the present invention provides a lightweight and efficient target detection method for small object aerial images, including: Step 10: Perform a convolution operation on the aerial image to extract basic features; selectively fuse the basic features to obtain a first joint feature; perform a convolution operation on the first joint feature, selectively fuse the convolved features to obtain the next joint feature, and convolve and selectively fuse each joint feature in turn until the last joint feature is obtained.
[0034] like Figure 2 The figure shows the overall architecture of a neural network model according to an embodiment of the present invention, which includes a backbone network, a neck network, and a head network. In the backbone network, convolutional layers perform convolution operations on aerial images to extract low-level features (such as edges and textures) and gradually reduce the resolution to ultimately obtain basic features. The selective convolution module selectively fuses corresponding feature maps.
[0035] like Figure 2As shown, in one embodiment, an aerial image serves as input to the backbone network and is first convolved through two convolutional layers to initially extract features. The second convolutional layer then outputs basic features. The first selective convolutional module then selectively fuses the basic features. In an optional embodiment, after each selective fusion, a convolutional layer may be used to convolve the selectively fused feature map to facilitate subsequent better feature extraction, ultimately yielding the first joint feature. The first joint feature is then fed into the second selective convolutional module, where it is selectively fused again to yield a selectively fused feature. After passing through another convolutional layer, the second joint feature is then fed into the third selective convolutional module, where it is selectively fused again to yield a selectively fused feature. After passing through another convolutional layer, the third joint feature is then fed into the fourth selective convolutional module, where it is selectively fused again to yield a final joint feature.
[0036] In an optional embodiment, the size of the aerial image may be 640×640 pixels. The convolution layer may be a standard convolution layer; the convolution layer may include a stride or pooling operation.
[0037] The embodiment of the present invention inserts multiple selective convolution modules at selected depth positions in the backbone network to specifically enhance the feature extraction capability of the neural network model of the embodiment of the present invention to extract small objects from aerial or drone images. Figure 3 As shown in the figure, the selective convolution module consists of three parallel branches, namely the identity mapping branch, the filament feature generator branch and the convolution branch with spatial attention, and the outputs of each branch are spliced and fused into a joint feature map in the channel dimension. The process of selective fusion will be explained below.
[0038] Step 20: Perform multi-scale pooling on the last joint feature, and perform weighted fusion on the pooled features based on learnable weights to obtain contextual features.
[0039] Since the backbone network performs feature extraction at the deepest level to obtain the last joint feature, the embodiment of the present invention performs spatial pooling on the last joint feature at multiple scales to capture contextual information of different receptive fields in the last joint feature, thereby enabling the neural network model of the embodiment of the present invention to focus on scales related to small object detection.
[0040] Step 30: Fuse all joint features and the context features to obtain a multi-scale feature map.
[0041] In the neck network, a multi-scale alignment module is used to resize and process feature maps from different scales (i.e., all joint features and context features), and perform feature splicing to generate aligned and fused multi-scale feature maps. Figure 2 In the figure, the symbol containing "C" in the circle represents feature splicing of the input features. The specific method of feature splicing is selected by those skilled in the art according to the specific usage scenario and is not limited here.
[0042] In one embodiment, if Figure 2 As shown, the features after feature splicing can also be input into the re-parameterization Visual Geometry Group (RepVGG) module, and then the features processed by the RepVGG module are upsampled, and the upsampled features and the first joint features are feature spliced to obtain the first-level features; and the features processed by the RepVGG module are directly used as the second-level features; wherein, the use of the RepVGG module will be explained below. A convolution operation is performed on the first-level features using a convolution kernel of a first preset size, so that the convolved first-level features can be subsequently input into the detection head; a convolution operation is performed on the second-level features using a convolution kernel of a second preset size, so that the convolutioned second-level features can be subsequently input into the detection head; a convolution operation is performed on the first-level features using a convolution kernel of a third preset size, so that the convolutioned second-level features can be subsequently input into the detection head, and the detection head receives features of each level, thereby better utilizing the features characterizing small objects captured from a complex background; wherein, the process of processing the second-level features using the convolution kernel of the second preset size is different from the processing process using the convolution kernel of the third preset size; in an optional embodiment, the third preset size can be 3×3, and in the processing process of the convolution kernel of the third preset size, the step size of the convolution operation can be 2; the first preset size, the second preset size and the third preset size are selected by technical personnel in this field according to the specific usage scenario.
[0043] Step 40: Input the multi-scale feature map into the detection head, which predicts the bounding box and category probability at each spatial position and scale to obtain the object detection result.
[0044] The detection head is selected by those skilled in the art according to the specific usage scenario and is not limited here.
[0045] In the head network of an embodiment of the present invention, multi-scale features are input into a detection head, which predicts the bounding box of small objects in aerial images and the category to which the small objects belong, thereby obtaining a category score for the possible category. Finally, the predicted bounding box and its corresponding category label are output as the object detection result.
[0046] In one embodiment, the detection head includes multiple convolutional layers, each of which is configured to output a set of bounding box offsets, objectness scores, and class probabilities for each spatial unit, thereby outputting bounding box predictions and class probabilities at each spatial location and scale. For example, three different feature map resolutions can be used to detect large, medium, and small objects. Non-maximum suppression is applied to the bounding box predictions and class probabilities corresponding to each convolutional layer to filter out overlapping predictions, ultimately resulting in object detection results.
[0047] In an optional embodiment, the first-level features, the second-level features, and the third-level features are all input into the detection head, and the convolutional layer of the detection head generates bounding box coordinates, objectness scores, and category probabilities at multiple spatial scales; wherein three parallel prediction heads can be used to detect objects of various sizes in the multi-scale feature map at different resolutions.
[0048] It should be noted that the lightweight and efficient target detection method for small object aerial images in the embodiment of the present invention is suitable for drone monitoring, remote sensing, traffic monitoring, disaster response and other scenarios that require real-time identification of small, densely clustered objects in high-altitude images.
[0049] The present invention selectively fuses features in aerial images at different levels to specifically enhance the representation of fine-grained texture and edge information corresponding to small objects. Multi-scale pooling and weighted fusion are performed on the last joint feature to capture contextual information of different receptive fields in the deepest joint features, thereby enabling the neural network model to focus on scales related to small object detection. The joint features and contextual features processed at different levels are aligned and fused, and then the detection head is used to predict the fused multi-scale feature map; since the features of small objects have been enhanced in the multi-scale feature map, the detection head can more accurately locate small objects, avoid missing small objects in the aerial image, and solve the problem of difficulty in detecting small objects from the complex background of the aerial image.
[0050] In order to enhance the features of small objects in aerial images so as to detect small objects more accurately, such as Figure 4 As shown, in step 10, the selective fusion of the basic features to obtain the first joint feature includes: Step 101: Convolve and preprocess the basic features to obtain features to be enhanced.
[0051] like Figure 3 Shown is a specific example of a selective convolution module, the basic features As input, it first passes through the convolution module for convolution, and then passes through the segmentation module for image preprocessing to obtain the features to be enhanced.
[0052] Step 102: Perform identity mapping on the feature to be enhanced to obtain the original feature.
[0053] The selective convolution module processes the features to be enhanced through an identity mapping branch, a filament feature generator branch, and a convolution branch with spatial attention. The identity mapping branch provides a skip connection, directly passing the input features to be enhanced, thereby preserving the original spatial details and contextual information and obtaining the original features.
[0054] Step 103: Generate additional feature channels according to the features to be enhanced, and obtain detail features based on the additional feature channels and the features to be enhanced.
[0055] In the filament feature generator branch, the filament module generates a small set of intrinsic feature maps through lightweight convolution (e.g., convolution with a convolution kernel of size 1×1), and expands additional new feature channels based on the intrinsic feature channels in the intrinsic feature maps through multiple low-computational transformations (e.g., depth-wise separable convolution or linear projection) to explicitly preserve and enhance the detail information of small objects.
[0056] Step 104: Perform feature enhancement on the feature to be enhanced based on the spatial attention weight to obtain an attention feature.
[0057] Since the noise introduced by the complex background is likely to drown out the signal of small objects, the embodiments of the present invention use the attention mechanism to suppress irrelevant background information through the convolution branch with spatial attention (i.e., the attention module), so as to effectively suppress background interference and highlight the relevant areas where small objects are located.
[0058] Step 105: Generate a first joint feature using the original feature, the detail feature and the attention feature.
[0059] The outputs of the above three branches are connected and fused through convolution to get the output In one embodiment, the size of the convolution kernel used for fusion can be 1×1. This design of the embodiment of the present invention preserves the original spatial details through the identity mapping branch, enriches the feature channels through the filament feature generator branch, and highlights the salient areas through the attention of the convolution branch with spatial attention, thereby improving the ability of the neural network model to detect small and overlapping objects without significantly increasing the computational complexity.
[0060] In order to enhance the features of small objects while keeping the network structure light enough to facilitate real-time or embedded deployment of the neural network model of the embodiment of the present invention, the embodiment of the present invention also provides a filament module, such as Figure 5 As shown, step 103 includes: Step 1031: Use a unit-size convolution kernel to perform a convolution operation on the feature to be enhanced to obtain an intrinsic feature map.
[0061] Among them, the unit size convolution kernel is: a convolution kernel with a size of 1×1.
[0062] like Figure 6 Shown is a specific example of the filament module, which first extracts the intrinsic feature map through 1×1 convolution to capture important object details.
[0063] Step 1032: Use multiple separable convolution kernels to convolve each intrinsic feature channel in the intrinsic feature map to obtain a new feature channel corresponding to each intrinsic feature channel.
[0064] Among them, the intrinsic feature map is composed of multiple intrinsic feature channels.
[0065] For example, the intrinsic feature map is divided into three parts, and the first part of the intrinsic feature map is processed by separable convolution kernel 1 to generate the corresponding new feature channel; the second part is processed by separable convolution kernel 2 to generate the corresponding new feature channel; the third part is processed by separable convolution kernel 3 to generate the corresponding new feature channel; thereby generating an additional new feature channel for each intrinsic channel in the intrinsic feature map, that is, obtaining an additional lightweight feature map, thereby improving the richness of feature representation with minimal overhead.
[0066] In an optional embodiment, each intrinsic channel in the intrinsic feature map may be expanded into multiple new feature channels through linear projection.
[0067] Step 1033: Fuse each of the intrinsic feature channels with the corresponding new feature channel to obtain detail features.
[0068] The original intrinsic feature map and the additional lightweight feature map composed of the new feature channels are concatenated along the channel dimension. The filament module enriches the feature representation, especially the detail feature representation of small objects, while maintaining low computational complexity. Figure 6 The lightweight design shown ensures the high efficiency and operating speed of the neural network model of the embodiment of the present invention; the corresponding neural network structure not only achieves high precision in small-scale aerial target detection tasks, but also has the advantages of compact structure and suitability for edge deployment.
[0069] To illustrate the convolution branch with spatial attention, Figure 7 As shown, step 104 includes: Step 1041: Perform a standard convolution operation on the features to be enhanced to obtain high-level features.
[0070] The attention module first performs a standard convolution operation on the input to obtain high-level features. The specific method of the standard convolution operation is selected by those skilled in the art based on the specific usage scenario and is not limited here.
[0071] Step 1042: Generate a single-channel attention map of the high-level features.
[0072] Among them, high-level features are used to generate a two-dimensional single-channel attention map. The specific generation method is selected by those skilled in the art according to the specific usage scenario and is not limited here.
[0073] Step 1043: Use an activation function to process the single-channel attention map to obtain a spatial attention mask.
[0074] The activation function is selected by those skilled in the art according to a specific usage scenario. In one embodiment, the activation function may be a sigmoid function.
[0075] The spatial attention mask has the same size as the high-level feature and contains the spatial attention weights for each feature channel in the high-level feature.
[0076] Step 1044: Calculate the product of the spatial attention weight of each channel in the spatial attention mask and the pixel value of the corresponding feature channel in the high-level feature to obtain the attention feature.
[0077] The attention mechanism is used to highlight salient areas, thereby enhancing the characteristics of the target area where small objects are located and suppressing background interference.
[0078] To illustrate the fusion process, Figure 8 As shown, step 105 includes: Step 1051: In the channel dimension, feature splicing is performed on the original features, the detail features, and the attention features to obtain multiple channel features.
[0079] like Figure 3 As shown in the figure, the outputs of the three branches are first directly concatenated.
[0080] Step 1052: Fuse the multiple channel features into a first joint feature.
[0081] In one embodiment, the concatenated features are fused through a convolution operation.
[0082] After generating the last joint feature through the last selective convolution module, in order to adaptively aggregate context information, such as Figure 9 As shown, the step 20 includes: Step 201: Use a lightweight convolution kernel to perform a convolution operation on the last joint feature to obtain a feature to be processed.
[0083] like Figure 10 FIG. 4 shows a specific example of a fast spatial pyramid pooling module. The fast spatial pyramid pooling module first performs preliminary processing through convolution to obtain features to be processed.
[0084] Step 202: Use pooling kernels of different sizes to pool the features to be processed respectively to obtain multiple pooled features; wherein a pooling kernel of one size corresponds to one pooled feature.
[0085] Multi-scale pooling is then performed on the feature map of the deepest joint feature.
[0086] The size of the pooling kernel is selected by those skilled in the art according to the specific usage scenario; in one embodiment, Figure 10 As shown in the figure, the number of pooling kernels can be 3, and the corresponding pooling kernel sizes can be 3×3, 5×5, and 7×7, respectively. Each pooling kernel performs pooling in a maximum pooling manner, and the corresponding pooling step and padding are adaptively adjusted according to the selected pooling kernel size.
[0087] Step 203: Taking the number of the plurality of pooled features as a total weight; determining a plurality of different learnable weights according to the total weight; wherein the sum of the plurality of different learnable weights is one.
[0088] Weighted fusion is performed via learnable weights to generate adaptive contextual features.
[0089] In the embodiment of the present invention, each pooled feature map is weighted by learnable weights (α, β, γ, etc.), and all learnable weights satisfy α+β+γ+…=1, thereby adaptively adjusting the contributions of different scales through the learnable weights.
[0090] Step 204: using different learnable weights to weight each of the pooled features to obtain a plurality of weighted features; wherein one learnable weight corresponds to one weighted feature.
[0091] Step 205: performing feature concatenation on the multiple weighted features to obtain multi-scale features; performing a convolution operation on the multi-scale features to obtain context features.
[0092] Each pooling output is scaled by a learnable weight and then combined into the final contextual feature.
[0093] In the neck network, embodiments of the present invention use spatial pooling with various convolution kernel sizes (3×3, 5×5, and 7×7) on the feature map of the deepest joint features to capture context in different receptive fields. Each pooled output is assigned a learnable weight, and the weighted features are combined into a single map. The learnable weights (e.g., α, β, γ) are normalized (e.g., α+β+γ=1) so that the neural network can adaptively emphasize the most informative scales, thereby balancing details with coarse background. This weighted downsampling strategy enhances flexibility, enabling the model to focus on scales relevant to small object detection and reducing redundant computation.
[0094] like Figure 2 As shown in , the feature maps output from multiple selective convolution modules and the context features contain a lot of effective information highlighting the features of small objects, which is helpful for small object target detection. However, the sizes of these feature maps are often not aligned. In order to utilize these feature maps, the scale mismatch problem is corrected, as shown in Figure 11 As shown, the step 30 includes: Step 301: assigning a preset number of joint features to at least one alignment group; performing multi-scale alignment on the joint features in each alignment group to obtain at least one first alignment feature map.
[0095] The preset number is selected by those skilled in the art according to the specific usage scenario.
[0096] like Figure 2 As shown in the figure, a multi-scale alignment module is first used to adjust the feature maps from the first three different depths of the backbone network to a uniform spatial resolution. Each feature map is refined through a convolutional layer and spliced in the channel dimension. Then, a convolutional fusion operation is performed to achieve multi-scale alignment, and finally the first aligned feature map is obtained.
[0097] Step 302: Obtain the remaining joint features from all the joint features, perform multi-scale alignment on the remaining joint features and the context features, and obtain a second aligned feature map.
[0098] like Figure 2 As shown, in one embodiment, the feature map output by the third selective convolution module can be directly used as a joint feature and multi-scale aligned with the last joint feature and the context feature.
[0099] Step 303: Fusing the at least one first aligned feature map and the second aligned feature map to obtain a multi-scale feature map.
[0100] For example, Figure 2 In
[15] , the aligned feature maps of the joint features from three backbone scales (e.g., early, mid, and late layers) are fused through two multi-scale alignment modules.
[0101] Specifically, to illustrate the processing flow of each multi-scale alignment module, as shown in Figure 12 As shown, in step 301, performing multi-scale alignment on the joint features in each alignment group to obtain at least one first alignment feature map includes: Step 3011: Adjust each joint feature in the alignment group to a predefined spatial size to obtain multiple features to be merged.
[0102] like Figure 13 The figure shows how three feature maps from different stages of the backbone network (i.e., the first three joint features) are resized to a common resolution, refined by convolution, concatenated, and then fused by the final convolution. For example, the joint features are first resized to a certain size using either evaluation pooling or interpolation.
[0103] The predefined spatial size is selected by those skilled in the art according to specific usage scenarios. The length and width of each feature to be merged are consistent.
[0104] Step 3012: Use a preset convolution kernel to perform a convolution operation on each of the features to be merged to adjust the feature channels of the features to be merged to obtain multiple adjusted features.
[0105] Among them, the preset convolution kernel is selected by those skilled in the art according to the specific usage scenario and is not limited here.
[0106] Each feature map is then processed by one or more convolutions (e.g., convolution operations using kernels of size 1×1 or 3×3) to refine and normalize the features.
[0107] Step 3013: In the channel dimension, perform feature splicing on the multiple adjusted features to obtain spliced features.
[0108] The processed feature maps are concatenated channel by channel and processed by a final fused convolution (usually 1×1).
[0109] Step 3014: Use a unit-size convolution kernel to perform a convolution operation on the spliced features to obtain a first aligned feature map.
[0110] By matching spatial sizes and channel contexts, the multi-scale alignment module prevents mismatches when fusing high-level and low-level joint features; this alignment ensures the coherent combination of multi-scale information and improves small object localization across scales.
[0111] The backbone network of the embodiment of the present invention improves the operating efficiency through the selective convolution module. In one embodiment, in order to further improve the operating efficiency, the RepVGG module can also be used in the backbone network.
[0112] like Figure 14 The figure shows a specific example of a RepVGG module. During training, each RepVGG module contains multiple (e.g., two or three) parallel 3×3 convolutional branches, each equipped with an independent BatchNorm layer and optionally with residual skip connections. These parallel branches increase network capacity while facilitating gradient flow during training. After training, the weights of these parallel branches and their BatchNorm layers are mathematically fused into an equivalent single 3×3 convolution kernel. Therefore, during inference, each RepVGG module effectively only needs to perform a single fast 3×3 convolution operation, achieving low latency. Overall, the backbone network alternates between RepVGG modules (for general feature extraction) and selective convolutional modules (for feature enhancement), achieving a balance between feature richness and computational speed. This provides the expressive power of a multi-branch network during learning while performing as an efficient single-path convolution at runtime, enabling low latency for edge deployment.
[0113] In an optional embodiment, the neural network of the embodiment of the present invention includes two variants: a standard variant, which reduces the channel width of convolutional layers and blocks (for example, using narrower width multipliers) and can use fewer selective convolution modules. The reduced channel width and fewer layers reduce resource consumption and are suitable for real-time deployment and low-resource environments. The other is an extended variant, which increases the number of channels (for example, wider layers or additional blocks) to improve the capability of high-precision tasks. When using more powerful hardware, the number of channels can be doubled or a selective convolution module can be added in key layers. The increased channel width and / or additional layers are used to improve detection accuracy if hardware resources allow, and are suitable for tasks with high precision requirements. The choice can be made based on deployment requirements.
[0114] The selective convolution module utilizes the selective fusion mechanism, and the fast spatial pyramid pooling module uses the last joint feature to generate contextual features, so that the lightweight and efficient target detection method for small object aerial images using the embodiment of the present invention can more accurately detect small-sized, densely distributed target objects in aerial images than the existing technology, and can achieve high-precision detection while maintaining a compact structure. At the same time, the filament module efficiently expands additional feature channels, making the neural network model light enough and maintaining the computing speed, so as to facilitate high-throughput reasoning using the neural network model of the embodiment of the present invention.
[0115] Example 2: like Figure 15 FIG. 1 is a schematic diagram of an architecture of a lightweight and efficient target detection device for small object aerial images according to an embodiment of the present invention. The lightweight and efficient target detection device for small object aerial images according to this embodiment includes one or more processors 21 and a memory 22. Figure 15 A processor 21 is taken as an example.
[0116] The processor 21 and the memory 22 may be connected via a bus or other means. Figure 15 The bus connection is taken as an example.
[0117] Memory 22, as a nonvolatile computer-readable storage medium, can be used to store nonvolatile software programs and nonvolatile computer-executable programs, such as the lightweight and efficient small object aerial imagery target detection method of this embodiment. Processor 21 executes the lightweight and efficient small object aerial imagery target detection method by running the nonvolatile software program and instructions stored in memory 22.
[0118] The memory 22 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 22 may optionally include a memory remotely located relative to the processor 21, and such remote memory may be connected to the processor 21 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0119] The program instructions / modules are stored in the memory 22. When executed by the one or more processors 21, the lightweight and efficient target detection method for small object aerial images in the above-mentioned embodiment is executed, for example, each step of the lightweight and efficient target detection method for small object aerial images in the embodiment of the present invention described above is executed.
[0120] An embodiment of the present invention further provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions are executed by one or more processors, for example Figure 15 A processor 21 can enable the above one or more processors to execute the lightweight and efficient target detection method for small objects in aerial images in a specific embodiment of the present invention, for example, to execute the various steps of the lightweight and efficient target detection method for small objects in aerial images in the embodiment of the present invention described above; it can also realize Figure 15 The various modules and units described above; or executing the lightweight and efficient target detection method for small object aerial images in a specific embodiment of the present invention, for example, executing the various steps of the lightweight and efficient target detection method for small object aerial images in the embodiment of the present invention described above; it can also be realized Figure 15The various modules and units described.
[0121] It is worth noting that the information interaction, execution process, etc. between the modules and units within the above-mentioned devices and systems are based on the same concept as the processing method embodiment of the present invention. The specific content can be found in the description of the method embodiment of the present invention and will not be repeated here.
[0122] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a disk or an optical disk, etc.
[0123] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A lightweight and efficient target detection method for small objects in aerial images, characterized by: include: Perform convolution operations on aerial images to extract basic features; The basic features are selectively fused to obtain a first joint feature; the first joint feature is convolved, and the convolved features are selectively fused to obtain the next joint feature, and each joint feature is convolved and selectively fused in turn until the last joint feature is obtained; Performing multi-scale pooling on the last joint feature, and weighted fusion of the pooled features based on learnable weights to obtain contextual features; Fusing all joint features and the context features to obtain a multi-scale feature map; The multi-scale feature map is input into the detection head, which predicts the bounding box and class probability at each spatial position and scale to obtain the object detection result.
2. The lightweight and efficient small object detection method for aerial images according to claim 1 is characterized in that: The method comprises: Convolution and preprocessing are performed on the basic features to obtain features to be enhanced; Performing identity mapping on the features to be enhanced to obtain original features; generating an additional feature channel according to the feature to be enhanced, and obtaining a detail feature based on the additional feature channel and the feature to be enhanced; Performing feature enhancement on the feature to be enhanced based on the spatial attention weight to obtain an attention feature; A first joint feature is generated using the original feature, the detail feature, and the attention feature.
3. The lightweight and efficient small object detection method for aerial images according to claim 2 is characterized in that: The method comprises: Using a unit-size convolution kernel to perform a convolution operation on the feature to be enhanced to obtain an intrinsic feature map; Using multiple separable convolution kernels to convolve each intrinsic feature channel in the intrinsic feature map, respectively, to obtain a new feature channel corresponding to each intrinsic feature channel; Each of the intrinsic feature channels is fused with the corresponding new feature channel to obtain detail features.
4. The lightweight and efficient small object detection method for aerial images according to claim 2 is characterized in that: The method comprises: Performing a standard convolution operation on the features to be enhanced to obtain high-level features; Generate a single-channel attention map of the high-level features; Processing the single-channel attention map using an activation function to obtain a spatial attention mask; The product of the spatial attention weight of each channel in the spatial attention mask and the pixel value of the corresponding feature channel in the high-level feature is calculated to obtain the attention feature.
5. The lightweight and efficient small object detection method for aerial images according to claim 2 is characterized in that: The method comprises: In the channel dimension, the original features, the detail features, and the attention features are concatenated to obtain multiple channel features; The multiple channel features are fused into a first joint feature.
6. The lightweight and efficient small object detection method for aerial images according to claim 1 is characterized in that: The method comprises: Assigning a preset number of joint features to at least one alignment group; performing multi-scale alignment on the joint features in each alignment group to obtain at least one first alignment feature map; Obtaining remaining joint features from all the joint features, and performing multi-scale alignment on the remaining joint features and the context features to obtain a second aligned feature map; The at least one first aligned feature map and the second aligned feature map are fused to obtain a multi-scale feature map.
7. The lightweight and efficient small object detection method for aerial images according to claim 6 is characterized in that: The method comprises: Adjusting each joint feature in the alignment group to a predefined spatial size to obtain a plurality of features to be merged; Using a preset convolution kernel to perform a convolution operation on each of the features to be merged to adjust the feature channels of the features to be merged to obtain a plurality of adjusted features; In the channel dimension, performing feature splicing on the multiple adjusted features to obtain spliced features; A convolution operation is performed on the concatenated features using a unit-size convolution kernel to obtain a first aligned feature map.
8. The lightweight and efficient small object detection method for aerial images according to any one of claims 1 to 7, characterized in that: The method comprises: Performing a convolution operation on the last joint feature using a lightweight convolution kernel to obtain a feature to be processed; Using pooling kernels of different sizes to pool the features to be processed respectively to obtain multiple pooled features; wherein a pooling kernel of one size corresponds to one pooled feature; Taking the number of the plurality of pooled features as a total weight; determining a plurality of different learnable weights according to the total weight; wherein the sum of the plurality of different learnable weights is one; Using different learnable weights, weighting each of the pooled features to obtain a plurality of weighted features; wherein one learnable weight corresponds to one weighted feature; Feature concatenation is performed on the multiple weighted features to obtain multi-scale features; and a convolution operation is performed on the multi-scale features to obtain context features.
9. A lightweight and efficient target detection device for small object aerial images, characterized in that: The lightweight and efficient target detection device for small object aerial images includes at least one processor and a memory, wherein the at least one processor and the memory are connected via a data bus, and the memory stores instructions that can be executed by the at least one processor. After being executed by the processor, the instructions are used to implement the lightweight and efficient target detection method for small object aerial images described in any one of claims 1-8.
10. A non-volatile computer storage medium, characterized in that The computer storage medium stores computer-executable instructions, which are executed by one or more processors to implement the lightweight and efficient target detection method for small object aerial images according to any one of claims 1 to 8.