A small target detection method, device, and storage medium
By using the combination of enhanced feature extraction network and detailed feature optimization module in small object detection, the problem of low detection accuracy of small object is solved, and efficient detection of small objects is achieved.
Patent Information
- Application Number
- CN202510397756.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing technology has low detection accuracy of small and medium-sized targets, insufficient fusion of multi-scale features, and easy loss of detailed information in deep features, resulting in frequent missed detection and missed detection.
The context-aware aggregation unit in the enhanced feature extraction network is used to perform multi-scale adaptive feature fusion, and the geometric and local details of deep features are enhanced by combining the detailed feature optimization module, and the prediction layer feature map is generated through the optimization processing of global enhanced feature maps and deep feature maps.
It improves the accuracy and robustness of small object detection, effectively retains local and global information of the image, and enhances the detection ability of small objects.
Smart Images

Figure CN119919739B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer vision object detection, and particularly to a small object detection method, device, and storage medium. Background Art
[0002] In the field of computer vision, small object detection is a highly challenging task that requires accurate identification and localization of small-sized objects in images. Due to their small size, blurred edges, and susceptibility to background interference, it is difficult to extract their effective information, making detection difficult. However, small object detection plays a fundamental role in many applications such as autonomous driving, marine monitoring, and medical image diagnosis. Therefore, how to effectively improve the performance of small object detection has always been a hot topic and a difficult point in the research of related fields.
[0003] In recent years, although significant progress has been made in deep learning and small object detection technologies, small object detection still faces many challenges. First, although multi-scale feature fusion technology can, to a certain extent, retain the fine-grained information of the shallow layer and the semantic information of the deep layer, improving object detection performance, there are still deficiencies in small object detection. The main reason is that multi-scale feature fusion is usually performed at the feature pyramid stage or in deeper network layers, rather than at the feature extraction stage of the early backbone network. Second, deep learning models can automatically extract multi-level and multi-scale features from input data by constructing and training complex neural networks, thereby achieving accurate detection of objects. However, as the number of layers of the deep neural network increases, the receptive field of the network gradually expands, and the resolution of the feature map decreases accordingly. This results in the easy loss or masking of the detailed information of small objects by the features of large objects, which seriously restricts the performance of small object detection and leads to frequent cases of missed detection and false detection.
[0004] Currently, there is no effective solution to the problem of low detection accuracy for small objects in related technologies. Summary of the Invention
[0005] Embodiments of the present application provide a small object detection method, device, system, electronic device, and storage medium to at least solve the problem of low detection accuracy for small objects in related technologies.
[0006] In a first aspect, embodiments of the present application provide a small object detection method, including:
[0007] Obtain an image to be processed;
[0008] Input the image to be processed into a trained small object detection model; wherein, the small object detection model includes an enhanced feature extraction network and a detail feature optimization module;
[0009] Through multiple context-aware aggregation units in the enhanced feature extraction network, perform local feature extraction processing on the image to be processed, generate multiple local feature maps and deep feature maps with different resolutions, and fuse the multiple local feature maps to generate a global enhanced feature map;
[0010] Through the detail feature optimization module, perform optimized detail processing on the input global enhanced feature map and the deep feature map to generate a prediction layer feature map;
[0011] Based on the prediction layer feature map, obtain the object detection result.
[0012] In some embodiments, the step of performing local feature extraction processing on the image to be processed through multiple context-aware aggregation units in the enhanced feature extraction network to generate multiple local feature maps with different resolutions includes:
[0013] Through multiple context-aware aggregation units in the enhanced feature extraction network, perform convolution processing on the image to be processed to obtain multiple initial feature maps with different resolutions;
[0014] Determine the target feature map from each of the initial feature maps, and perform upsampling on the remaining feature maps in the initial feature maps except the target feature map to obtain multiple local feature maps with the same scale.
[0015] In some embodiments, the step of performing convolution processing on the image to be processed through multiple context-aware aggregation units in the enhanced feature extraction network to obtain multiple initial feature maps with different resolutions includes:
[0016] Perform a convolution operation on the image through the context-aware aggregation unit to generate a first initial feature map, and perform a depthwise separable convolution operation on the first initial feature map to obtain a second initial feature map; the initial feature map includes the first initial feature map and the second initial feature map;
[0017] The step of determining the target feature map from each of the initial feature maps includes:
[0018] Determine the first initial feature map in the initial feature map as the target feature map.
[0019] In some embodiments, the step of fusing the multiple local feature maps to generate a global enhanced feature map includes:
[0020] Through the context-aware aggregation unit, global pooling is respectively performed on each of the local feature maps to generate a plurality of pooled local feature maps; and spatial adaptive weight assignment processing is respectively performed on the pooled local feature maps to obtain first weight parameters corresponding to each of the local feature maps, and based on the first weight parameters, the plurality of local feature maps are fused to generate the global enhanced feature map.
[0021] In some embodiments, the fusing the plurality of local feature maps based on the first weight parameters to generate the global enhanced feature map includes:
[0022] Performing weight normalization processing on each of the first weight parameters to obtain second weight parameters;
[0023] Based on the second weight parameters, the plurality of local feature maps are fused to generate the global enhanced feature map.
[0024] In some embodiments, the detail feature optimization module includes a spatial attention mechanism and a channel attention mechanism; the optimizing the details of the input global enhanced feature map and the deep feature map through the detail feature optimization module to generate a prediction layer feature map includes:
[0025] Performing a spatial recombination operation on the global enhanced feature map through the detail feature optimization module to generate a shallow recombination feature map; wherein, the size of the shallow recombination feature map is the same as that of the deep feature map;
[0026] Let the shallow recombination feature map pass through the spatial attention mechanism and the channel attention mechanism respectively to extract a spatial weight and a channel weight; based on the spatial weight, the channel weight and the deep feature map, a prediction layer feature map is generated.
[0027] In some embodiments, the letting the shallow recombination feature map pass through the spatial attention mechanism and the channel attention mechanism respectively to extract a spatial weight and a channel weight includes:
[0028] Performing a convolution operation and global pooling on the shallow recombination feature map through the spatial attention mechanism to generate the spatial weight;
[0029] Performing a convolution operation and global average pooling on the shallow recombination feature map through the channel attention mechanism, and integrating through a fully connected layer in the detail feature optimization module to generate the channel weight.
[0030] In some embodiments, the method further includes:
[0031] Obtaining training images;
[0032] Input the training image into the initial feature extraction network for prediction to obtain a training enhanced feature map, and input the training enhanced feature map into the initial detail optimization module for prediction to obtain a training layer feature map;
[0033] Based on the training layer feature map, construct a loss function, and adjust the parameters of the detection model with the goal of minimizing the calculation result of the loss function to obtain the trained small target detection model.
[0034] In a second aspect, an embodiment of the present application provides a small target detection device, including:
[0035] An image acquisition module, configured to acquire an image to be processed;
[0036] A model detection module, configured to input the image to be processed into the trained small target detection model; wherein, the small target detection model includes an enhanced feature extraction network and a detail feature optimization module;
[0037] The model detection module is further configured to perform local feature extraction processing on the image to be processed via a plurality of context-aware aggregation units in the enhanced feature extraction network to generate a plurality of local feature maps and deep feature maps with different resolutions, and fuse the plurality of local feature maps to generate a global enhanced feature map;
[0038] The model detection module is further configured to perform optimized detail processing on the input global enhanced feature map and the deep feature map via the detail feature optimization module to generate a prediction layer feature map;
[0039] A result generation module, configured to obtain a target detection result based on the prediction layer feature map.
[0040] In a third aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the small target detection method as described in the first aspect above.
[0041] Compared with the related art, a small target detection method, device and storage medium provided by an embodiment of the present application balance the detail features of local small targets and the global context through the enhanced feature extraction network, and at the same time use the detail feature optimization module to strengthen the geometric and local detail features in the deep features, so that the network can retain the local and global information of the image at an earlier stage, reduce information loss, and strengthen the geometric and local detail features in the deep features, solve the problem of low detection accuracy of small targets, enhance the detection ability of small targets, and significantly improve the accuracy of small target detection.
[0042] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. Description of the Drawings
[0043] The drawings described herein are provided to further understand the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not unduly limit the present application. In the drawings:
[0044] Figure 1 is a hardware structure block diagram of a terminal of the small target detection method according to an embodiment of the present invention;
[0045] Figure 2 is a flowchart of the small target detection method according to an embodiment of the present application;
[0046] Figure 3 is an overall flowchart of the small target detection method according to a preferred embodiment of the present application;
[0047] Figure 4 is a flowchart of an enhanced feature extraction network of the small target detection method according to a preferred embodiment of the present application;
[0048] Figure 5 is a flowchart of a detailed feature optimization module of the small target detection method according to a preferred embodiment of the present application;
[0049] Figure 6 is a structure block diagram of a small target detection device according to an embodiment of the present application. Detailed Embodiments
[0050] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without making creative efforts fall within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and time-consuming, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.
[0051] Reference to "embodiment" in this application means that specific features, structures, or characteristics described in connection with an embodiment can be included in at least one embodiment of this application. The phrase may not necessarily refer to the same embodiment when it appears in various positions in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art will explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.
[0052] Unless otherwise defined, technical terms or scientific terms involved in this application should have the ordinary meaning understood by those of ordinary skill in the technical field to which this application belongs. The words "a", "an", "one kind", "the", and the like involved in this application do not indicate a quantity limitation and can represent a singular or plural number. The terms "include", "comprise", "have" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products, or devices. The terms "connect", "be connected", "couple" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0053] The method embodiment provided in this embodiment can be executed on a terminal, a computer, or a similar computing device. Taking running on a terminal as an example, Figure 1 is a hardware structure block diagram of a terminal for the small target detection method according to an embodiment of the present invention. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than Figure 1 shown in the figure, or have a structure different from Figure 1The different configurations shown.
[0054] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the small target detection method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0055] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0056] This embodiment provides a small target detection method. Figure 2 It is a flowchart of the small target detection method according to the embodiments of the present application, as Figure 2 shown, and this process includes the following steps:
[0057] Step S201, obtain the image to be processed;
[0058] Step S202, input the image to be processed into the trained small target detection model; wherein, the small target detection model includes an enhanced feature extraction network and a detail feature optimization module;
[0059] Among them, the image to be processed can come from various image acquisition devices, such as cameras, scanners, etc. The image to be processed contains small target objects, which need to be accurately detected and recognized through subsequent steps; the small target object specifically refers to a moving object with a small body size that appears in the required scenario; the small body size can be a body size smaller than a preset body size value. Usually, the small target object can be a small animal (such as a mouse, etc.), small sundries (such as falling leaves, paper, plastic bags, etc.) or other objects. In addition, although pedestrians and vehicles are important objects in video structured analysis, when they are far away from the camera, resulting in only occupying a very small part of the screen, the accuracy of their recognition and analysis may be affected; in this case, according to specific requirements, objects such as pedestrians and vehicles in the distance can also be considered as small target objects. The small target detection model includes an enhanced feature extraction network and a detailed feature optimization module. The enhanced feature extraction network is a Global-Local Enhanced Feature Extraction Network (GLEFENet), and the enhanced feature extraction network is composed of multiple stacked Context-aware hierarchical feature aggregation modules (CAHFAM). Inputting the image to be processed into the trained small target detection model, this step is the starting point of image processing and provides a basis for subsequent feature extraction and optimization.
[0060] Step S203: Through multiple context-aware aggregation units in the enhanced feature extraction network, perform local feature extraction processing on the image to be processed, generate multiple local feature maps and deep feature maps with different resolutions, and fuse the multiple local feature maps to generate a global enhanced feature map;
[0061] Among them, the image to be processed is first processed by multiple Context-Aware Hierarchical Feature Aggregation Modules (CAHFAMs) in the Global-Local Enhanced Feature Extraction Network (GLEFENet). Each CAHFAM uses a convolutional layer on the input features to obtain a feature map, and then continuously uses depthwise separable convolutions to successively obtain feature maps of different resolutions, generating multiple local feature maps of different resolutions. These feature maps have receptive fields of different sizes and context-aware capabilities at different scales. Then, multiple local feature maps are fused to generate a global enhanced feature map, which includes multi-scale local details and global semantic information. The deep feature map is gradually formed in the deep stage of the feature extraction network through successive operations such as convolution and pooling. In the GLEFENet, since CAHFAMs are used instead of traditional convolutional blocks, the deep feature map is gradually formed during the process of fusing local details and global semantic information and retains rich context information. In this step, through the Context-Aware Hierarchical Feature Aggregation Modules (CAHFAMs) of the Global-Local Enhanced Feature Extraction Network (GLEFENet), multi-scale adaptive feature fusion can be performed at the feature extraction stage, which can retain the local and global information of the image at an earlier stage and balance the local information and global information. Since it fuses multi-scale local details and global semantic information, GLEFENet can reduce the possibility that small target features are gradually submerged in the network, thereby improving the performance of small target detection.
[0062] Step S204: Through the Fine Characterization Refinement Module, perform fine-detail processing on the input global enhanced feature map and deep feature map to generate a prediction layer feature map.
[0063] Among them, the Fine Characterization Refinement Module (FCRM) uses the rich detail information in the global enhanced feature map (shallow feature map) to supplement the deep feature map, so as to enhance the geometric and local detail features in the deep feature map. The global enhanced feature map and the deep feature map are input into the trained Fine Characterization Refinement Module (FCRM) for fine-detail processing. The FCRM uses the rich detail information in the global enhanced feature map to supplement the deep features, obtains the geometric features of the target, and enhances the interaction between semantic information and detail information to generate a prediction layer feature map. Since the deep feature map usually contains more abstract and global feature information but is prone to losing detail information, in this step, the FCRM module fuses the detail information in the shallow feature map, effectively enhancing the geometric and local detail features in the deep feature map and improving the detection accuracy of the model for small targets.
[0064] Step S205: Based on the prediction layer feature map, obtain the target detection result.
[0065] Among them, based on the feature map of the prediction layer, the final object detection result is obtained by using existing object detection algorithms (such as bounding box regression, classification, etc.). The object detection result includes the position information of the object (such as the center point coordinates, length and width, etc.) and the classification information (such as the category to which the object belongs).
[0066] Through the above steps, the present application inputs the image to be processed into the trained small object detection model. The enhanced feature extraction network in the small object detection model replaces the traditional convolutional block with the context-aware hierarchical feature aggregation module (CAHFAM), and performs multi-scale adaptive feature fusion at the early stage of feature extraction. In the feature extraction layer of this model, the image to be processed is processed into multiple local feature maps with different resolutions. Then, these multiple local feature maps are fused to achieve the feature enhancement effect, so that the generated global enhanced feature map not only contains the detailed information of the local features, but also incorporates the global context information, generating the global enhanced feature map. Therefore, the global enhanced feature map pays attention to both local information and global information of the image, achieving the balance between local and global information; the deep feature map is gradually formed in the process of fusing local details and global semantic information. The deep feature map has a larger receptive field and can capture more abstract and global feature information, but it is often easy to lose the detailed information of small objects. Therefore, in order to make up for this defect, the global enhanced feature map and the deep feature Figure 1 are input into the trained fine-grained detail refinement module (FCRM) together. This model uses the rich detailed information in the global enhanced feature map to supplement the deep features, strengthens the geometric and local detailed features in the deep features, enhances the interaction between semantic information and detailed information, and finally generates the feature map of the prediction layer. The feature map of the prediction layer contains more complete and accurate object feature information. Finally, the object detection result is output based on the feature map of the prediction layer. By jointly using the enhanced feature extraction network and the fine-grained detail refinement module, the present application solves the problem of low detection accuracy of small objects caused by difficult feature extraction, easy loss of detailed information, etc. in the existing small object detection methods, effectively improves the performance of the model in small object detection, and makes the detection result more accurate and reliable.
[0067] In some embodiments, through multiple context-aware hierarchical feature aggregation modules in the enhanced feature extraction network, the image to be processed is subjected to local feature extraction processing to generate multiple local feature maps with different resolutions, including:
[0068] Through multiple context-aware hierarchical feature aggregation modules in the enhanced feature extraction network, the image to be processed is subjected to convolutional processing to obtain multiple initial feature maps with different resolutions;
[0069] The target feature map is determined from each initial feature map, and the remaining feature maps in the initial feature map except the target feature map are upsampled to obtain multiple local feature maps with the same scale.
[0070] Among them, in the enhanced feature extraction network (Global-Local Enhanced Feature Extraction Network, GLEFENet), a Context-Aware Hierarchical Feature Aggregation Module (CAHFAM) is used to replace the traditional convolutional block. The input data passes through multiple convolutional layers to extract features from the image to be processed, obtaining multiple initial feature maps with different resolutions. One of the multiple initial feature maps is selected as the target feature map. Generally, the feature map obtained after the first convolutional processing is determined as the target feature map. Then, through upsampling processing, the scales of the remaining feature maps except the target feature map are adjusted to be the same as that of the target feature Figure 1 map, so as to facilitate subsequent feature fusion. Specifically, upsampling can be implemented by interpolation methods (such as bilinear interpolation, nearest neighbor interpolation, etc.) or other neural network layers (such as transposed convolutional layers). In this embodiment, through upsampling processing, local feature maps with different resolutions have the same scale, which provides convenience for subsequent feature fusion, ensures the consistency between feature maps during the fusion process, and enables the network to learn richer and more diverse feature representations by fusing feature information of different scales, which helps to improve the accuracy and robustness of small object detection.
[0071] In some embodiments, the image to be processed is subjected to convolutional processing through multiple Context-Aware Hierarchical Feature Aggregation Modules in the enhanced feature extraction network, obtaining multiple initial feature maps with different resolutions, including:
[0072] The image is subjected to a convolutional operation through the Context-Aware Hierarchical Feature Aggregation Module to generate a first initial feature map, and a depthwise separable convolutional operation is performed on the first initial feature map to obtain a second initial feature map; the initial feature maps include the first initial feature map and the second initial feature map;
[0073] Determining the target feature map from each initial feature map includes:
[0074] Determining the first initial feature map in the initial feature maps as the target feature map.
[0075] Among them, for the standard convolutional operation on the image to be processed, this step usually uses multiple convolutional kernels (or called filters) to slide on the image, perform weighted summation on each local area, and may add a bias term, and then pass through a non-linear activation function (such as ReLU) to obtain the first initial feature map. It should be noted that taking the first initial feature map obtained after the first convolutional processing as the target feature map is to ensure that the sizes of the feature maps output during the adjacent multi-scale adaptive feature fusion process do not differ too much, otherwise too much information will be lost, making it difficult to maintain the integrity of the target details and context information contained in the feature maps.
[0076] Next, a depthwise separable convolution operation can be performed on the first initial feature map to generate a second initial feature map. The depthwise separable convolution operation includes two steps: depthwise convolution and pointwise convolution. Depthwise convolution uses different convolution kernels for each input channel and outputs a feature map with the same number of channels as the input; pointwise convolution then uses a 1x1 convolution kernel to perform a linear combination between channels on the output of the depthwise convolution, thereby obtaining a new feature map. By adjusting the number of convolution kernels of the pointwise convolution, the number of output feature maps can be controlled. Based on the first initial feature map (target feature map), a depthwise separable convolution operation can generate multiple local feature maps with different resolutions, that is, the second initial feature maps. These feature maps are used for subsequent multi-scale feature fusion. Through the convolution operation in this embodiment, multi-level and multi-scale feature information can be effectively extracted from the original image. And through the depthwise separable convolution operation, multiple local feature maps with different resolutions can be generated. These feature maps contain rich local information and context information, which helps to improve the accuracy of small object detection. Figure 1 This is used for subsequent multi-scale feature fusion. Through the convolution operation in this embodiment, multi-level and multi-scale feature information can be effectively extracted from the original image. And through the depthwise separable convolution operation, multiple local feature maps with different resolutions can be generated. These feature maps contain rich local information and context information, which helps to improve the accuracy of small object detection.
[0077] In some embodiments, multiple local feature maps are fused to generate a globally enhanced feature map, including:
[0078] Via the context-aware aggregation unit, global pooling is respectively performed on each local feature map to generate multiple pooled local feature maps; and spatial adaptive weight assignment processing is respectively performed on the pooled local feature maps to obtain the first weight parameter corresponding to each local feature map. And based on the first weight parameter, multiple local feature maps are fused to generate a globally enhanced feature map.
[0079] Among them, for each local feature map, global average pooling (GAP) and global max pooling (GMP) operations are respectively performed. The global pooling operation can extract the global information of the feature map, that is, perform an average or maximum operation on the entire feature map to obtain a fixed-length output vector. Through global pooling, each local feature map is converted into a pooled local feature map with global context information. For each pooled local feature map, a convolutional layer and an activation function (such as the Sigmoid function) are used to perform adaptive weight adjustment for spatial positions. This process assigns an adaptive weight to each spatial position, that is, the first weight parameter, and the first weight parameter reflects the importance of this position in the feature fusion process. The local feature maps are weighted and fused with their corresponding first weight parameters. Specifically, each local feature map is multiplied element-wise with its spatial adaptive weight to emphasize or suppress the features at specific positions. The weighted local feature maps are fused to generate a globally enhanced feature map, so that the globally enhanced feature map combines the information from feature maps of different scales and forms a global feature representation containing richer context information. During the fusion process, the weight parameter of each enhanced feature will determine its contribution degree in the final fused feature map. Through the fusion process, the network can generate a globally enhanced feature map containing features of multiple scales and multiple positions. This feature map not only retains the global context information but also highlights the key local details. The adaptive weight adjustment process is shown in the following formula:
[0080] ;
[0081] In the above formula, (i = 0, 1, 2) represents the local feature map F i The feature map after size adjustment; GAP and GMP respectively represent the global average pooling and global max pooling operations; CAT is the feature concatenation operation; Conv represents the convolutional operation; σ is the Sigmoid activation function; W i (i = 0, 1, 2) represents the first weight parameter of the local feature map F i .
[0082] In this embodiment, through the global pooling operation, each local feature map is converted into a feature representation with global context information, which helps the network better understand and utilize the global structure information of the image during the feature fusion process; the spatial adaptive weight assignment process assigns different weights to different spatial positions of each local feature map, enabling more accurate retention and emphasis of key information and suppression of irrelevant information during the feature fusion process; the globally enhanced feature map fuses the information from feature maps of different scales and emphasizes the feature representation of the key regions, enabling the network to better identify and locate small targets in the image, thereby improving the performance of small target detection.
[0083] In some of these embodiments, based on the first weight parameter, multiple local feature maps are fused to generate a global enhanced feature map, including:
[0084] Perform weight normalization on each first weight parameter to obtain a second weight parameter;
[0085] Based on the second weight parameter, fuse multiple local feature maps to generate a global enhanced feature map.
[0086] Among them, after determining each local feature map and its first weight parameter, perform normalization processing on these first weight parameters to obtain second weight parameters. Normalization processing usually uses the Softmax function or a similar mechanism to map the weight parameters to the interval [0, 1] and ensure that the sum of all weight parameters is 1. The weight normalization processing in this embodiment is to ensure that in the feature fusion process, the contribution of each global enhanced feature map can be balanced according to its importance, avoiding the loss or over-amplification of feature information caused by too large or too small weights.
[0087] The specific process is shown in the following formula:
[0088] ;
[0089] ;
[0090] In the above formula, H represents the normalized second weight parameter; Softmax refers to the normalization exponential function; CAT is the feature concatenation operation; W i (i = 0, 1, 2) represents the first weight parameter of the local feature map F i ; (i = 0, 1, 2) represents the feature map after size adjustment of the local feature map F i ; F out is the global enhanced feature map output after passing through the enhanced feature extraction network.
[0091] In some of these embodiments, the detail feature optimization module includes a spatial attention mechanism and a channel attention mechanism; through the detail feature optimization module, perform optimized detail processing on the input global enhanced feature map and deep feature map to generate a prediction layer feature map, including:
[0092] Through the detail feature optimization module, perform spatial reorganization operations on the global enhanced feature map to generate a shallow reorganization feature map; among them, the size of the shallow reorganization feature map is the same as that of the deep feature map;
[0093] Let the shallow reorganization feature map pass through the spatial attention mechanism and the channel attention mechanism respectively to extract the spatial weight and the channel weight; based on the spatial weight, the channel weight and the deep feature map, generate a prediction layer feature map.
[0094] Among them, the Detail Feature Optimization Module (FCRM) uses a recombination operation (i.e., the Split operation) to replace the traditional convolution or pooling operation to retain more shallow feature information. The Split operation rearranges the spatial elements of the feature into the channel dimension, so as to adjust the feature size to be consistent with the deep feature Figure 1 while being able to preserve the spatial encoding information, and a shallow recombined feature map is generated by this operation. The spatial attention mechanism makes the shallow recombined feature map go through a convolutional layer and a global pooling operation to generate spatial weights. This process focuses on the spatial position information of the feature map, and emphasizes or suppresses the features of specific regions by assigning different weights to different positions. The channel attention mechanism makes the shallow recombined feature map go through a convolutional layer and a global average pooling operation, and integrates through a fully connected layer to generate channel weights. This process focuses on the channel information of the feature map, and emphasizes or suppresses the features of specific channels by assigning different weights to different channels. Then, the extracted spatial weights and channel weights are used to re-weight the deep feature map. Specifically, the spatial weights and channel weights are respectively multiplied or added element-wise to the deep feature map to highlight the key information and suppress the irrelevant information. The re-weighted deep feature map is combined with the global enhanced feature map (or the appropriately processed shallow recombined feature map) to form the final prediction layer feature map. This process is shown in the following formula:
[0095] ;
[0096] In the above formula, refers to the prediction layer feature map output after being strengthened by the Detail Feature Optimization Module; represents the deep feature map; T S represents the function to obtain the spatial weights; T C represents the function to obtain the channel weights; represents the shallow feature map.
[0097] In this embodiment, the Detail Feature Optimization Module (FCRM) performs recombination and two-way weight extraction operations on the global enhanced feature map, which can effectively supplement the missing small target detail information in the deep feature map, enhance the interaction between semantic information and detail information, thereby improving the key feature extraction ability of small targets, significantly enhancing the accuracy of small target detection. The Detail Feature Optimization Module retains more shallow feature information and re-weights the deep feature map, which enables the model to more accurately capture the key features of small targets with different scales and complex backgrounds.
[0098] In some of these embodiments, the shallow recombined feature map is respectively passed through the spatial attention mechanism and the channel attention mechanism to extract the obtained spatial weights and channel weights, including:
[0099] Perform convolution operations and global pooling on the shallowly recombined feature map via the spatial attention mechanism to generate spatial weights.
[0100] Perform convolution operations and global average pooling on the shallowly recombined feature map via the channel attention mechanism, and integrate them via the fully connected layer in the detailed feature optimization module to generate channel weights.
[0101] Among them, after obtaining the shallowly recombined feature map, generate spatial weights via the spatial attention mechanism, perform convolution operations on the shallowly recombined feature map to extract the spatial location information of the feature map. The convolution operation can retain the local structural information of the feature map while reducing the computational amount. Then perform global pooling operations (such as global average pooling or global max pooling) on the convolved feature map to obtain the global context information of the feature map. Global pooling compresses the spatial dimension of the feature map into a vector of a fixed length, enabling each position's feature to contribute to the entire feature map. Finally, process the globally pooled feature vector through a convolutional layer and an activation function (such as the Sigmoid function) to generate spatial weights. The spatial weights reflect the importance of different positions in the feature map, emphasizing or suppressing the features of specific regions by assigning different weights to different positions. Generate channel weights via the channel attention mechanism. Similar to the spatial attention mechanism, first perform convolution operations on the shallowly recombined feature map to extract the channel information of the feature map. Then perform global average pooling operations on the convolved feature map to obtain the channel-level global context information of the feature map. Global average pooling compresses each channel of the feature map into a scalar value, reflecting the importance of that channel in the entire feature map. Input the globally averaged pooled feature vector into the fully connected layer for integration to generate channel weights. The fully connected layer can perform non-linear transformations on the feature vector and enhance the discriminability of the feature representation through activation functions (such as the ReLU or Sigmoid function). The channel weights reflect the importance of different channels in the feature map, emphasizing or suppressing the features of specific channels by assigning different weights to different channels. Simply put, the spatial weights can be adaptively adjusted through a convolutional layer, and global pooling (which can be global average pooling or global max pooling, global max pooling is usually used to capture the most significant features, and global average pooling is used to capture the overall features) is performed along the channels. Finally, spatial weights are obtained. The spatial weights reflect the importance of different spatial positions and are used to weight the feature map subsequently. At the same time, the channel weights are adaptively adjusted through another convolutional layer, and then global average pooling is used to aggregate the spatial information of each channel and obtain the global information. Finally, after passing through two fully connected layers, the channel weights are adaptively adjusted again. The channel weights reflect the importance of different channels and are also used to weight the feature map subsequently. The process is shown as follows:
[0102] ;
[0103] ;
[0104] Among them, CAP represents global pooling along the channel; GAP refers to the global average pooling operation; fc refers to the fully connected layer; δ represents the ReLu activation function; σ represents the Sigmoid activation function.
[0105] In this embodiment, through the generation of spatial weights, the model can pay attention to the importance of different spatial positions in the image, suppress irrelevant regions while highlighting key regions, which is particularly important for small object detection because small objects usually only occupy a small part of the image, and spatial weights can help the model focus more on the positions where these small objects are located; and through the generation of channel weights, the model can distinguish the importance of different channels, so as to more effectively utilize the information in the feature map. For the small object detection task, some channels may contain more key information about small objects, and channel weights can help the model capture these information more accurately; by combining spatial weights and channel weights, the Fine Detail Feature Refinement Module (FCRM) can strengthen the geometric and local detail features in the deep features, enhance the interaction between semantic information and detail information, improve the key feature extraction ability of small objects, thus significantly improving the accuracy of small object detection. At the same time, by adaptively adjusting the feature weights of different spaces and channels, the model can more robustly handle complex and changing scenarios.
[0106] In some of these embodiments, the method further includes:
[0107] Obtain training images;
[0108] Input the training images into the initial feature extraction network for prediction to obtain training enhanced feature maps, and input the training enhanced feature maps into the initial detail refinement module for prediction to obtain training layer feature maps;
[0109] Based on the training layer feature maps, construct a loss function, and adjust the parameters of the detection model with the goal of minimizing the calculation result of the loss function to obtain a trained small object detection model.
[0110] Among them, images containing small targets are selected from the training dataset. These images should cover different scenarios, lighting conditions, and target postures to ensure the generalization ability of the model. The training images are first passed through an initial feature extraction model, which uses a context-aware hierarchical feature aggregation module (CAHFAM) for multi-scale feature fusion to extract a training enhanced feature map containing local and global information. The training enhanced feature map is then input into an initial detail optimization model, which uses the rich detail information in the training enhanced feature map (shallow features) to supplement the deep features. Through steps such as the Split operation, channel and spatial bidirectional weight extraction, an enhanced training layer feature map is generated. The training layer feature map is used to predict object classification and location regression, and the loss function value is calculated based on the difference between the prediction result and the ground truth label. The loss function can include multiple parts such as classification loss, confidence loss, and regression loss. Through the backpropagation algorithm, the gradient of the loss function result is passed back to the initial feature extraction model and the initial detail optimization model. According to the gradient information, the weight parameters of the model are updated to minimize the loss function value. This process is carried out through multiple iterations until the model converges or reaches the preset number of training epochs. After iterative training, it gradually converges to form the final enhanced feature extraction network (GLEFENet) and the detail feature optimization module (FCRM), that is, a trained small target detection model is obtained.
[0111] Specifically, the loss function includes three categories: classification loss, confidence loss, and bounding box regression loss, and the formula is expressed as follows:
[0112] ;
[0113] ;
[0114] ;
[0115] ;
[0116] In the above formula, 、 and are the classification loss, confidence loss, and regression loss respectively; classes represents the target type; is the true / predicted confidence (score); represents (horizontal center position, vertical center position, width, height) of the predicted / true box; refers to the cross-entropy loss; H / W represents the height / width of the final prediction map; M represents the number of anchor boxes; indicates whether there is an object in the i-th cell and whether the j-th bounding box predictor in the i-th cell is responsible for this prediction.
[0117] In this embodiment, through the joint training of the enhanced feature extraction network and the detailed feature optimization module, the model can more effectively extract and utilize the local detailed features and global context information of small targets, thereby improving the accuracy and robustness of small target detection. And through the iterative training process using the loss function, the structure of the model (including weight parameters) is continuously optimized, making the model more efficient and accurate in feature extraction, classification, and location regression, etc.
[0118] The embodiments of the present application will be described and illustrated below through preferred embodiments.
[0119] Figure 3 is the overall flowchart of the small target detection method according to the preferred embodiment of the present application. As Figure 3 shown, the enhanced feature extraction network (i.e., the global-local enhanced feature extraction network GLEFENet) is integrated into the feature extraction network by using a context-aware aggregation unit instead of a convolutional block. In the early stage of the network, that is, during the feature extraction stage, multi-scale feature adaptive fusion is performed, and progressive feature fusion is continuously carried out at this stage to achieve a balance between local details and global semantics, reducing the possibility that small target features are gradually submerged in the network. In a convolutional neural network, as the network depth increases, its receptive field also becomes larger, and small target feature information is often submerged by large target feature information. How to retain small target features without losing them is crucial for small target detection models.
[0120] In the enhanced feature extraction network of the present application, a context-aware aggregation unit CAHFAM is proposed. This module performs interactive fusion through multi-level feature maps, enabling the network to not only focus on local detailed information such as the edges and textures of small targets but also possess rich global context information.
[0121] Figure 4 is the flowchart of the enhanced feature extraction network of the small target detection method according to the preferred embodiment of the present application. As Figure 4 shown, first, the input feature F is used with a convolutional layer to obtain the feature map F0. Then, depthwise separable convolutions are continuously used to obtain the feature maps F1 and F2 in sequence. At this time, three feature maps F 0、 F1 and F2 with different resolution receptive fields and different scales of context awareness are obtained. Then, an upsampling layer is used to adjust the sizes of F1 and F2 to be the same as the size of the F0 feature map, and finally, the adjusted feature maps are obtained. Subsequently, global average pooling and global max pooling are respectively performed on the feature obtained after adjusting the size to obtain sufficient global information, and then convolution and activation functions are performed to perform adaptive weight adjustment of the spatial positions, assigning an adaptive weight to each spatial position. This process is shown in the following formula:
[0122] ;
[0123] In the above formula, (i = 0, 1, 2) represents the local feature map F i the feature map after size adjustment; GAP and GMP respectively represent global average pooling and global max pooling operations; CAT is the feature concatenation operation; Conv represents the convolution operation; σ is the Sigmoid activation function; W i (i = 0, 1, 2) represents the local feature map F i of the first weight parameter.
[0124] At this time, W i (i = 0, 1, 2) is concatenated, and spatial position weight normalization is performed to obtain the normalized spatial position weight, that is, the second weight parameter. The second weight parameter is used to re-weight the feature map, so as to enable the network to achieve the perception ability of different spatial positions in the image, suppress irrelevant regions and highlight key regions at the same time. This process is shown in the following formula:
[0125] ;
[0126] ;
[0127] In the above formula, H represents the normalized second weight parameter; Softmax refers to the normalized exponential function; CAT is the feature concatenation operation; W i (i = 0, 1, 2) represents the local feature map F i of the first weight parameter; (i = 0, 1, 2) represents the local feature map F i the feature map after size adjustment; F out is the shallow feature map output after passing through the enhanced feature extraction network.
[0128] The detailed feature optimization module FCRM proposed in this application considers the problem that deep features have a larger receptive field, are more focused on generating abstract and global features for the target, and are more likely to lose the detailed information of the target. Losing detailed information is catastrophic for small targets, which means that the features of small targets are more likely to be lost in the deep feature layer than in the shallow feature layer. The detailed feature optimization module uses shallow features containing rich detailed information to supplement the local information of deep features, obtains the geometric features of the target, enhances the interaction between semantic information and detailed information, and strengthens the key features of small targets. Figure 5 is the flowchart of the detailed feature optimization module of the small target detection method according to the preferred embodiment of the present application, as shown in Figure 5As shown in the figure, in order to retain as much shallow feature information as possible, this module first uses the Split operation to replace the convolution or pooling operation, that is, rearranging the spatial elements of the feature into the channel dimension, so as to save the spatial encoding information while achieving the purpose of being consistent with the deep feature size, and at this time the feature is obtained. Next, channel and spatial bidirectional weight extraction is performed on the feature, the deep feature is reweighted, and the key information is highlighted bidirectionally, and finally the deep feature supplemented with shallow feature information is obtained. This process is shown in the following formula:
[0129] ;
[0130] In the above formula, refers to the predicted layer feature map output after being strengthened by the detailed feature optimization module; represents the deep feature map; T S represents the function to obtain the spatial weight; T C represents the function to obtain the channel weight; represents the shallow feature map.
[0131] The spatial weight and the channel weight are generated by the following operations. First, the spatial weight is adaptively adjusted through the convolutional layer, then global pooling is performed along the channel, and finally the spatially adaptive weight is obtained. At the same time, the channel weight is adaptively adjusted through another convolutional layer, and then global average pooling is used to aggregate the spatial information of each channel and obtain the global information. Finally, after two fully connected layers, the channel weight is adaptively adjusted again. This process is shown in the following formula:
[0132] ;
[0133] ;
[0134] Among them, CAP represents global pooling along the channel; GAP refers to the global average pooling operation; fc refers to the fully connected layer; δ represents the ReLu activation function; σ represents the Sigmoid activation function.
[0135] Next is the specific process of iteration through the loss function during the model training process:
[0136] Object detection mainly includes two purposes: classification and localization. Therefore, the loss function of the model proposed in this application includes three categories: classification loss, confidence loss, and bounding box regression loss. The formula is expressed as follows:
[0137] ;
[0138] ;
[0139] ;
[0140] ;
[0141] In the above formula, , and are the classification loss, confidence loss, and regression loss respectively; classes represents the target type; is the true / predicted confidence (score); represents (horizontal center position, vertical center position, width, height) of the predicted / true bounding box; refers to the cross-entropy loss; H / W represents the height / width of the final prediction map; M represents the number of anchor boxes; indicates whether there is an object in the i-th cell and whether the j-th bounding box predictor in the i-th cell is responsible for the prediction.
[0142] In summary, first, the enhanced feature extraction network proposed in this application (i.e., the global-local enhanced feature extraction network, GLEFENet) replaces the traditional convolutional block with a context-aware aggregation unit (CAHFAM), performs multi-scale adaptive feature fusion in the feature extraction stage, continuously fuses local details and global semantic information, and the context-aware aggregation unit balances the detailed features of local small targets and global context, reducing the risk that the features of small targets are overwhelmed by the features of large targets in the deep network, thereby improving the small target detection performance. Second, the detailed feature optimization module (FCRM) proposed in this application supplements the deep features with rich detailed information in the shallow features, solves the problem that the deep features are prone to losing the detailed information of small targets. This module strengthens the geometric and local detailed features in the deep features, enhances the interaction between semantic information and detailed information, thereby improving the key feature extraction ability of small targets and significantly improving the accuracy of small target detection. Finally, the small target detection method proposed in this application combines the enhanced feature extraction network and the detailed feature optimization module, effectively improving the performance of the model in small target detection.
[0143] This embodiment also provides a small target detection device, which is used to implement the above embodiment and preferred implementation manners, and those that have been described will not be repeated. As used below, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0144] Figure 6 is the structural block diagram of the small target detection device according to the embodiment of the present application, asFigure 6 As shown, the device includes:
[0145] An image acquisition module 10 for acquiring an image to be processed;
[0146] A model detection module 20 for inputting the image to be processed into a trained small object detection model; wherein, the small object detection model includes an enhanced feature extraction network and a detail feature optimization module;
[0147] The model detection module 20 is further configured to perform local feature extraction processing on the image to be processed via multiple context-aware aggregation units in the enhanced feature extraction network, generate multiple local feature maps and deep feature maps with different resolutions, and fuse the multiple local feature maps to generate a global enhanced feature map;
[0148] The model detection module 20 is further configured to perform optimized detail processing on the input global enhanced feature map and deep feature map via the detail feature optimization module to generate a prediction layer feature map;
[0149] A result generation module 30 for obtaining an object detection result based on the prediction layer feature map.
[0150] The above-mentioned small object detection device further includes a training module; the training module is configured to acquire training images; input the training images into an initial feature extraction network for prediction to obtain a training enhanced feature map, and input the training enhanced feature map into an initial detail optimization module for prediction to obtain a training layer feature map; construct a loss function based on the training layer feature map, and adjust the parameters of the detection model with the goal of minimizing the calculation result of the loss function to obtain a trained small object detection model.
[0151] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.
[0152] In addition, in combination with the small object detection method in the above-mentioned embodiments, the embodiments of the present application can be implemented by providing a storage medium. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the small object detection methods in the above-mentioned embodiments is implemented.
[0153] Those skilled in the art should understand that the technical features of the above-mentioned embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0154] The embodiments described above merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A small target detection method, characterized in that Including: Obtain the image to be processed; Input the image to be processed into the trained small object detection model; wherein, the small object detection model includes an enhanced feature extraction network and a detail feature optimization module; Via multiple context-aware aggregation units in the enhanced feature extraction network, perform local feature extraction processing on the image to be processed, generate multiple local feature maps and a deep feature map with different resolutions, and fuse the multiple local feature maps to generate a global enhanced feature map, including: Via the context-aware aggregation unit, perform global pooling on each local feature map respectively to generate multiple pooled local feature maps; and perform spatial adaptive weight assignment processing on each pooled local feature map respectively to obtain the first weight parameter corresponding to each local feature map, and based on the first weight parameter, fuse the multiple local feature maps to generate the global enhanced feature map, including: Perform weight normalization processing on each first weight parameter to obtain a second weight parameter; Based on the second weight parameter, fuse the multiple local feature maps to generate the global enhanced feature map; Via multiple context-aware aggregation units in the enhanced feature extraction network, perform local feature extraction processing on the image to be processed, generate multiple local feature maps with different resolutions, including: Via multiple context-aware aggregation units in the enhanced feature extraction network, perform convolution processing on the image to be processed to obtain multiple initial feature maps with different resolutions, including: Via the context-aware aggregation unit, perform a convolution operation on the image to generate a first initial feature map, and perform a depthwise separable convolution operation on the first initial feature map to obtain a second initial feature map; the initial feature map includes the first initial feature map and the second initial feature map; Determine the target feature map from each initial feature map, and perform upsampling on the remaining feature maps in the initial feature map except the target feature map to obtain multiple local feature maps with the same scale; The determining the target feature map from each initial feature map includes: Determine the first initial feature map in the initial feature map as the target feature map; Via the detail feature optimization module, perform detail optimization processing on the input global enhanced feature map and the deep feature map to generate a prediction layer feature map; Based on the prediction layer feature map, obtain the object detection result.
2. The small target detection method according to claim 1, wherein The detail feature optimization module includes a spatial attention mechanism and a channel attention mechanism; via the detail feature optimization module, perform detail optimization processing on the input global enhanced feature map and the deep feature map to generate a prediction layer feature map, including: Via the detail feature optimization module, perform spatial recombination operation on the global enhanced feature map to generate a shallow recombination feature map; wherein, the size of the shallow recombination feature map is the same as that of the deep feature map; Let the shallow recombined feature maps pass through the spatial attention mechanism and the channel attention mechanism respectively to extract spatial weights and channel weights; generate a predicted layer feature map based on the spatial weights, the channel weights, and the deep feature map.
3. The small target detection method according to claim 2, wherein The step of letting the shallow recombined feature maps pass through the spatial attention mechanism and the channel attention mechanism respectively to extract spatial weights and channel weights includes: Perform a convolution operation and global pooling on the shallow recombined feature map via the spatial attention mechanism to generate the spatial weights; Perform a convolution operation and global average pooling on the shallow recombined feature map via the channel attention mechanism, and integrate via the fully connected layer in the detail feature optimization module to generate the channel weights.
4. The small target detection method according to claim 1, wherein The method further includes: Obtain training images; Input the training images into the initial feature extraction network for prediction to obtain training enhanced feature maps, and input the training enhanced feature maps into the initial detail optimization module for prediction to obtain training layer feature maps; Construct a loss function based on the training layer feature maps, and adjust the parameters of the detection model with the goal of minimizing the calculation result of the loss function to obtain the trained small target detection model.
5. A small target detection device, characterized in that, It includes: An image acquisition module for acquiring an image to be processed; A model detection module for inputting the image to be processed into the trained small target detection model; wherein, the small target detection model includes an enhanced feature extraction network and a detail feature optimization module; The model detection module is further configured to perform local feature extraction processing on the image to be processed via multiple context-aware aggregation units in the enhanced feature extraction network to generate multiple local feature maps and deep feature maps with different resolutions, and fuse the multiple local feature maps to generate a global enhanced feature map, including: Perform global pooling on each of the local feature maps via the context-aware aggregation unit to generate multiple pooled local feature maps; and perform spatial adaptive weight assignment processing on each of the pooled local feature maps to obtain first weight parameters corresponding to each of the local feature maps, and fuse the multiple local feature maps based on the first weight parameters to generate the global enhanced feature map, including: Perform weight normalization processing on each of the first weight parameters to obtain second weight parameters; Fuse the multiple local feature maps based on the second weight parameters to generate the global enhanced feature map; The step of performing local feature extraction processing on the image to be processed via multiple context-aware aggregation units in the enhanced feature extraction network to generate multiple local feature maps with different resolutions includes: Perform convolution processing on the image to be processed via multiple context-aware aggregation units in the enhanced feature extraction network to obtain multiple initial feature maps with different resolutions, including: The context-aware aggregation unit performs a convolution operation on the image to generate a first initial feature map, and performs a depthwise separable convolution operation on the first initial feature map to obtain a second initial feature map; the initial feature maps include the first initial feature map and the second initial feature map; Determine a target feature map from each of the initial feature maps, and upsample the remaining feature maps in the initial feature maps except the target feature map to obtain a plurality of local feature maps of the same scale; The determining a target feature map from each of the initial feature maps includes: Determining the first initial feature map in the initial feature maps as the target feature map; The model detection module is further configured to perform detail feature optimization processing on the input global enhanced feature map and the deep feature map via a detail feature optimization module to generate a prediction layer feature map; A result generation module, configured to obtain a target detection result based on the prediction layer feature map.
6. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the small target detection method according to any one of claims 1 to 4 when running.
Citation Information
Patent Citations
Infrared small target detection method based on attention fusion feature pyramid network
CN116524312A
Model training method and device, infrared weak and small target detection method and device and electronic equipment
CN119295740A