A scale imbalance remote sensing image target detection method, device and equipment
By introducing the frequency channel attention network and Soft-NMS algorithm to optimize remote sensing image target detection, the computational complexity and accuracy problems of multi-scale target detection are solved, efficient detection of small targets is achieved, and the impact of scale imbalance is alleviated.
Patent Information
- Application Number
- CN202510187279.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Existing remote sensing image target detection technology has problems such as high computational complexity, slow speed, noise interference and poor small target detection when processing multi-scale targets, especially in the case of scale imbalance, the detection accuracy is not high.
The frequency channel attention network is introduced for feature extraction, the Soft-NMS algorithm is used to eliminate redundant bounding boxes, and the EIoU loss function is used to optimize the predicted bounding box position. The improved Soft-EIoU-NMS method is used to improve the recall rate and reduce the false detection rate.
The model's ability to extract multi-scale target features is improved, the scale imbalance problem is alleviated, and detection accuracy and efficiency are improved, especially the detection ability of small targets.
Smart Images

Figure CN120107555B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image target detection, and in particular to a method, device, equipment and storage medium for scale-imbalanced remote sensing image target detection. Background Art
[0002] Early target detection techniques in remote sensing images mainly relied on image pyramids, sliding windows, and feature-based segmentation methods, but these methods are often computationally complex and have poor processing effects on small targets or low-resolution images.
[0003] In recent years, the emergence of deep learning, particularly convolutional neural networks (CNNs) and object detection frameworks such as YOLO, Faster R-CNN, and RetinaNet, has significantly improved the performance of multi-scale object detection. However, due to the large scale variation of objects, these technologies still face challenges such as high computational complexity, slow processing speed, noise interference, and small object detection. Many researchers are currently working to improve the efficiency and robustness of remote sensing imagery object detection through methods such as lightweight networks, optimization algorithms, and data augmentation. Summary of the Invention
[0004] The present invention provides a method, apparatus, device and storage medium for detecting targets in scale-imbalanced remote sensing images, aiming to improve the model's ability to extract multi-scale target features and alleviate the scale imbalance problem.
[0005] In a first aspect, an embodiment of the present invention provides a method for detecting targets in scale-imbalanced remote sensing images, comprising:
[0006] Extracting targets from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network;
[0007] Redundant bounding boxes are eliminated through the soft non-maximum suppression Soft-NMS algorithm, where EIoU is used instead of IoU in NMS;
[0008] EIoU loss is used as the regression loss function to determine the position of the predicted bounding box.
[0009] In a second aspect, an embodiment of the present invention provides a device for detecting objects in scale-imbalanced remote sensing images, comprising:
[0010] The target extraction module is used to extract targets from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network;
[0011] The bounding box elimination module is used to eliminate redundant bounding boxes through the soft non-maximum suppression Soft-NMS algorithm, in which EIoU is used instead of IoU in NMS;
[0012] The predicted bounding box position module is used to determine the position of the predicted bounding box using EIoU loss as the regression loss function.
[0013] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0014] one or more processors;
[0015] a memory for storing one or more programs;
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the scale-imbalanced remote sensing image target detection method provided by any embodiment of the present invention.
[0017] In a fourth aspect, an embodiment of the present invention provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform a scale-unbalanced remote sensing image target detection method as provided in any embodiment of the present invention.
[0018] The embodiments of the present invention provide a method, apparatus, device, and storage medium for detecting targets in scale-imbalanced remote sensing images. By improving the feature extraction network, loss function, and inference post-processing steps of the basic model, the present invention solves the problems of low multi-scale target detection accuracy, excessive computational complexity, and poor small target detection in the prior art. This improves the model's ability to extract multi-scale target features and alleviates the scale imbalance problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A flowchart of a method for detecting targets in scale-imbalanced remote sensing images provided in Example 1 of the present invention;
[0020] Figure 2 A schematic structural diagram of a device for detecting targets in scale-imbalanced remote sensing images provided in the second embodiment of the present invention;
[0021] Figure 3 A schematic structural diagram of an electronic device provided in a third embodiment of the present invention;
[0022] Figure 4 Schematic diagram of an improved frequency channel attention module in Example 1 of the present invention;
[0023] Figure 5 Schematic diagram of EIoU parameters in embodiment 1 of the present invention. DETAILED DESCRIPTION
[0024] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0025] Example 1
[0026] Figure 1 This is a flowchart of a method for detecting targets in scale-imbalanced remote sensing images provided in Example 1 of the present invention. This embodiment is applicable to extracting targets from remote sensing images, especially remote sensing images with large differences in target scales. The method can be performed by a device for detecting targets in scale-imbalanced remote sensing images. The device can be implemented by hardware and / or software and can generally be integrated into an electronic device, such as a computer. The method specifically includes:
[0027] Step 110: extracting the target from the remote sensing image frame by introducing a feature extraction network with a frequency channel attention network;
[0028] During the feature extraction phase, a frequency channel attention network (FcaNet) can be introduced. Furthermore, in the embodiments of this application, the frequency channel attention network is improved and integrated into the feature extraction network to fully capture rich features across all channel dimensions, such as feature information in different directions and frequencies. This is used to enhance the neural network's ability to extract semantic information and improve the model's ability to detect small-scale objects.
[0029] Step 120: Eliminate redundant bounding boxes using a soft non-maximum suppression (Soft-NMS) algorithm.
[0030] Among them, Soft-NMS is introduced in the reasoning stage and EIoU is used instead of IoU in NMS, that is, the Soft-EIoU-NMS method is used in the post-reasoning processing stage to improve the recall rate of objects of different scales in overlapping boxes and reduce the false detection rate.
[0031] Step 130: Use EIoU loss as the regression loss function to determine the position of the predicted bounding box
[0032] Continuously refining the predicted bounding box position is achieved by defining a bounding box regression loss function, calculating the regression loss value, and then performing reverse gradient descent. The EIoU loss is introduced as a regression loss function to mitigate the impact of scale-imbalanced object instances on the accuracy of loss function optimization.
[0033] The technical solution of this embodiment solves the problems of low multi-scale target detection accuracy, excessive computational complexity, and poor small target detection effect in the prior art by improving the feature extraction network, loss function, and inference post-processing steps of the basic model. It improves the model's ability to extract multi-scale target features and alleviates the scale imbalance problem.
[0034] Optionally, before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network, the method further includes:
[0035] Use the CSPDarknet53 network module as the backbone network module;
[0036] An improved frequency channel attention module is added between the SPPF module and the last layer C3 module to construct a feature extraction network.
[0037] Among them, the improved frequency channel attention module Att (such as Figure 4 As shown in the figure, the weighted channel and spatial feature information are input into the feature fusion network. Since small targets have higher energy in the frequency channel, they will be given a larger weight to increase the network's attention to small-scale target instances, improve the feature extraction ability of the neural network, and suppress interfering background information. The Att module is based on the FcaNet idea. First, the input X∈R C×H×W Segmentation is performed on channel C, where H and W represent height and width respectively, and the channel is divided into multiple small blocks [X0, X1, X2, ..., X n-1 ], each X i ∈R C’×H×W , i belongs to {0, 1, 2, ..., n-1}; where C' is the number of channels in the segmented patch Xi, C' = C / n. The segmented features are then processed through a linear layer for dimensionality reduction, resulting in a feature map that is downsampled 16 times. The Leaky ReLU activation function is then used to increase nonlinear expression capability. The activated feature map is then processed through a linear layer for dimensionality increase, resulting in a feature map that is upsampled 16 times, ensuring the same number of channels as the original input channels at the same layer. Finally, a Sigmoid function is used to obtain attention weights between [0, 1]. The resulting adaptive weight factor is then weighted and combined with the original feature map at the same layer.
[0038] Optionally, before eliminating redundant bounding boxes using the soft non-maximum suppression (Soft-NMS) algorithm, the method further includes:
[0039] Soft-NMS is introduced into the network model to eliminate redundant bounding boxes;
[0040] EIoU is used instead of IoU in NMS, where the EIoU parameter diagram is as follows Figure 5 As shown, EIoU is expressed as
[0041]
[0042] Among them, c w and c h b and b represent the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively. gt Denote the center point coordinates of the predicted bounding box and the true annotation bounding box, w and w respectively. gt and h and h gt Denote the width and height of the predicted bounding box and the true annotation bounding box, respectively, ρ denotes the Euclidean distance between the center points of the predicted bounding box and the true annotation bounding box, and b gt =(x gt ,y gt ) represents the coordinates of the center point of the true annotation bounding box; b = (x1, y1) represents the coordinates of the center point of the predicted bounding box, and the gt superscript is used to represent the true annotation (ground truth).
[0043] Optionally, the using EIoU loss as a regression loss function to determine the position of the predicted bounding box includes:
[0044] When using EIOU loss as the positioning regression loss function, it is defined as follows:
[0045]
[0046] Optionally, the using EIoU loss as a regression loss function to determine the position of the predicted bounding box includes:
[0047] When the EIoU between the predicted bounding boxes is greater than the threshold, a confidence score is given, where the larger the EIoU, the lower the confidence score given. The confidence score is obtained by the following formula:
[0048]
[0049] Among them, N t is the set EIoU threshold, b max is the predicted bounding box with the highest confidence score in a certain class, b i is the i-th predicted bounding box, s i is the confidence score of the i-th predicted bounding box, S i The above formula (i.e. the capital S on the left side of the equal sign) i This calculation formula) calculates the confidence score of the i-th predicted bounding box, σ is a parameter, for example, σ = 0.5 in the experiment.
[0050] The Soft-EIoU-NMS algorithm used in the NMS method directly removes the predicted bounding boxes whose IoU is greater than the threshold, which will cause the problem of missing detection of overlapping and dense target instances; when the EIoU of two predicted bounding boxes is greater than the threshold, it does not directly remove them, but gives a relatively low confidence score (the removal described above is to set the confidence value to 0), and follows the principle that the larger the EIoU, the lower the confidence score.
[0051] The technical solution of the embodiment of the present application can not only improve the accuracy and efficiency of target detection, but also be applicable to multiple fields such as remote sensing, unmanned driving, security monitoring, etc., and has broad application prospects and market value. The frequency channel attention network is introduced and improved to fully obtain feature information in different directions and frequencies; the semantic information extraction capability of the neural network is improved, and the model's detection capability for small-scale targets is improved. The Soft-EIoU-NMS method is used in the post-inference processing stage to improve the recall rate of targets of different scales within overlapping bounding boxes and reduce the false detection rate. The EIoU loss is introduced as a regression loss function to alleviate the accuracy impact of scale-unbalanced target instances on the loss function optimization.
[0052] Example 2
[0053] Figure 2 A schematic diagram of a device for detecting a target in a scale-unbalanced remote sensing image according to the second embodiment of the present invention is shown in FIG. Figure 2 As shown, the scale-unbalanced remote sensing image target detection device includes: a target extraction module 210, a bounding box elimination module 220 and a bounding box position prediction module 230, wherein:
[0054] The target extraction module 210 is used to extract the target from the remote sensing image frame by introducing a feature extraction network of a frequency channel attention network;
[0055] A bounding box elimination module 220 is configured to eliminate redundant bounding boxes using a soft non-maximum suppression (Soft-NMS) algorithm, wherein EIoU is used instead of IoU in NMS.
[0056] The predicted bounding box position module 230 is configured to determine the position of the predicted bounding box using the EIoU loss as a regression loss function.
[0057] Optionally, the scale-imbalanced remote sensing image target detection device further includes:
[0058] A backbone network module determination module, configured to use the CSPDarknet53 network module as a backbone network module before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network;
[0059] The network improvement module is used to add an improved frequency channel attention module between the SPPF module and the last layer C3 module to build a feature extraction network.
[0060] Optionally, the scale-imbalanced remote sensing image target detection device further includes:
[0061] A Soft-EIoU-NMS algorithm building module is used to introduce Soft-NMS into the network model to eliminate redundant bounding boxes before eliminating redundant bounding boxes through the soft non-maximum suppression Soft-NMS algorithm;
[0062] EIoU is used instead of IoU in NMS, where EIoU is expressed as
[0063]
[0064] Among them, c w and c h b and b represent the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively. gt Denote the center point coordinates of the predicted bounding box and the true annotation bounding box, w and w respectively. gt and h and h gt Represent the width and height of the predicted bounding box and the true annotation bounding box respectively.
[0065] Optional, predicted bounding box location module, used to:
[0066] When using EIOU loss as the positioning regression loss function, it is defined as follows:
[0067]
[0068] Optional, predicted bounding box location module, used to:
[0069] When the EIoU between the predicted bounding boxes is greater than the threshold, a confidence score is given, where the larger the EIoU, the lower the confidence score given. The confidence score is obtained by the following formula:
[0070]
[0071] Among them, N t is the set EIoU threshold, b max is the predicted bounding box with the highest confidence score in a certain class, b i is the i-th predicted bounding box, s i is the confidence score of the i-th predicted bounding box.
[0072] The scale-imbalanced remote sensing image target detection device provided by the embodiment of the present invention can execute the scale-imbalanced remote sensing image target detection method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0073] Example 3
[0074] Figure 3 This is a structural diagram of an electronic device provided in the third embodiment of the present invention, such as Figure 3 As shown, the electronic device includes a processor 310, a memory 320, an input device 330 and an output device 340; the number of processors 310 in the electronic device can be one or more. Figure 3 In the figure, a processor 310 is used as an example; the processor 310, memory 320, input device 330 and output device 340 in the electronic device can be connected via a bus or other means. Figure 3 The bus connection is taken as an example.
[0075] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the scale-imbalanced remote sensing image target detection method in the embodiments of the present invention (for example, the target extraction module 210, the bounding box elimination module 220, and the predicted bounding box position module 230 in the scale-imbalanced remote sensing image target detection apparatus). The processor 310 executes the software programs, instructions, and modules stored in the memory 320 to execute various functional applications and data processing of the electronic device, thereby implementing the scale-imbalanced remote sensing image target detection method described above.
[0076] The memory 320 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal, etc. Furthermore, the memory 320 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include a memory remotely located relative to the processor 310, and these remote memories may be connected to the electronic device via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0077] The input device 330 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. The output device 340 may include a display device such as a display screen.
[0078] Example 4
[0079] Embodiment 4 of the present invention further provides a storage medium containing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to perform a method for detecting targets in scale-imbalanced remote sensing images, including:
[0080] Extracting targets from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network;
[0081] Redundant bounding boxes are eliminated through the soft non-maximum suppression Soft-NMS algorithm, where EIoU is used instead of IoU in NMS;
[0082] EIoU loss is used as the regression loss function to determine the position of the predicted bounding box.
[0083] Of course, the storage medium containing computer-executable instructions provided in an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in the scale-unbalanced remote sensing image target detection method provided in any embodiment of the present invention.
[0084] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0085] It is worth noting that in the embodiment of the above-mentioned scale-unbalanced remote sensing image target detection device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0086] Although the present invention has been described in detail above using general explanations, specific embodiments, and experiments, it will be apparent to those skilled in the art that modifications and improvements may be made based on the present invention. Therefore, such modifications and improvements, which do not depart from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A method for detecting targets in scale-imbalanced remote sensing images, characterized in that: include: Extracting targets from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network; Redundant bounding boxes are eliminated through the soft non-maximum suppression Soft-NMS algorithm, where EIoU is used instead of IoU in NMS; Use EIoU loss as the regression loss function to determine the position of the predicted bounding box; Before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network, the method further includes: Use the CSPDarknet53 network module as the backbone network module; An improved frequency channel attention module is added between the SPPF module and the last layer C3 module to build a feature extraction network; Before eliminating redundant bounding boxes by the soft non-maximum suppression Soft-NMS algorithm, the method further includes: Soft-NMS is introduced into the network model to eliminate redundant bounding boxes; EIoU is used instead of IoU in NMS, where EIoU is expressed as ; in, c w and c h Represents the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively, b and b gt Represent the center point coordinates of the predicted bounding box and the true annotation bounding box respectively, w and w gt as well as h and h gt Represent the width and height of the predicted bounding box and the true annotation bounding box respectively, ρ represents the Euclidean distance between the center points of the predicted bounding box and the true annotation bounding box, b gt =(x gt ,y gt ), b= (x1, y1); The EIoU loss is used as the regression loss function to determine the position of the predicted bounding box, including: When using EIOU loss as the positioning regression loss function, it is defined as follows: ; The EIoU loss is used as the regression loss function to determine the position of the predicted bounding box, including: When the EIoU between the predicted bounding boxes is greater than the threshold, a confidence score is given, where the larger the EIoU, the lower the confidence score given. The confidence score is obtained by the following formula: ; in, N t is the EIoU threshold set, b max is the predicted bounding box with the highest confidence score in a certain class, b i is the i-th predicted bounding box, s i is the confidence score of the i-th predicted bounding box, and σ is a parameter set in the experiment.
2. A scale-unbalanced remote sensing image target detection device, characterized in that: include: The target extraction module is used to extract targets from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network; The bounding box elimination module is used to eliminate redundant bounding boxes through the soft non-maximum suppression Soft-NMS algorithm, in which EIoU is used instead of IoU in NMS; The predicted bounding box position module is used to determine the position of the predicted bounding box using EIoU loss as the regression loss function; The scale-unbalanced remote sensing image target detection device also includes: A backbone network module determination module, configured to use the CSPDarknet53 network module as a backbone network module before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network; The network improvement module is used to add an improved frequency channel attention module between the SPPF module and the last layer C3 module to build a feature extraction network; The scale-unbalanced remote sensing image target detection device also includes: A Soft-EIoU-NMS algorithm building module is used to introduce Soft-NMS into the network model to eliminate redundant bounding boxes before eliminating redundant bounding boxes through the soft non-maximum suppression Soft-NMS algorithm; EIoU is used instead of IoU in NMS, where EIoU is expressed as ; in, c w and c h Represents the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively, b and b gt Represent the center point coordinates of the predicted bounding box and the true annotation bounding box respectively, w and w gt as well as h and h gt Represent the width and height of the predicted bounding box and the true annotation bounding box respectively, ρ represents the Euclidean distance between the center points of the predicted bounding box and the true annotation bounding box, b gt =(x gt ,y gt ), b= (x1, y1); Predict bounding box locations module, which: When using EIOU loss as the positioning regression loss function, it is defined as follows: ; When the EIoU between the predicted bounding boxes is greater than the threshold, a confidence score is given, where the larger the EIoU, the lower the confidence score given. The confidence score is obtained by the following formula: ; in, N t is the EIoU threshold set, b max is the predicted bounding box with the highest confidence score in a certain class, b i is the i-th predicted bounding box, s i is the confidence score of the i-th predicted bounding box, and σ is a parameter set in the experiment.
3. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the scale-unbalanced remote sensing image target detection method as described in claim 1.
4. A storage medium containing computer-executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to perform the scale-imbalanced remote sensing image target detection method according to claim 1.