Scale imbalance remote sensing image target detection method, device and equipment
By introducing frequency channel attention network and Soft-NMS algorithm in remote sensing image object detection, combined with EIoU loss function, the problem of scale imbalance in multi-scale object detection is solved, and the detection accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510187279.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
When processing multi-scale targets, existing remote sensing image object detection technology has high computational complexity, slow processing speed, poor noise interference and small target detection effects, making it difficult to effectively alleviate the problem of scale imbalance.
Feature extraction is performed by introducing a frequency channel attention network, the soft non-maximum suppression Soft-NMS algorithm is used to eliminate redundant bounding boxes, and the location of the predicted bounding boxes is determined using EIoU loss as a regression loss function.
The model's ability to extract multi-scale target features is improved, scale imbalance problem is alleviated, detection accuracy and efficiency are improved, and detection effect is improved on small targets.
Smart Images

Figure CN120107555A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image target detection, and in particular to a scale-imbalanced remote sensing image target detection method, device, equipment and storage medium. Background Art
[0002] Early target detection techniques in remote sensing images mainly relied on image pyramids, sliding windows, and feature-based segmentation methods, but these methods are often computationally complex and do not perform well on small targets or low-resolution images.
[0003] In recent years, the emergence of deep learning, especially the application of convolutional neural networks (CNN) and target detection frameworks such as YOLO, Faster R-CNN and RetinaNet, has significantly improved the performance of multi-scale target detection. However, due to the large scale differences of targets, these technologies still face challenges such as high computational complexity, slow processing speed, noise interference and small target detection. Many researchers are now working on improving them through lightweight networks, optimization algorithms and data enhancement methods to improve the efficiency and robustness of remote sensing image target detection. Summary of the invention
[0004] The present invention provides a method, device, equipment and storage medium for detecting a scale-imbalanced remote sensing image target, aiming to improve the model's ability to extract multi-scale target features and alleviate the scale imbalance problem.
[0005] In a first aspect, an embodiment of the present invention provides a method for detecting a target in a scale-unbalanced remote sensing image, comprising:
[0006] Extracting objects from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network;
[0007] Redundant bounding boxes are eliminated through the soft non-maximum suppression Soft-NMS algorithm, where EIoU is used instead of IoU in NMS;
[0008] The EIoU loss is used as the regression loss function to determine the location of the predicted bounding box.
[0009] In a second aspect, an embodiment of the present invention provides a scale-unbalanced remote sensing image target detection device, comprising:
[0010] The target extraction module is used to extract targets from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network;
[0011] The bounding box elimination module is used to eliminate redundant bounding boxes through the soft non-maximum suppression Soft-NMS algorithm, in which EIoU is used instead of IoU in NMS;
[0012] The predicted bounding box position module is used to determine the position of the predicted bounding box using EIoU loss as the regression loss function.
[0013] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0014] one or more processors;
[0015] A memory for storing one or more programs;
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the scale-unbalanced remote sensing image target detection method provided by any embodiment of the present invention.
[0017] In a fourth aspect, an embodiment of the present invention provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute a scale-unbalanced remote sensing image target detection method as provided in any embodiment of the present invention.
[0018] The embodiments of the present invention provide a method, device, equipment and storage medium for detecting targets in scale-imbalanced remote sensing images. By improving the feature extraction network, loss function and inference post-processing links of the basic model, the problems of low multi-scale target detection accuracy, high computational complexity and poor small target detection effect in the prior art are solved, thereby improving the model's ability to extract multi-scale target features and alleviating the scale imbalance problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A flowchart of a method for detecting a target in a scale-imbalanced remote sensing image provided in the first embodiment of the present invention;
[0020] Figure 2 A schematic diagram of the structure of a device for detecting a target in a scale-unbalanced remote sensing image provided by the second embodiment of the present invention;
[0021] Figure 3 A schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention;
[0022] Figure 4 Schematic diagram of an improved frequency channel attention module in Embodiment 1 of the present invention;
[0023] Figure 5 Schematic diagram of EIoU parameters in Embodiment 1 of the present invention. DETAILED DESCRIPTION
[0024] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0025] Embodiment 1
[0026] Figure 1 A flowchart of a method for detecting a target in a scale-unbalanced remote sensing image provided in Embodiment 1 of the present invention is provided. This embodiment is applicable to the case of extracting a target in a remote sensing image, especially a remote sensing image with a large target scale difference. The method can be performed by a device for detecting a target in a scale-unbalanced remote sensing image. The device can be implemented by hardware and / or software and can generally be integrated in an electronic device, such as a computer device. The method specifically includes:
[0027] Step 110, extracting the target from the remote sensing image frame by introducing a feature extraction network of a frequency channel attention network;
[0028] Among them, in the feature extraction stage, the frequency channel attention network FcaNet can be introduced. In addition, in the embodiment of the present application, the frequency channel attention network is improved and then integrated into the feature extraction network to fully obtain rich features in each channel dimension, such as: feature information in different directions and frequencies; to improve the semantic information extraction ability of the neural network and improve the model's detection ability for small-scale targets.
[0029] Step 120, eliminating redundant bounding boxes by using a soft non-maximum suppression Soft-NMS algorithm;
[0030] Among them, Soft-NMS is introduced in the reasoning stage and EIoU is used instead of IoU in NMS, that is, the Soft-EIoU-NMS method is used in the post-reasoning processing stage to improve the recall rate of objects of different scales in overlapping borders and reduce the false detection rate.
[0031] Step 130: Use EIoU loss as the regression loss function to determine the position of the predicted bounding box
[0032] Among them, the position of the predicted bounding box is continuously refined by defining the bounding box regression loss function, calculating the regression loss value Loss, and then performing reverse gradient descent. The EIoU loss is introduced as the regression loss function to alleviate the accuracy impact of scale-imbalanced target instances on the loss function optimization.
[0033] The technical solution of this embodiment solves the problems of low multi-scale target detection accuracy, excessive computational complexity, and poor small target detection effect in the prior art by improving the feature extraction network, loss function, and inference post-processing links of the basic model, thereby improving the model's ability to extract multi-scale target features and alleviating the scale imbalance problem.
[0034] Optionally, before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network, the method further includes:
[0035] Use the CSPDarknet53 network module as the backbone network module;
[0036] An improved frequency channel attention module is added between the SPPF module and the last layer C3 module to build a feature extraction network.
[0037] Among them, the improved frequency channel attention module Att (such as Figure 4 As shown in the figure, the weighted channel and spatial feature information are input into the feature fusion network. Since small targets have higher energy in the frequency channel, they will be given a larger weight to increase the network's attention to small-scale target instances, improve the feature extraction ability of the neural network, and suppress interfering background information. The Att module is based on the FcaNet idea. First, the input X∈R C×H×W Segmentation is performed on channel C, where H and W represent height and width respectively, and are divided into multiple small blocks [X 0 ,X 1 ,X 2 ,……,X n-1 ], each X i ∈R C’×H×W , i belongs to {0,1,2,…,n-1}; where C' is the number of channels of the segmented small block Xi, C'=C / n. Then the segmented features are processed by linear layer for dimensionality reduction to obtain the feature map after 16 times of downsampling; then the Leaky ReLU activation function is used to increase the nonlinear expression ability, and then the feature map after the activation function is processed by linear layer for dimensionality increase to obtain the feature map after 16 times of upsampling, so that the number of channels is the same as the original input channel of the same layer; finally, the attention weight value between [0,1] is obtained by Sigmoid function, and the adaptive weight factor is obtained and then weighted combined with the original feature map of the same layer.
[0038] Optionally, before eliminating redundant bounding boxes by using the soft non-maximum suppression Soft-NMS algorithm, the method further includes:
[0039] Soft-NMS is introduced into the network model to eliminate redundant bounding boxes;
[0040] EIoU is used to replace IoU in NMS, where the EIoU parameter diagram is as follows Figure 5 As shown, EIoU is expressed as
[0041]
[0042] Among them, c w and c h b and b represent the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively. gt Denote the center point coordinates of the predicted bounding box and the true labeled bounding box, w and w respectively. gt and h and h gt Represent the width and height of the predicted bounding box and the true labeled bounding box, ρ represents the Euclidean distance between the center points of the predicted bounding box and the true labeled bounding box, and b gt =(x gt ,y gt ) represents the coordinates of the center point of the real labeled bounding box; b = (x1, y1) represents the coordinates of the center point of the predicted bounding box, and the gt superscript is used to represent the real annotation (ground truth).
[0043] Optionally, the using EIoU loss as a regression loss function to determine the position of the predicted bounding box includes:
[0044] When using EIOU loss as the positioning regression loss function, it is defined as follows:
[0045]
[0046] Optionally, the using EIoU loss as a regression loss function to determine the position of the predicted bounding box includes:
[0047] When the EIoU between the predicted bounding boxes is greater than the threshold, a confidence score is given, where the larger the EIoU, the lower the confidence score given. The confidence score is obtained by the following formula:
[0048]
[0049] Among them, N t is the set EIoU threshold, b max is the predicted bounding box with the highest confidence score in a class, b i is the i-th predicted bounding box, s i is the confidence score of the i-th predicted bounding box, S i The above formula (i.e. the capital S on the left side of the equal sign) i This calculation formula) calculates the confidence score of the i-th predicted bounding box, σ is a parameter, for example, σ = 0.5 in the experiment.
[0050] The Soft-EIoU-NMS algorithm used in the NMS method directly removes the predicted bounding boxes whose IoU is greater than the threshold, which will cause the problem of missing overlapping and dense target instances; when the EIoU of two predicted bounding boxes is greater than the threshold, it is not directly removed, but a relatively low confidence score is given (the removal described in the previous article is to set the confidence value to 0), and follows the principle that the larger the EIoU, the lower the confidence score.
[0051] The technical solution of the embodiment of the present application can not only improve the accuracy and efficiency of target detection, but also be applicable to multiple fields such as remote sensing, unmanned driving, security monitoring, etc., and has broad application prospects and market value. The frequency channel attention network is introduced and improved to fully obtain feature information in different directions and frequencies; the semantic information extraction ability of the neural network is improved, and the model's detection ability for small-scale targets is improved. The Soft-EIoU-NMS method is used in the post-inference processing stage to improve the recall rate of targets of different scales in overlapping borders and reduce the false detection rate. The EIoU loss is introduced as a regression loss function to alleviate the accuracy impact of scale-unbalanced target instances on the optimization of the loss function.
[0052] Embodiment 2
[0053] Figure 2 A schematic diagram of a device for detecting a target in a scale-unbalanced remote sensing image is provided in the second embodiment of the present invention. Figure 2 As shown, the scale-unbalanced remote sensing image target detection device includes: a target extraction module 210, a bounding box elimination module 220 and a bounding box position prediction module 230, wherein:
[0054] A target extraction module 210 is used to extract targets from remote sensing image frames by introducing a feature extraction network of a frequency channel attention network;
[0055] A bounding box elimination module 220 is used to eliminate redundant bounding boxes by using a soft non-maximum suppression Soft-NMS algorithm, wherein EIoU is used instead of IoU in NMS;
[0056] The predicted bounding box position module 230 is used to determine the position of the predicted bounding box by using the EIoU loss as a regression loss function.
[0057] Optionally, the scale-unbalanced remote sensing image target detection device further includes:
[0058] A backbone network module determination module, used to use the CSPDarknet53 network module as a backbone network module before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network;
[0059] The network improvement module is used to add an improved frequency channel attention module between the SPPF module and the last layer C3 module to build a feature extraction network.
[0060] Optionally, the scale-unbalanced remote sensing image target detection device further includes:
[0061] A Soft-EIoU-NMS algorithm building module is used to introduce Soft-NMS into the network model to eliminate redundant bounding boxes before eliminating redundant bounding boxes by the soft non-maximum suppression Soft-NMS algorithm;
[0062] EIoU is used to replace IoU in NMS, where EIoU is expressed as
[0063]
[0064] Among them, c w and c h b and b represent the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively. gt Denote the center point coordinates of the predicted bounding box and the true labeled bounding box, w and w respectively. gt and h and h gt Represent the width and height of the predicted bounding box and the true labeled bounding box respectively.
[0065] Optional, predicted bounding box location module, used to:
[0066] When using EIOU loss as the positioning regression loss function, it is defined as follows:
[0067]
[0068] Optional, predicted bounding box location module, used to:
[0069] When the EIoU between the predicted bounding boxes is greater than the threshold, a confidence score is given, where the larger the EIoU, the lower the confidence score given. The confidence score is obtained by the following formula:
[0070]
[0071] Among them, N t is the set EIoU threshold, b max is the predicted bounding box with the highest confidence score in a class, b i is the i-th predicted bounding box, s i is the confidence score of the i-th predicted bounding box.
[0072] The scale-unbalanced remote sensing image target detection device provided in the embodiment of the present invention can execute the scale-unbalanced remote sensing image target detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0073] Embodiment 3
[0074] Figure 3 A schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention is shown in FIG. Figure 3 As shown, the electronic device includes a processor 310, a memory 320, an input device 330 and an output device 340; the number of the processor 310 in the electronic device can be one or more. Figure 3 A processor 310 is taken as an example; the processor 310, the memory 320, the input device 330 and the output device 340 in the electronic device can be connected via a bus or other means. Figure 3 The example of connecting through bus is taken in the following.
[0075] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the scale-unbalanced remote sensing image target detection method in the embodiment of the present invention (for example, the target extraction module 210, the bounding box elimination module 220 and the predicted bounding box position module 230 in the scale-unbalanced remote sensing image target detection device). The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions and modules stored in the memory 320, that is, realizing the above-mentioned scale-unbalanced remote sensing image target detection method.
[0076] The memory 320 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 320 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include a memory remotely arranged relative to the processor 310, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0077] The input device 330 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. The output device 340 may include a display device such as a display screen.
[0078] Embodiment 4
[0079] Embodiment 4 of the present invention further provides a storage medium containing computer executable instructions, wherein the computer executable instructions are used to execute a method for detecting a scale-imbalanced remote sensing image target when executed by a computer processor, including:
[0080] Extracting objects from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network;
[0081] Redundant bounding boxes are eliminated through the soft non-maximum suppression Soft-NMS algorithm, where EIoU is used instead of IoU in NMS;
[0082] The EIoU loss is used as the regression loss function to determine the location of the predicted bounding box.
[0083] Of course, the storage medium containing computer executable instructions provided in an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in the scale-unbalanced remote sensing image target detection method provided in any embodiment of the present invention.
[0084] Through the above description of the implementation methods, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0085] It is worth noting that in the above-mentioned embodiment of the scale-unbalanced remote sensing image target detection device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0086] Although the present invention has been described in detail above by means of general description, specific implementation methods and tests, it is obvious to those skilled in the art that some modifications or improvements may be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection claimed by the present invention.
Claims
1. A method for detecting targets in scale-unbalanced remote sensing images, characterized in that: include: Extracting objects from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network; Redundant bounding boxes are eliminated through the soft non-maximum suppression Soft-NMS algorithm, where EIoU is used instead of IoU in NMS; The EIoU loss is used as the regression loss function to determine the location of the predicted bounding box.
2. The method for detecting targets in scale-unbalanced remote sensing images according to claim 1, characterized in that: Before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network, it also includes: Use the CSPDarknet53 network module as the backbone network module; An improved frequency channel attention module is added between the SPPF module and the last layer C3 module to build a feature extraction network.
3. The method according to claim 1, characterized in that: Before eliminating redundant bounding boxes by the soft non-maximum suppression Soft-NMS algorithm, the method further includes: Soft-NMS is introduced into the network model to eliminate redundant bounding boxes; EIoU is used to replace IoU in NMS, where EIoU is expressed as Among them, c w and c h b and b represent the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively. gt Denote the center point coordinates of the predicted bounding box and the true labeled bounding box, w and w respectively. gt and h and h gt They represent the width and height of the predicted bounding box and the true labeled bounding box respectively, and ρ represents the Euclidean distance between the center points of the predicted bounding box and the true labeled bounding box.
4. The method according to claim 3, characterized in that The method of using EIoU loss as a regression loss function to determine the position of the predicted bounding box includes: When using EIOU loss as the positioning regression loss function, it is defined as follows:
5. The method according to claim 4, characterized in that The method of using EIoU loss as a regression loss function to determine the position of the predicted bounding box includes: When the EIoU between the predicted bounding boxes is greater than the threshold, a confidence score is given, where the larger the EIoU, the lower the confidence score given. The confidence score is obtained by the following formula: Among them, N t is the set EIoU threshold, b max is the predicted bounding box with the highest confidence score in a class, b i is the i-th predicted bounding box, s i is the confidence score of the i-th predicted bounding box.
6. A scale-unbalanced remote sensing image target detection device, characterized in that: include: The target extraction module is used to extract targets from remote sensing image frames by introducing a feature extraction network with a frequency channel attention network; The bounding box elimination module is used to eliminate redundant bounding boxes through the soft non-maximum suppression Soft-NMS algorithm, in which EIoU is used instead of IoU in NMS; The predicted bounding box position module is used to determine the position of the predicted bounding box using EIoU loss as the regression loss function.
7. The device according to claim 6, characterized in that Also includes: A backbone network module determination module, used to use the CSPDarknet53 network module as a backbone network module before extracting the target from the remote sensing image frame by introducing the feature extraction network of the frequency channel attention network; The network improvement module is used to add an improved frequency channel attention module between the SPPF module and the last layer C3 module to build a feature extraction network.
8. The device according to claim 6, characterized in that Also includes: A Soft-EIoU-NMS algorithm building module is used to introduce Soft-NMS into the network model to eliminate redundant bounding boxes before eliminating redundant bounding boxes by the soft non-maximum suppression Soft-NMS algorithm; EIoU is used to replace IoU in NMS, where EIoU is expressed as Among them, c w and c h b and b represent the width and height of the minimum bounding rectangle of the predicted bounding box and the true annotation bounding box, respectively. gt Denote the center point coordinates of the predicted bounding box and the true labeled bounding box, w and w respectively. gt and h and h gt They represent the width and height of the predicted bounding box and the true labeled bounding box respectively, and ρ represents the Euclidean distance between the center points of the predicted bounding box and the true labeled bounding box.
9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the scale-unbalanced remote sensing image target detection method as described in any one of claims 1-5.
10. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the scale-unbalanced remote sensing image target detection method as described in any one of claims 1 to 5 when executed by a computer processor.
Citation Information
Patent Citations
Remote sensing image target detection method
CN116258953A
Small target detection method based on improved YOLOv5
CN117710965A
Multi-scale feature optimized remote sensing image segmentation model and method
CN118154868A
Remote sensing image target detection method based on multi-scale feature extraction
CN118230180A
System and Method for Data Forwarding
US20130114506A1