Remote sensing image directed target detection method

By using a non-angle periodic loss function based on aspect ratio excitation strategy in the directed object detection of remote sensing images, the problems of angular periodicity and aspect ratio are solved, and the detection performance and stability are significantly improved.

CN120125997APending Publication Date: 2025-06-10付亦凡
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510183862.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has problems with angular periodicity and excessive aspect ratio in directed object detection of remote sensing images, resulting in poor detection performance.

Method used

A non-angle periodic loss function based on aspect ratio excitation strategy is used to train a directed detection model of remote sensing images, solve the angular periodic problem and deal with the goal of excessive aspect ratio.

Benefits of technology

It effectively improves the detection performance of the model for directed targets of remote sensing images, especially when processing large aspect ratio targets, improves the accuracy of bounding box positioning and enhances the stability and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125997A_ABST
    Figure CN120125997A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image directed target detection method, which belongs to the technical field of remote sensing image processing, and comprises the following steps: acquiring to-be-detected remote sensing image data, and obtaining a prediction result based on a remote sensing image directed prediction model; the training method of the remote sensing image directed detection model comprises the steps of obtaining and labeling a large number of labeled remote sensing images as a training set, extracting multi-scale features of the remote sensing images, obtaining feature maps of different scales for fusion, constructing a feature pyramid map with multi-scale information, inputting the feature pyramid map into a detection head, and performing classification, centrality, bounding box and angle prediction. And according to the prediction result and the labeling information, calculating classification loss, centrality loss, bounding box regression loss and angle parameter regression loss to obtain a loss function, and calculating angle parameter regression loss by using a non-angle periodic loss function based on a length-width ratio excitation strategy. According to the invention, the problems of angle periodicity and overlarge length-width ratio during detection of directed targets in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting oriented targets in remote sensing images, belonging to the technical field of remote sensing image processing. Background Art

[0002] As an important application direction of image processing, the purpose of detecting rotated targets in remote sensing images is to find the specific positions of all targets to be detected and predict the corresponding class information for each target instance. As an efficient analysis method for image interpretation tasks, the results of remote sensing image target detection can clearly reflect the spatial distribution of ground objects, facilitating people to recognize and discover laws from them. Compared with traditional natural image target detection, remote sensing image target detection has the following difficulties: the targets to be detected are densely distributed and have various angular orientations; the background is redundant and complex, and is similar in shape to the target instances; the size of remote sensing images is huge, and the scale of the targets to be detected is extremely small; the image background is complex and the detection difficulty is high; the aspect ratios of the targets to be detected vary greatly. Therefore, the detection models and methods for conventional natural images cannot meet the needs of detecting rotated targets in remote sensing images. It is necessary to further develop remote sensing image target detection technology based on technologies such as pattern recognition, artificial intelligence, and image processing according to the distribution characteristics of target ground objects in remote sensing images, so as to effectively detect and identify the ground object categories that are difficult to detect in the images.

[0003] In the past decade, deep learning has achieved excellent performance in various computer vision tasks represented by target detection due to its powerful feature modeling ability. Compared with traditional methods, the target detection algorithm based on a deep convolutional neural network directly calibrates the position and class results of the target in the form of parameter iterative update learning, effectively solving the problem of poor generalization of manually designed features. However, due to the influence of its internal working mode, the convolution operation is difficult to model the rotation orientation information of the targets to be detected, resulting in the fact that the rotated target detection task in the remote sensing scenario has always been difficult to reach a high level. Some scholars have designed various extraction modules for rotation-invariant features and prediction modules for the rotation angles of targets to enhance the detection network's perception ability of the target rotation angles on the basis of the convolutional neural network, thereby improving the detection performance of rotated targets in remote sensing images.

[0004] However, the above methods still have problems: the angle prediction module regards the angle data as linear data and does not consider the angle periodicity problem; if the aspect ratio of the target is too large, a small angle offset will cause a large IoU loss, and it is necessary to redesign the loss function to make the network focus on the targets with too large aspect ratios. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for detecting oriented targets in remote sensing images, by means of a non-angle periodic loss function based on an aspect ratio excitation strategy, to solve the angle periodicity problem and the problem of too large aspect ratio that occur in the prior art when detecting oriented targets.

[0006] To solve the above technical problems, the present invention is implemented by the following technical solutions.

[0007] The present invention provides a method for detecting oriented targets in remote sensing images, including:

[0008] Obtaining remote sensing image data to be detected;

[0009] Based on the remote sensing image data to be detected and a pre-trained remote sensing image oriented prediction model, making a prediction to obtain a prediction result;

[0010] The prediction result includes the class probability, centrality, bounding box coordinates, and angle value of the target;

[0011] Among them, the training method of the remote sensing image oriented detection model includes:

[0012] Obtaining and annotating a large number of annotated remote sensing images as a training set, and the annotation information includes the class probability, centrality, bounding box coordinates, and angle value of the target;

[0013] Feeding the training set into a deep learning network, and using a convolutional neural network to extract multi-scale features of the remote sensing image to obtain feature maps of different scales;

[0014] Using a feature pyramid network to fuse the feature maps of different scales to construct a feature pyramid map with multi-scale information;

[0015] Inputting the feature pyramid map into a detection head to perform classification, centrality, bounding box, and angle prediction to obtain a prediction result;

[0016] Among them, according to the prediction result and the annotation information, calculating the classification loss, centrality loss, bounding box regression loss, and angle parameter regression loss to obtain a loss function;

[0017] Among them, using a non-angle periodic loss function based on an aspect ratio excitation strategy to calculate the angle parameter regression loss.

[0018] Further, feeding the training set into a deep learning network, and using a convolutional neural network to extract multi-scale features of the remote sensing image to obtain feature maps of different scales, where the feature map is expressed as:

[0019] ;

[0020] In the formula, represents the th feature map, where represents the real number space, represents the width of the remote sensing image, represents the The number of channels of a feature map indicates that the remote sensing images in the training set are processed through a residual network indicates the th feature map passes through the th residual block structure in the residual network indicates a pooling operation indicates a two-dimensional convolution operation

[0021] Further, the feature maps include a first feature map, a second feature map, a third feature map, and a fourth feature map. Among them, the number of channels of the first feature map, the second feature map, the third feature map, and the fourth feature map decreases in sequence. The feature pyramid network is used to fuse feature maps of different scales to construct a feature pyramid with multi-scale information, including:

[0022] Use 256 convolutional kernels of size 1 to perform convolution on the second feature map, the third feature map, and the fourth feature map to obtain a fifth feature map, a sixth feature map, and a seventh feature map

[0023] After performing an upsampling operation on the seventh feature map, add it element-wise to the sixth feature map to obtain a new seventh feature map

[0024] After performing an upsampling operation on the sixth feature map, add it element-wise to the fifth feature map to obtain a new sixth feature map

[0025] After performing an upsampling operation on the fifth feature map, add it element-wise to the fourth feature map to obtain a new fifth feature map

[0026] Perform a two-dimensional convolution operation on the new fifth feature map, the new sixth feature map, and the seventh feature map to eliminate the aliasing effect of upsampling, and obtain an eighth feature map, a ninth feature map, and a tenth feature map

[0027] Repeat performing a two-dimensional convolution operation on the new seventh feature map and then introducing a ReLU activation function for processing until the number of channels is the same as that of the new sixth feature map to obtain an eleventh feature map

[0028] According to the eighth feature map, the ninth feature map, the tenth feature map, and the eleventh feature map, perform construction to obtain a feature pyramid map with multi-scale information

[0029] Further, input the feature pyramid map into a detection head to perform classification, centrality, bounding box, and angle prediction to obtain prediction results, including:

[0030] Perform four consecutive two-dimensional convolution operations on the feature pyramid diagram, then perform group normalization operation, and then introduce the ReLU activation function for processing to obtain the processed feature pyramid diagram;

[0031] Input the processed feature pyramid diagram into the classification branch, regression branch, and centerness branch of the prediction head in parallel for prediction to obtain the prediction result.

[0032] Further, the classification branch, the regression branch, and the centerness branch are all connected to a group of shared convolutional layers. Among them, the classification branch takes the result of the shared convolutional layer as input, uses a convolutional layer with the number of convolutional kernels matching the number of remote sensing image categories for prediction, and after being processed by the convolutional layer, it is activated through the Sigmoid function to output the class probability of the target;

[0033] The centerness branch takes the result of the shared convolutional layer as input, uses a convolutional layer with the number of convolutional kernels being 1 for prediction, and after being processed by the convolutional layer, it is activated through the Sigmoid function to output the centerness of the target;

[0034] The regression branch takes the result of the shared convolutional layer as input, uses a convolutional layer with the number of convolutional kernels being 5 for prediction, and after being processed by the convolutional layer, it is activated using the ReLU function to output the bounding box coordinates and angle values of the target.

[0035] Further, the loss function is expressed as:

[0036] ;

[0037] In the formula, represents the loss value obtained by calculating the loss function, represents the classification loss, represents the hyperparameter used to balance the centerness loss of, represents the hyperparameter used to balance the bounding box regression loss of, represents the hyperparameter used to balance the angle parameter regression loss of.

[0038] Further, the focal loss function is used to calculate the classification loss, and the classification loss is used to measure the difference between the predicted value of the class probability of the target and the true class label of the target, which is expressed as:

[0039] ;

[0040] In the formula, represents the classification loss, represents the total number of samples in the training set, represents the set of positive samples, Indicates the probability that the th target is a positive class, represents the logarithmic function, indicates the th true class label of the target.

[0041] Furthermore, the IoU loss function is used to calculate the bounding box regression loss, which is used to measure the difference between the predicted value and the true value of the target's bounding box, expressed as:

[0042] ;

[0043] In the formula, represents the bounding box regression loss, represents the total number of samples in the training set, indicates the th intersection over union of the predicted value and the true value of the target's bounding box, represents the set of positive samples, represents the logarithmic function.

[0044] Furthermore, the binary cross-entropy loss function is used to calculate the centrality loss, which is used to measure the difference between the predicted value and the true value of the target's centrality, expressed as:

[0045] ;

[0046] In the formula, represents the centrality loss, represents the total number of samples in the training set, represents the set of positive samples, indicates the th predicted value of the target's centrality, indicates the th true value of the target's centrality, represents the logarithmic function.

[0047] Furthermore, the angular parameter regression loss is calculated using the non-angle periodic loss function based on the aspect ratio excitation strategy, which is used to measure the difference between the predicted value and the true value of the target's angle, expressed as:

[0048] ;

[0049] In the formula, represents the angular parameter regression loss, represents the penalty factor, represents the non-angle periodic loss function;

[0050] Among them, the non-angle periodic loss function is expressed as:

[0051] ;

[0052] In the formula, represents the non-angle periodic loss function, represents a parameter used to control the magnitude of the function gradient, which is used to affect the convergence speed of the model, represents the sine function, represents the predicted angle value, represents the true angle value, represents except for other cases.

[0053] Among them, the penalty factor is expressed as:

[0054] ;

[0055] In the formula, represents the exponential function, and are hyperparameters, represents a hyperparameter used to control the magnitude of the penalty factor, represents a hyperparameter used to control the aspect ratio threshold of the bounding box, and AR represents the aspect ratio of the bounding box;

[0056] Among them, the calculation formula of the aspect ratio AR of the bounding box is expressed as:

[0057] ;

[0058] In the formula, represents the value of the aspect ratio AR of the bounding box, represents taking the maximum value, represents taking the minimum value, , respectively represent the abscissa and ordinate of the upper left corner of the bounding box, , respectively represent the abscissa and ordinate of the lower right corner of the bounding box.

[0059] Compared with the prior art, the beneficial effects achieved by the present invention:

[0060] By adopting a non-angle periodic loss function based on the aspect ratio excitation strategy, the present invention not only effectively solves the angle periodicity problem in the detection of oriented targets in remote sensing images, avoids jumps and misguidance during angle prediction, but also significantly improves the accuracy of the model in the positioning of the target bounding box for targets with too large aspect ratios. Moreover, it does not require the introduction of a complex encoding and decoding structure, is simple and efficient to implement, greatly improves the detection effect on remote sensing images and targets with large aspect ratios, and enhances the stability and reliability of the detection. Brief Description of the Drawings

[0061] Figure 1 It is a schematic flowchart of a method for detecting oriented targets in remote sensing images provided by an embodiment of the present invention. Detailed Embodiment

[0062] The technical solution of the present invention will be described in detail below through the drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0063] The term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.

[0064] Embodiment 1

[0065] As Figure 1 shown, this embodiment introduces a method for detecting oriented targets in remote sensing images, including:

[0066] Step 1: Obtain remote sensing image data to be detected.

[0067] Step 2: Based on the remote sensing image data to be detected and a pre-trained remote sensing image oriented prediction model, perform prediction to obtain a prediction result.

[0068] This embodiment uses a trained remote sensing image oriented prediction model to process the remote sensing image to be detected. The remote sensing image oriented prediction model can automatically extract image features and perform classification, centrality, bounding box, and angle prediction operations.

[0069] The prediction result includes the class probability, centrality, bounding box coordinates, and angle value of the target.

[0070] In this embodiment, the class probability of the target is used to determine which class the target belongs to; the centrality is used to evaluate the confidence that the target is located at the center of the bounding box; the bounding box coordinates define the position of the target in the image; and the angle value represents the orientation of the target.

[0071] Among them, the training method of the remote sensing image oriented detection model includes:

[0072] Obtain and label a large number of labeled remote sensing images as a training set, and the labeling information includes the class probability, centrality, bounding box coordinates, and angle value of the target.

[0073] Remote sensing images contain rich surface information, such as buildings, vehicles, ships, etc. By annotating the category to which the target belongs, the bounding box coordinates, and the angle value, the model can learn the features of these targets and accurately identify and locate them in new images.

[0074] In this embodiment, in addition to the basic category annotation, the centrality, bounding box coordinates, and angle value of the target are also annotated, which helps the model to more precisely understand the position and pose of the target in the image, thereby improving the accuracy and robustness of detection.

[0075] The training set is fed into a deep learning network, and a convolutional neural network is used to extract multi-scale features of the remote sensing image to obtain feature maps of different scales. The convolutional neural network can automatically extract the features in the image, and the multi-scale feature maps help to capture targets of different sizes, enhance the detection ability of the remote sensing image directed prediction model for targets of different sizes, and improve the detection accuracy.

[0076] A feature pyramid network is used to fuse the feature maps of different scales to construct a feature pyramid map with multi-scale information. The feature pyramid network can fuse the feature maps of different scales to generate a feature pyramid map with rich information, improve the detection performance of the remote sensing image directed prediction model for targets of different scales, and enhance the robustness of the model.

[0077] The feature pyramid map is input into the detection head for classification, centrality, bounding box, and angle prediction to obtain the prediction results. The detection head uses the feature pyramid map for prediction and outputs information such as the category, centrality, bounding box, and angle of the target, realizing the accurate detection and positioning of the target, and providing strong support for practical applications.

[0078] Among them, according to the prediction results and the annotation information, the classification loss, centrality loss, bounding box regression loss, and angle parameter regression loss are calculated to obtain the loss function. The loss function is used to measure the difference between the model prediction results and the true annotation information, guide the training process of the model, ensure that the model can be continuously optimized during the training process, and improve the prediction accuracy.

[0079] Among them, the angle parameter regression loss is calculated using a non-angle periodic loss function based on the aspect ratio excitation strategy. The non-angle periodic loss function solves the periodic problem in angle prediction. At the same time, the aspect ratio excitation strategy helps to handle targets with too large aspect ratios, improve the accuracy of angle prediction, enhance the detection ability of the model for long and narrow targets, and improve the stability of angle prediction.

[0080] Embodiment 2

[0081] Based on the same inventive concept as in Embodiment 1, this embodiment introduces the implementation steps of a method for detecting oriented targets in remote sensing images, including:

[0082] Step 1: Obtain the remote sensing image data to be detected.

[0083] Step 2: Based on the obtained remote sensing image data and a pre-trained remote sensing image oriented prediction model, perform prediction to obtain a prediction result;

[0084] The prediction result includes the class probability, centrality, bounding box coordinates, and angle value of the target;

[0085] Step 2.1: Train the remote sensing image oriented detection model.

[0086] Among them, the training method of the remote sensing image oriented detection model includes:

[0087] Step 2.1.1: Obtain and label a large number of labeled remote sensing images as the training set, and the labeling information includes the class probability, centrality, bounding box coordinates, and angle value of the target;

[0088] Step 2.1.2: Feed the training set into a deep learning network, and use a convolutional neural network to extract multi-scale features of the remote sensing image to obtain feature maps of different scales.

[0089] In this embodiment, the feature map is expressed as:

[0090] ;

[0091] In the formula, represents the th feature map, where represents the real number space, represents the width of the remote sensing image, represents the th channel number of the th feature map, represents that the remote sensing image in the training set is processed through a residual network, represents the th residual block structure in the residual network through which the th feature map passes, represents a pooling operation,

[0092] Step 2.1.3: Use a feature pyramid network to fuse the feature maps of different scales to construct a feature pyramid map with multi-scale information.

[0093] In some embodiments, the feature maps include a first feature map, a second feature map, a third feature map, and a fourth feature map, wherein the number of channels of the first feature map, the second feature map, the third feature map, and the fourth feature map decreases in sequence.

[0094] In this embodiment, a feature pyramid network is used to fuse feature maps of different scales to construct a feature pyramid with multi-scale information, including:

[0095] Convolve the second feature map, the third feature map, and the fourth feature map with 256 convolutional kernels of size 1 to obtain a fifth feature map, a sixth feature map, and a seventh feature map;

[0096] After performing an upsampling operation on the seventh feature map, add it element-wise to the sixth feature map to obtain a new seventh feature map;

[0097] After performing an upsampling operation on the sixth feature map, add it element-wise to the fifth feature map to obtain a new sixth feature map;

[0098] After performing an upsampling operation on the fifth feature map, add it element-wise to the fourth feature map to obtain a new fifth feature map;

[0099] Perform a two-dimensional convolution operation on the new fifth feature map, the new sixth feature map, and the seventh feature map to eliminate the aliasing effect of upsampling, and obtain an eighth feature map, a ninth feature map, and a tenth feature map;

[0100] Repeat performing a two-dimensional convolution operation on the new seventh feature map and then introducing a ReLU activation function for processing until the number of channels is the same as that of the new sixth feature map to obtain an eleventh feature map;

[0101] Based on the eighth feature map, the ninth feature map, the tenth feature map, and the eleventh feature map, construct a feature pyramid map with multi-scale information. The feature pyramid map has both high-level semantic information and low-level detail information.

[0102] Step 2.1.4: Input the feature pyramid map into a detection head to perform classification, centerness, bounding box, and angle prediction to obtain a prediction result.

[0103] In some embodiments, inputting the feature pyramid map into a detection head to perform classification, centerness, bounding box, and angle prediction to obtain a prediction result includes:

[0104] Perform four consecutive two-dimensional convolution operations on the feature pyramid map, then perform a group normalization operation, and then introduce a ReLU activation function for processing to obtain a processed feature pyramid map;

[0105] The processed feature pyramid map is input into the classification branch, regression branch, and centerness branch of the prediction head in parallel for prediction to obtain the prediction results.

[0106] In some embodiments, the classification branch, the regression branch, and the centerness branch are all connected to a group of shared convolutional layers. Among them, the classification branch takes the result of the shared convolutional layer as input, uses a convolutional layer with the number of convolutional kernels matching the number of remote sensing image categories for prediction. After being processed by the convolutional layer, it is activated by the Sigmoid function to output the class probability of the target.

[0107] The centerness branch takes the result of the shared convolutional layer as input, uses a convolutional layer with the number of convolutional kernels being 1 for prediction. After being processed by the convolutional layer, it is activated by the Sigmoid function to output the centerness of the target.

[0108] The regression branch takes the result of the shared convolutional layer as input, uses a convolutional layer with the number of convolutional kernels being 5 for prediction. After being processed by the convolutional layer, it is activated by the ReLU function to output the bounding box coordinates and angle values of the target.

[0109] Among them, according to the prediction results and the annotation information, the classification loss, centerness loss, bounding box regression loss, and angle parameter regression loss are calculated to obtain the loss function.

[0110] In this embodiment, the loss function is expressed as:

[0111] ;

[0112] In the formula, represents the loss value obtained by calculating the loss function, represents the classification loss, represents the hyperparameter used to balance the centerness loss of, represents the hyperparameter used to balance the bounding box regression loss of, represents the hyperparameter used to balance the angle parameter regression loss of.

[0113] In this embodiment, the focal loss function is used to calculate the classification loss. The classification loss is used to measure the difference between the predicted class probability value of the target and the true class label of the target, and is expressed as:

[0114] ;

[0115] In the formula, represents the classification loss, represents the total number of samples in the training set, represents the set of positive samples, Indicates the probability that the th target is a positive class, represents the logarithmic function, indicates the th true class label of the target.

[0116] In this embodiment, the IoU loss function is used to calculate the bounding box regression loss, which is used to measure the difference between the predicted value and the true value of the bounding box of the target, and is expressed as:

[0117] ;

[0118] In the formula, represents the bounding box regression loss, represents the total number of samples in the training set, indicates the th intersection over union of the predicted value and the true value of the bounding box of the target, represents the set of positive samples, represents the logarithmic function.

[0119] In this embodiment, the binary cross-entropy loss function is used to calculate the centrality loss, which is used to measure the difference between the predicted value and the true value of the centrality of the target, and is expressed as:

[0120] ;

[0121] In the formula, represents the centrality loss, represents the total number of samples in the training set, represents the set of positive samples, indicates the th predicted value of the centrality of the target, indicates the th true value of the centrality of the target, represents the logarithmic function.

[0122] In this embodiment, the non-angle periodic loss function based on the aspect ratio excitation strategy is used to calculate the angle parameter regression loss. Among them, the angle parameter regression loss is used to measure the difference between the predicted value and the true value of the angle of the target, and is expressed as:

[0123] ;

[0124] In the formula, represents the angle parameter regression loss, represents the penalty factor, represents the non-angle periodic loss function;

[0125] Among them, the non-angle periodic loss function is expressed as:

[0126] ;

[0127] In the formula, represents the non-angle periodic loss function, represents a parameter used to control the magnitude of the function gradient and is used to affect the convergence speed of the model, represents the sine function, represents the predicted angle value, represents the true angle value, represents except for other cases.

[0128] Among them, the penalty factor is expressed as:

[0129] ;

[0130] In the formula, represents the exponential function, and are hyperparameters, represents a hyperparameter used to control the magnitude of the penalty factor, represents a hyperparameter used to control the aspect ratio threshold of the bounding box, and AR represents the aspect ratio of the bounding box;

[0131] Among them, the calculation formula of the aspect ratio AR of the bounding box is expressed as:

[0132] ;

[0133] In the formula, represents the value of the aspect ratio AR of the bounding box, represents taking the maximum value, represents taking the minimum value, , respectively represent the abscissa and ordinate of the upper left corner of the bounding box, , respectively represent the abscissa and ordinate of the lower right corner of the bounding box.

[0134] Embodiment 3

[0135] Based on the same inventive concept as other embodiments, this embodiment introduces a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the steps of the method in Embodiment 1 or 2 above are implemented.

[0136] Embodiment 4

[0137] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method in Embodiment 1 or 2 above.

[0138] In summary of the above embodiments, the present invention is different from the mainstream deep learning-based object detection frameworks. It adopts an aspect ratio excitation strategy to make the model focus on the learning of large aspect ratio objects, improves the detection effect of the model on large aspect ratio objects, and enhances the bounding box localization ability of the model. By using a non-angle periodic loss function, regarding the angle data as circular data rather than linear data, it solves the periodic problem that occurs in the angle prediction process, and does not require the introduction of additional encoding and decoding structures, which is simple to use. Moreover, the aspect ratio excitation strategy can be well combined with the non-angle periodic loss function, improving the detection effect of the model on oriented objects in remote sensing images.

[0139] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0141] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps for realizing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 steps for realizing the functions specified in one block or multiple blocks.

[0143] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.

Claims

1. A method for detecting directed targets in remote sensing images, characterized in that: include: Acquire remote sensing image data to be detected; According to the remote sensing image data to be detected, based on a pre-trained remote sensing image directional prediction model, prediction is performed to obtain a prediction result; The prediction result includes the category probability, center degree, bounding box coordinates and angle value to which the target belongs; The training method of the remote sensing image directional detection model includes: Obtain and annotate a large number of remote sensing images as training sets. The annotation information includes the category probability, center, bounding box coordinates and angle value of the target. The training set is fed into a deep learning network, and a convolutional neural network is used to extract multi-scale features of the remote sensing image to obtain feature maps of different scales; The feature pyramid network is used to fuse feature maps of different scales to construct a feature pyramid map with multi-scale information; Input the feature pyramid image into the detection head to perform classification, center degree, bounding box and angle prediction to obtain a prediction result; According to the prediction results and the annotation information, the classification loss, the center loss, the bounding box regression loss and the angle parameter regression loss are calculated to obtain a loss function; Among them, the angle parameter regression loss is calculated using a non-angle periodic loss function based on an aspect ratio excitation strategy.

2. The method for detecting directed targets in remote sensing images according to claim 1, characterized in that: The training set is sent to a deep learning network, and a convolutional neural network is used to extract multi-scale features of the remote sensing image to obtain feature maps of different scales, wherein the feature map is represented as: ; In the formula, Indicates feature maps, where represents the real number space, Represents the width of the remote sensing image, Indicates The number of channels of the feature map, Represents the remote sensing images in the training set Processed by residual network, Indicates The feature map is passed through the residual network The residual block structure, represents the pooling operation, Represents a two-dimensional convolution operation.

3. The method for detecting directed targets in remote sensing images according to claim 2, characterized in that: The feature map includes a first feature map, a second feature map, a third feature map and a fourth feature map, wherein the number of channels of the first feature map, the second feature map, the third feature map and the fourth feature map decreases in sequence, and feature maps of different scales are fused using a feature pyramid network to construct a feature pyramid with multi-scale information, including: Convolve the second feature map, the third feature map, and the fourth feature map using 256 convolution kernels of size 1 to obtain the fifth feature map, the sixth feature map, and the seventh feature map; After upsampling the seventh feature map, the seventh feature map is added element by element to the sixth feature map to obtain a new seventh feature map; After upsampling the sixth feature map, the sixth feature map is added element by element to the fifth feature map to obtain a new sixth feature map; After performing an upsampling operation on the fifth feature map, the fifth feature map is added element by element to the fourth feature map to obtain a new fifth feature map; Perform a two-dimensional convolution operation on the new fifth feature map, the new sixth feature map, and the seventh feature map to eliminate the aliasing effect of upsampling, and obtain an eighth feature map, a ninth feature map, and a tenth feature map; Repeat the two-dimensional convolution operation on the new seventh feature map and then introduce the ReLU activation function for processing until the number of channels is the same as the number of channels of the new sixth feature map, thereby obtaining the eleventh feature map; According to the eighth feature map, the ninth feature map, the tenth feature map and the eleventh feature map, a feature pyramid map with multi-scale information is constructed.

4. The method for detecting directed targets in remote sensing images according to claim 3, characterized in that: The feature pyramid image is input into the detection head for classification, center degree, bounding box and angle prediction to obtain prediction results, including: The feature pyramid image is subjected to four consecutive two-dimensional convolution operations, and then a group normalization operation is performed, and then a ReLU activation function is introduced for processing to obtain a processed feature pyramid image; The processed feature pyramid image is input into the classification branch, regression branch, and centrality branch of the prediction head in parallel to perform prediction and obtain the prediction result.

5. The method for detecting directed targets in remote sensing images according to claim 4, characterized in that: The classification branch, the regression branch, and the center branch are all connected to a group of shared convolutional layers, wherein the classification branch takes the result of the shared convolutional layer as input, uses a convolutional layer whose number of convolutional kernels matches the number of remote sensing image categories for prediction, and after being processed by the convolutional layer, is activated by a Sigmoid function to output the probability of the category to which the target belongs; The center branch takes the result of the shared convolution layer as input, uses a convolution layer with 1 convolution kernel to make predictions, and after being processed by the convolution layer, is activated by a Sigmoid function to output the center of the target; The regression branch takes the result of the shared convolution layer as input, uses a convolution layer with 5 convolution kernels for prediction, and after being processed by the convolution layer, uses the ReLU function for activation to output the bounding box coordinates and angle values ​​of the target.

6. The method for detecting directed targets in remote sensing images according to claim 1, characterized in that: The loss function is expressed as: ; In the formula, Represents the loss value obtained by calculating the loss function, represents the classification loss, Indicates the balance of centrality loss The hyperparameters of Represents the loss used to balance the bounding box regression The hyperparameters of Represents the regression loss used to balance the angle parameters Hyperparameters of .

7. The method for detecting directed targets in remote sensing images according to claim 6, characterized in that: The classification loss is calculated using the focal loss function, which is used to measure the difference between the predicted value of the class probability to which the target belongs and the target's true class label, and is expressed as: ; In the formula, represents the classification loss, represents the total number of samples in the training set, represents the positive sample set, Indicates The probability that a target is a positive class, represents the logarithmic function, Indicates The true category labels of the targets.

8. The method for detecting directed targets in remote sensing images according to claim 6, characterized in that: The IoU loss function is used to calculate the bounding box regression loss, which is used to measure the difference between the bounding box prediction value and the true value of the bounding box of the target, and is expressed as: ; In the formula, represents the bounding box regression loss, represents the total number of samples in the training set, Indicates The intersection-over-union ratio of the bounding box prediction value and the true value of the bounding box of each target, represents the positive sample set, Represents a logarithmic function.

9. The method for detecting directed targets in remote sensing images according to claim 6, characterized in that: The binary cross entropy loss function is used to calculate the centrality loss, which is used to measure the difference between the predicted centrality value and the true centrality value of the target, and is expressed as: ; In the formula, represents the center loss, represents the total number of samples in the training set, represents the positive sample set, Indicates The predicted value of the centrality of the target, Indicates The true value of the centrality of the target, Represents a logarithmic function.

10. The method for detecting directed targets in remote sensing images according to claim 6, characterized in that: The angle parameter regression loss is calculated using a non-angle periodic loss function based on an aspect ratio excitation strategy. The angle parameter regression loss is used to measure the difference between the angle prediction value and the true angle value of the target, and is expressed as: ; In the formula, represents the angle parameter regression loss, represents the penalty factor, represents the non-angular periodic loss function; Wherein, the non-angular periodicity loss function is expressed as: ; In the formula, represents the non-angular periodic loss function, Represents the parameter used to control the size of the function gradient, which is used to affect the convergence speed of the model. represents the sine function, represents the angle prediction value, Represents the true value of the angle, Indicates except Other situations; Wherein, the penalty factor is expressed as: ; In the formula, represents the exponential function, and is a hyperparameter, represents a hyperparameter used to control the size of the penalty factor, represents the hyperparameter used to control the aspect ratio threshold of the bounding box, AR represents the aspect ratio of the bounding box; The calculation formula of the aspect ratio AR of the bounding box is expressed as: ; In the formula, The value representing the aspect ratio AR of the bounding box, Indicates taking the maximum value, Indicates taking the minimum value, , Respectively represent the horizontal and vertical coordinates of the upper left corner of the bounding box, , Respectively represent the horizontal and vertical coordinates of the lower right corner of the bounding box.

Citation Information

Cited By

  • Few-sample target detection method and device for remote sensing image

    CN122176535A