Method for detecting target in any direction of remote sensing image
By combining the improved PWC-Net network with a dual-branch feature extraction network, spatiotemporal feature maps are generated and redundant detection frames are eliminated, which solves the problem of detecting targets in arbitrary directions in remote sensing images and achieves high-precision and stable detection effects.
Patent Information
- Application Number
- CN202510753733.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing remote sensing image target detection methods have difficulty capturing targets in arbitrary directions and lack an effective temporal information fusion mechanism, resulting in insufficient detection performance when processing continuous frame data and prone to false detection and missed detection in complex backgrounds.
An improved PWC-Net network is used for optical flow field estimation. A dual-branch feature extraction network is combined to extract static semantic features and temporal correlation features to generate spatiotemporal feature maps. The weighted fusion is performed through a gated attention mechanism to construct a comprehensive similarity matrix, calculate the joint loss function, generate angle-sensitive basis vectors, perform convolution and Softmax activation, and finally eliminate redundant detection boxes through a soft-core non-maximum suppression algorithm.
It achieves accurate detection of targets in any direction, improves detection accuracy and robustness, and performs particularly well when dealing with tilted or rotating targets, enhancing detection capabilities in complex environments.
Smart Images

Figure CN120747765A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image processing and computer vision technology, in particular to a method for detecting targets in arbitrary directions using remote sensing images. Background Art
[0002] With the development of remote sensing technology, remote sensing image detection methods have been widely used in military reconnaissance, environmental monitoring, disaster assessment, and other fields. Traditional remote sensing image object detection methods rely primarily on hand-crafted feature extraction and classifiers, such as those based on low-level features such as edges and textures. In recent years, the rise of deep learning technology has brought new opportunities for remote sensing image object detection. Convolutional neural networks (CNNs), through end-to-end learning, can automatically extract rich feature representations and have achieved significant progress in various vision tasks.
[0003] However, existing technologies still have some shortcomings. For one thing, most current mainstream remote sensing image target detection methods rely on static image features, ignoring information in the temporal dimension. This makes it difficult to fully exploit the correlation between adjacent frames to improve detection performance when processing continuous frame data. Furthermore, while existing multi-scale feature fusion strategies can enhance the model's expressive power to a certain extent, they are still prone to false detections and missed detections when faced with complex background interference. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for detecting targets in arbitrary directions in remote sensing images to solve the problems of difficulty in capturing targets in arbitrary directions and lack of an effective temporal information fusion mechanism.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for detecting targets in arbitrary directions in remote sensing images, which comprises inputting a remote sensing image dataset into an improved PWC-Net network to perform optical flow field estimation and obtain an optical flow density map;
[0008] The optical flow density map is input into a dual-branch feature extraction network to extract static semantic features and temporal correlation features, and a gated attention mechanism is applied to weighted fusion to generate a spatiotemporal feature map.
[0009] Generate triplet samples based on the spatiotemporal feature map and input them into the dual encoder architecture to obtain the comprehensive similarity matrix by constructing the spatial similarity matrix and the temporal similarity matrix;
[0010] Calculate the self-supervision loss value and the detection loss value based on the comprehensive similarity matrix, and perform weighted summation to generate a joint loss function;
[0011] Based on the joint loss function, the angle-sensitive basis vector is generated through Fourier transform, and convolved with the spatiotemporal feature map to generate a feature response map. The feature response map is activated by Softmax and parameter decoding is performed to obtain a set of detection boxes.
[0012] The soft-core non-maximum suppression algorithm is used to eliminate redundant detection frames and output a target detection report that satisfies any direction.
[0013] As a preferred solution of the method for detecting targets in any direction in remote sensing images of the present invention, the steps for obtaining the optical flow density map are as follows:
[0014] Preprocess the input remote sensing image data for radiation correction and geometric registration;
[0015] Based on the multi-spectral flow and optical flow pyramid structure, an improved PWC-Net network is constructed by embedding an adaptive learning mechanism and a contrastive learning loss function.
[0016] The preprocessed remote sensing image data is input into the improved PWC-Net network, and multi-scale features are extracted layer by layer through the optical flow pyramid structure. The first layer uses convolution operations to generate preliminary optical flow predictions, and the second layer uses residual connections to fuse the optical flow predictions of the first layer with the features of the current layer to obtain optical flow field estimates.
[0017] The optical flow density map is generated by refining the optical flow field estimation layer by layer and optimizing it in combination with multispectral optical flow.
[0018] As a preferred solution of the method for detecting targets in any direction in remote sensing images of the present invention, the steps for obtaining the spatiotemporal feature map are as follows:
[0019] The optical flow density map is input into a dual-branch feature extraction network. The convolutional neural network branch extracts static semantic features of shape and texture, and the temporal correlation feature branch extracts temporal correlation features of motion trajectory and change trend.
[0020] The extracted static semantic features and temporal correlation features are applied with a gated attention mechanism to calculate the feature importance scores. The fully connected layers and activation functions are then connected to generate different weights, and the fused feature map is generated through weighted summation.
[0021] The fused feature map is refined to generate a spatiotemporal feature map.
[0022] As a preferred solution of the method for detecting targets in any direction in remote sensing images of the present invention, the steps for obtaining the triplet samples are as follows:
[0023] Perform random rotation transformation and frequency domain perturbation on the spatiotemporal feature map to generate a candidate sample set;
[0024] Select target region features from the candidate sample set, construct triple samples containing anchor points, positive samples, and negative samples, and optimize the triple samples through difficult sample mining method;
[0025] The target area features are defined based on the features of the position, shape, texture and motion trajectory of the target in the spatiotemporal feature map.
[0026] As a preferred solution of the method for detecting targets in any direction in remote sensing images of the present invention, the steps for obtaining the comprehensive similarity matrix are as follows:
[0027] The triplet sample is input into the dual encoder architecture. The spatial feature encoder extracts the spatial information of the sample through a convolutional neural network and forms a spatial similarity matrix by calculating the Euclidean distance between the anchor point and the spatial feature vectors of the positive sample and the negative sample.
[0028] The temporal feature encoder uses a recurrent neural network to extract the temporal information of the sample and forms a temporal similarity matrix by calculating the Euclidean distance between the temporal feature vectors of the anchor point and the positive and negative samples.
[0029] The spatial similarity matrix and the temporal similarity matrix are weightedly fused to generate a comprehensive similarity matrix.
[0030] As a preferred solution of the method for detecting targets in any direction in remote sensing images according to the present invention, the method comprises the following steps: calculating the self-supervision loss value and the detection loss value based on the comprehensive similarity matrix, and performing weighted summation to generate a joint loss function.
[0031] The similarity values between the anchor points and the positive and negative samples in the comprehensive similarity matrix are used to calculate the contrast loss, obtain the self-supervised loss value, and calculate the detection loss value through the cross entropy loss function and the smooth loss function;
[0032] Using a weighted learning algorithm, the weights of self-supervision loss and detection loss are automatically adjusted according to the importance of the task, and a weighted sum is performed to generate a joint loss function.
[0033] The steps for obtaining the detection frame set are as follows:
[0034] The angle features in the joint loss function are mapped to the frequency domain using Fourier transform, and the angle-sensitive basis vectors in the spatial domain are generated through inverse transform.
[0035] The angle-sensitive basis vectors are convolved with the spatiotemporal feature map at multiple scales through a convolutional neural network to generate a feature response map reflecting the target angle and position.
[0036] The Softmax function is used to activate the feature response map to generate a probability distribution, and the target detection box set is obtained through parameter decoding and non-maximum suppression methods.
[0037] As a preferred solution of the method for detecting targets in any direction in remote sensing images of the present invention, the steps for obtaining the detection frame set are as follows:
[0038] The angle features in the joint loss function are mapped to the frequency domain using Fourier transform, and the angle-sensitive basis vectors in the spatial domain are generated through inverse transform.
[0039] The angle-sensitive basis vectors are convolved with the spatiotemporal feature map at multiple scales through a convolutional neural network to generate a feature response map reflecting the target angle and position.
[0040] The Softmax function is used to activate the feature response map to generate a probability distribution, and the target detection box set is obtained through parameter decoding and non-maximum suppression methods.
[0041] As a preferred solution of the method for detecting targets in any direction in remote sensing images of the present invention, wherein: the detection frame uses a soft-core non-maximum suppression algorithm to eliminate redundant detection frames, and outputs a detection report that satisfies any direction targets, specifically comprising the following steps:
[0042] Use the Soft-NMS algorithm to eliminate redundant detection boxes on the detection box set;
[0043] The position, size, rotation angle, confidence score and category information of the target in the remote sensing image are extracted from the detection frame set after Soft-NMS processing, and compiled into the final detection report.
[0044] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the method for detecting targets in any direction in remote sensing images as described in the first aspect of the present invention is implemented.
[0045] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the method for detecting targets in any direction in remote sensing images as described in the first aspect of the present invention is implemented.
[0046] The beneficial effects of the present invention are as follows: optical flow field estimation is performed through an improved PWC-Net network, and static semantic features and temporal correlation features are extracted in combination with a dual-branch feature extraction network to generate a spatiotemporal feature map. At the same time, a soft-core non-maximum suppression algorithm is used to eliminate redundant detection frames, achieving accurate detection of targets in any orientation. This innovation greatly improves the accuracy of target detection, especially when dealing with tilted or rotated targets. Based on the improved PWC-Net network, multi-scale features are extracted layer by layer and optimized in combination with multispectral optical flow to generate a high-precision optical flow density map. The spatial and temporal information of the samples is extracted through a dual encoder architecture, which not only improves the robustness and stability of detection, but also maintains efficient detection capabilities in complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 This is a flow chart of the method for detecting targets in any direction in remote sensing images in Example 1.
[0049] Figure 2 This is a flowchart of obtaining an optical flow density map using the improved PWC-Net network in Example 1.
[0050] Figure 3 This is a flowchart for obtaining the optical flow density map in Example 1.
[0051] Figure 4 This is a flowchart of target detection in any direction of remote sensing images in Example 1. DETAILED DESCRIPTION
[0052] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0053] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0054] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0055] Example 1, with reference to Figures 1 to 4 , which is the first embodiment of the present invention, provides a method for detecting targets in any direction in a remote sensing image, comprising the following steps:
[0056] S1. Input the remote sensing image dataset into the improved PWC-Net network to estimate the optical flow field and obtain the optical flow density map.
[0057] The specific steps include:
[0058] Preprocess the input remote sensing image data. This step mainly includes radiometric correction and geometric registration. Radiometric correction uses mathematical models to adjust the pixel values in the image to eliminate deviations caused by factors such as atmospheric conditions and sensor characteristics, ensuring that images taken at different time points are comparable. Common radiometric correction methods include the 6S model. Then, geometric registration is performed to accurately align images acquired at different times or perspectives. This usually involves feature point detection and matching techniques, such as SIFT and SURF, to align remote sensing images to a common spatial reference frame.
[0059] After preprocessing the remote sensing image data, an improved PWC-Net network is constructed. The improved PWC-Net network is optimized based on the classic PWC-Net and is suitable for the multispectral characteristics and optical flow estimation requirements of remote sensing images. The improved PWC-Net network mainly includes multispectral flow, optical flow pyramid structure, adaptive learning mechanism, and contrastive learning loss function.
[0060] The core principle of the adaptive learning mechanism is to enable the improved PWC-Net network to dynamically adjust its parameters and structure according to the characteristics of the input, thereby improving its adaptability to complex or changing environments. In specific implementations, the adaptive learning mechanism usually combines the advanced technologies of attention mechanism and deformable convolution. During the construction process, the attention mechanism is first introduced to enhance the focus on important features. For example, spatial attention or channel attention mechanism is used to enable the improved PWC-Net network to automatically learn which areas or features are more critical. Then, deformable convolution is used to replace the traditional convolution operation, so that the convolution kernel can be deformed according to the local information of the input data, thereby enhancing the ability to process diverse data.
[0061] The optical flow pyramid structure allows the improved PWC-Net network to extract multi-scale features layer by layer from low resolution to high resolution. Each layer uses convolution operations to generate preliminary optical flow predictions and uses residual connections to fuse the optical flow prediction results of the previous layer with the feature map of the current layer. This can gradually refine the optical flow field estimation. The adaptive learning mechanism enables the network to automatically adjust parameters based on the input data, improving feature extraction. The contrastive learning loss function aims to maximize the differences between different samples while minimizing the distance between similar samples, enhancing the discriminative ability of the improved PWC-Net network.
[0062] After the improved PWC-Net network is built, the pre-processed remote sensing image is input into it for optical flow field estimation. A rough optical flow estimate is generated through the network bottom layer of the optical flow pyramid, that is, the image layer with the lowest resolution.
[0063] As the improved PWC-Net network goes deeper, each layer receives the optical flow prediction results of the previous layer and combines them with the feature map of the current layer to further refine the optical flow field estimation through convolution operations. In the optical flow estimation process, the information of multispectral flow is combined to obtain a high-precision optical flow density map. The optical flow density map is obtained by iteratively correcting the optical flow prediction results of each layer and reflects the changes of ground objects over time. Multispectral flow helps to accurately capture the motion characteristics of ground objects in remote sensing images and is crucial to improving the quality of optical flow density maps.
[0064] Through the above operations, optical flow information can be extracted efficiently and accurately from remote sensing images, providing strong support for land feature change monitoring, disaster assessment and other tasks.
[0065] S2. Input the optical flow density map into the dual-branch feature extraction network to extract static semantic features and temporal correlation features, and apply the gated attention mechanism to weighted fusion to generate a spatiotemporal feature map.
[0066] The specific steps include:
[0067] To extract high-quality spatiotemporal features from the optical flow density map, a dual-branch feature extraction network is designed. It consists of two branches: a convolutional neural network branch and a temporal correlation feature branch. In the convolutional neural network branch, static semantic features in the image are gradually extracted through a series of convolutional layers, batch normalization layers, and residual connections. These static features include information about the shape, texture, and color of the object. Each convolutional layer scans the local receptive field of the input optical flow density map and enhances the expressive power of the convolutional neural network branch through the ReLU nonlinear activation function. In addition, to reduce the gradient vanishing problem and accelerate the training process, a residual block is introduced into the convolutional neural network, allowing the convolutional neural network to learn the difference between the input and output, thereby building a deeper convolutional neural network structure.
[0068] In the temporal correlation feature branch, a recurrent neural network is used to capture the dynamic trends between consecutive frames of the optical flow density map. This not only focuses on the information within a single frame, but also on the motion patterns and correlations between frames. By using the hidden state of the previous frame as one of the inputs of the current frame, it can effectively model dependencies in the temporal dimension. At the same time, the use of an attention mechanism allows the recurrent neural network model to automatically focus on the most important areas when processing each frame, further improving the performance of the recurrent neural network model and ensuring that the recurrent neural network model can accurately identify key action sequences.
[0069] Based on static semantic features and temporal correlation features, a gated attention mechanism is adopted. The core of the gated attention mechanism is to calculate the importance score of each feature and assign corresponding weights to each feature accordingly. In specific operations, the gated attention mechanism will first perform dimensionality reduction on the input static semantic features and temporal correlation features to reduce the amount of calculation and improve efficiency. Then, MLP (multi-layer perceptron) is used to calculate the feature correlation matrix between static semantic features and temporal correlation features. The feature correlation matrix reflects the degree of interaction between different features and is the key basis for determining the final weighting coefficient. In the process of calculating the feature correlation matrix, regularization technology is applied to prevent overfitting. Finally, the softmax function is applied to the similarity matrix feature correlation matrix to obtain the weighting coefficient and complete the weighted fusion of static semantic features and temporal correlation features. The weighted fusion process allows the convolutional neural network and recurrent neural network models to flexibly adjust the contribution of different features according to the specific situation, so as to better adapt to complex scene changes.
[0070] After obtaining the fused feature map, further refinement is performed to generate the final spatiotemporal feature map. Since the feature extraction process may cause the resolution of the feature map to decrease, it is necessary to restore its original size through deconvolution operation. The deconvolution layer can upsample the low-resolution feature map to the target size by learning a set of filter parameters while retaining as much detail information as possible. Secondly, in order to reduce information loss and improve the quality of feature representation, multiple residual blocks are used in the refinement process. Each residual block contains several convolution layers and batch normalization layers, which work together to enhance the learning ability and generalization performance of the dual-branch feature extraction network. The use of residual connections makes the dual The branch feature extraction network makes it easier to train deep structures because the residual connection allows signals to be passed directly from one layer to another without going through all the intermediate layers. Finally, considering that there may be objects and actions of various scales in the scene, the spatial pyramid pooling technology is used to capture multi-scale spatial information. The spatial pyramid pooling technology generates a series of feature descriptors with different receptive fields by pooling feature maps at different scales. These descriptors are then spliced together to form a spatiotemporal feature map that comprehensively reflects the information of various scales in the scene. In addition, the entire refinement process also uses the ReLU activation function to increase nonlinear expression capabilities.
[0071] Through the above operations, rich spatiotemporal features are extracted from the optical flow density map, and high-quality spatiotemporal feature maps are generated, which not only greatly improves the ability to understand and analyze complex scenes, but also provides strong support for subsequent tasks.
[0072] S3. Generate triplet samples based on the spatiotemporal feature map and input them into the dual encoder architecture to obtain the comprehensive similarity matrix by constructing the spatial similarity matrix and the temporal similarity matrix.
[0073] The specific steps include:
[0074] In order to obtain high-quality triple samples from the spatiotemporal feature maps, data augmentation operations are first performed on these feature maps to generate a rich set of candidate samples. Specifically, random rotation transformation and frequency domain perturbation are used to increase sample diversity. Random rotation transformation uses OpenCV to rotate the image within a certain angle range (such as ±30 degrees) to simulate targets in different directions. Frequency domain perturbation is performed by Fourier transform using the FFT function in the SciPy library, and Gaussian noise is added in the frequency domain. Then, the inverse Fourier transform is used to restore it to the spatial domain to simulate different lighting conditions and noise environments.
[0075] Based on the generated candidate sample set, target region features are selected to construct a triplet sample consisting of an anchor point, a positive sample, and a negative sample. The target region features are defined based on multi-dimensional information such as the target's position, shape, texture, and motion trajectory in the spatiotemporal feature map. The most representative target region is screened out using a convolutional neural network combined with the non-maximum suppression (NMS) algorithm. For each anchor point, targets belonging to the same category are selected as positive samples, ensuring that the two are similar but not identical. At the same time, targets of different categories are selected as negative samples, ensuring that they are as different from the anchor point as possible.
[0076] We use difficult sample mining technology to optimize the triplet sample selection process. We employ the Online Hard Sample Mining (OHEM) method to dynamically select the samples with the largest loss value for training based on the model's current prediction results. Specifically, we calculate the loss value for each triplet sample during each training round, and select the top K samples with the largest loss values as difficult samples for priority updating. This method can focus more on samples that are easily misclassified, thereby improving overall performance and robustness.
[0077] The triplet sample is input into the spatial feature encoder of the dual-encoder architecture, which is mainly composed of a convolutional neural network to extract the spatial information of the triplet sample. In specific operations, a series of convolutional layers gradually capture low-level features in the image, such as edges and corners, and high-level features, such as object contours and textures. Each convolution operation is usually followed by batch normalization and ReLU activation functions to enhance the expressive power and accelerate the training process. In this process, each triplet sample will generate a high-dimensional spatial feature vector, which contains rich spatial information.
[0078] Next, the Euclidean distance between the anchor point and the positive sample and between the anchor point and the negative sample in the spatial feature vector is calculated. The mathematical expression is:
[0079]
[0080] Where D(p,q) represents the Euclidean distance between sample p and sample q, α represents the balance parameter, n represents the number of dimensions of the spatial feature vector, i represents the index of the dimension of the spatial feature vector, and w i Represents the weight of the eigenvector in the i-dimensional space, p i Represents the value of the feature vector of sample p in the i-th dimension space, q i represents the value of the eigenvector of sample q in the i-th dimension space, and θ represents the angle between sample p and sample q;
[0081] The calculated Euclidean distance can be used to quantify the similarity between triplet samples. For example, if the distance between the anchor point and the positive sample is small, it indicates that they are very similar in spatial features. Conversely, if the distance between the anchor point and the negative sample is large, it indicates that they have significant differences in spatial features, thereby obtaining a spatial similarity matrix.
[0082] The same triplet samples are then input into the temporal feature encoder, and the long short-term memory network is used to extract the temporal information of the triplet samples. Specifically, the time series data is input into the long short-term memory network to capture the temporal correlation between the triplet samples. Since remote sensing images usually contain multiple frames of data, the long short-term memory network can effectively model the dynamic change relationship between these frames. After each frame of data is processed by the long short-term memory network, a temporal feature vector is generated, which reflects the motion trajectory and change trend in the time dimension.
[0083] Similarly, the Euclidean distance between the temporal feature vectors of the anchor point and the positive sample, as well as between the anchor point and the negative sample, is calculated to form a temporal similarity matrix. This allows us to identify which samples are more similar in temporal features and which are significantly different. For example, when the distance between the temporal feature vectors of two triplet samples is small, it indicates that their behavior patterns in the temporal dimension are relatively consistent. Otherwise, it indicates that their behavior patterns are significantly different, thus obtaining a matrix reflecting the temporal similarity between samples.
[0084] The spatial similarity matrix and the temporal similarity matrix are weightedly fused to generate the final comprehensive similarity matrix. The key to this step is to determine the appropriate weight distribution strategy to balance the importance of spatial and temporal information. A common practice is to adjust the weight based on the experience or experimental results of specific application scenarios. For example, in some scenarios, the position and shape of the target in the remote sensing image may be more important than its motion trajectory. In this case, the weight of the spatial similarity matrix can be increased. In other scenarios, the change trend and dynamic behavior of the target may be the key factors, so the weight of the temporal similarity matrix should be increased.
[0085] After completing the weight assignment, the two similarity matrices are added together according to the weights to obtain a comprehensive similarity matrix. This matrix not only takes into account the spatial similarity of the samples, but also combines their correlation in the temporal dimension, providing a more comprehensive target representation. For example, when detecting fast-moving targets, information in the temporal dimension is particularly important because it can help distinguish different motion patterns. In static target detection, spatial features may dominate. Through this comprehensive similarity matrix, similarity can be accurately evaluated and provide strong support for subsequent target detection tasks.
[0086] S4. Calculate the self-supervision loss value and the detection loss value based on the comprehensive similarity matrix, and perform weighted summation to generate a joint loss function.
[0087] The specific steps include:
[0088] The comprehensive similarity matrix is used to calculate the contrastive loss to obtain the self-supervised loss value. The core of this process is to distinguish the differences between similar and dissimilar samples. Specifically, for each triple sample including the anchor point, positive sample, and negative sample, the corresponding similarity score is extracted from the comprehensive similarity matrix. The similarity score reflects the similarity between the samples in terms of spatial and temporal features. The contrastive loss function ensures that the distance between similar samples can be effectively shortened and the distance between dissimilar samples can be expanded. This step not only improves the quality of the representation of static semantic features and temporal correlation features, but also enhances the ability to understand complex data.
[0089] Next is the process of calculating the detection loss value, which mainly consists of two parts: cross entropy loss and Smooth L1 loss. Cross entropy loss is mainly used for classification tasks and measures the difference between the predicted category and the true category. This loss function can identify the target category in the image, while Smooth L1 loss is used for regression tasks, such as bounding box regression. By calculating the error between the predicted box and the true box, the position and size of the bounding box are optimized. This loss function is more robust to outliers and helps improve the accuracy of bounding box positioning. By combining self-supervised loss and detection loss, the classification and regression tasks in target detection can be optimized simultaneously to ensure the best performance on different tasks.
[0090] A single loss function may not be able to fully meet the complex needs of remote sensing image target detection. Therefore, an adaptive weight learning algorithm is used to automatically adjust the weights of the self-supervision loss and the detection loss according to the importance of the task. The core of this algorithm is to dynamically adjust the weights of the two loss terms to balance the needs of different tasks. For example, during training, if the detection loss value is found to decrease slowly, its weight can be appropriately increased. Conversely, if the self-supervision loss value is high, its weight can be increased. The optimization direction can be flexibly adjusted according to actual performance to better adapt to different application scenarios.
[0091] The self-supervision loss and detection loss are weighted and summed to generate the final joint loss function. The key to this step is to reasonably set the initial loss weight and continuously optimize the initial loss weight during training to achieve the best performance. The joint loss function not only considers the similarity relationship between samples, but also combines the specific requirements of classification and regression tasks, providing a more comprehensive optimization goal. In this way, effective training can be obtained in multiple dimensions, thereby significantly improving the accuracy and robustness of arbitrary target detection.
[0092] S5. Based on the joint loss function, the angle-sensitive basis vector is generated through Fourier transform, and convolved with the spatiotemporal feature map to generate a feature response map. The feature response map is activated by Softmax and parameter decoding is performed to obtain a set of detection boxes.
[0093] The specific steps include:
[0094] Fast Fourier transform (FFT) is used to map the angular features in the joint loss function to the frequency domain. FFT is an efficient algorithm for calculating discrete Fourier transforms, which can convert signals in the time or spatial domain into frequency domain representations. In the specific operation, Python is used to perform a two-dimensional Fourier transform to capture the characteristic information of different frequency components, especially those frequency components related to specific angles.
[0095] Next, the corresponding frequency components are selected according to the direction of the target to generate angle-sensitive basis vectors. In the frequency domain, those frequency components that can reflect specific angle characteristics are selected. For example, some frequency components may correspond to features in the horizontal direction, while other components may correspond to features in the vertical or other inclined angles. In this way, a set of basis vectors with direction sensitivity is generated. These basis vectors are then converted back to the spatial domain through inverse Fourier transform for further processing with the spatiotemporal feature map.
[0096] After generating the angle-sensitive basis vectors, the next step is to perform multi-scale convolution on them with the spatiotemporal feature map. The core of this step is to enhance the ability to recognize objects of different scales. Multi-scale convolution uses convolution kernels of different sizes, making it possible to focus on both small and large objects in the same layer. For example, smaller convolution kernels (such as 3x3) help capture detailed information, while larger convolution kernels (such as 5x5 or 7x7) are more suitable for capturing global structures. This multi-level convolution operation can significantly improve detection accuracy.
[0097] In terms of specific implementation, convolutional neural networks can be built through deep learning frameworks such as TensorFlow. In PyTorch, the torch.nn.Conv2d function can be used to define the convolution layer, and multi-scale convolution can be achieved by adjusting the size and step size of the convolution kernel. The result of the convolution operation is a feature response map, which reflects the response intensity of the target at different angles and positions. Each pixel value represents the possibility of the target at the corresponding position. The feature response map not only provides rich detailed information, but also lays the foundation for subsequent probability distribution calculations.
[0098] Softmax activation of the feature response map is a key step. The Softmax function can convert the input real value into a probability value in the range of [0,1] and ensure that the sum of the probabilities of all categories is 1. This means that the feature response map after Softmax activation is actually a probability distribution map, in which each pixel value represents the possibility of the existence of a target at that location. In practical applications, Softmax activation can help better understand the information in the feature response map and provide support for subsequent target positioning. In PyTorch, you can use the torch.nn.Softmax function to implement Softmax activation;
[0099] Then, parameter decoding is used to convert the feature response map after Softmax activation into the location and size information of the target. The decoding process involves converting the output prediction box coordinates and confidence scores into the actual target detection box. Usually, a series of prediction box coordinates and confidence scores are output. This information needs to be converted into the actual target detection box through the decoder. The actual size and proportion of the image also need to be considered during the decoding process to ensure the accuracy of the detection box.
[0100] Finally, non-maximum suppression is applied to remove redundancy, avoid repeated detection of the same target, and improve the accuracy of the detection results.
[0101] S6. Use the soft-core non-maximum suppression algorithm to eliminate redundant detection frames and output a target detection report that satisfies any direction.
[0102] The specific steps include:
[0103] First, remove those redundant detection frames with high overlap from the detection frame set. Although the traditional non-maximum suppression (NMS) algorithm can effectively remove some redundant frames, it sets a fixed overlap threshold to decide whether to retain a detection frame. This method may mistakenly delete some detection frames that overlap with high-confidence frames but are still important. To solve this problem, the soft-core non-maximum suppression (Soft-NMS) algorithm is adopted. The core idea of Soft-NMS is to reduce the confidence scores of detection frames with high overlap rather than directly delete them. In the specific operation, a custom function is written in Python to implement the Soft-NMS algorithm. First, the intersection over union (IoU) between all detection frames is calculated, and then the confidence score of each detection frame is adjusted according to the IoU value. After this step, an optimized detection frame set is obtained, which not only has a higher confidence but also has a lower overlap with each other.
[0104] Based on the set of detection frames processed by Soft-NMS, the key information of the target in the remote sensing image is extracted and organized into the final detection report. Each detection frame usually contains multiple important attributes, such as the target's position, size, rotation angle, confidence score, and category label. The position and size information can help us determine the specific position of the target in the image, the rotation angle is used to describe the direction of the target, the confidence score reflects the accuracy of the detection result, and the category label indicates the specific type of the target. When extracting key information, it is necessary to ensure that each key information contains important attributes. In order to facilitate subsequent analysis and display, this information can be organized into a structured format, such as CSV file or JSON format. Each record represents a detected target and lists all its attributes in detail. For example, in a JSON format report, each record will clearly list the target's center point coordinates, width, height, rotation angle, confidence score, and category label information;
[0105] Finally, the detection frame is drawn on the original remote sensing image to generate an image with a detection frame. This process not only enhances the visual effect and makes the detection results more intuitive, but also provides users with a convenient way to verify the accuracy of the detection results. Use OpenCV's function to load the original remote sensing image, and then call the corresponding drawing function for each detection frame to draw it on the image. Since the target in the remote sensing image may be a rotated rectangle, special attention should be paid to the processing of the rotation angle when drawing. By calculating the coordinates of the four vertices of the rotated rectangle, the polylines function is used to draw the polygon, thereby accurately depicting the shape and direction of the target. In addition to drawing the detection frame, the corresponding category label and confidence score can be added next to each detection frame by calling the putText function. In this way, the user can obtain detailed information of the target directly from the image without having to refer to additional report files. After completing all drawing operations, use the imwrite function to save the final image for subsequent viewing or further analysis.
[0106] This embodiment also provides a computer device suitable for the method of detecting targets in any direction in remote sensing images, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the method of detecting targets in any direction in remote sensing images proposed in the above embodiment.
[0107] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0108] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for detecting targets in any direction in remote sensing images as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0109] In summary, the present invention uses an improved PWC-Net network to estimate optical flow fields, combines it with a dual-branch feature extraction network to extract static semantic features and temporal correlation features, generates a spatiotemporal feature map, and combines it with a soft-core non-maximum suppression algorithm to eliminate redundant detection frames, achieving accurate detection of targets in any orientation. This innovation greatly improves target detection accuracy, especially when dealing with tilted or rotated targets. Based on the improved PWC-Net network, multi-scale features are extracted layer by layer and optimized in combination with multispectral optical flow to generate a high-precision optical flow density map. The spatial and temporal information of the samples is extracted through a dual encoder architecture, which not only improves the robustness and stability of detection, but also maintains efficient detection capabilities in complex dynamic environments.
[0110] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for detecting targets in arbitrary directions in remote sensing images, characterized by: include, The remote sensing image dataset is input into the improved PWC-Net network to estimate the optical flow field and obtain the optical flow density map; The optical flow density map is input into a dual-branch feature extraction network to extract static semantic features and temporal correlation features, and a gated attention mechanism is applied to weighted fusion to generate a spatiotemporal feature map. Generate triplet samples based on the spatiotemporal feature map and input them into the dual encoder architecture to obtain the comprehensive similarity matrix by constructing the spatial similarity matrix and the temporal similarity matrix; Calculate the self-supervision loss value and the detection loss value based on the comprehensive similarity matrix, and perform weighted summation to generate a joint loss function; Based on the joint loss function, the angle-sensitive basis vector is generated through Fourier transform, and convolved with the spatiotemporal feature map to generate a feature response map. The feature response map is activated by Softmax and parameter decoding is performed to obtain a set of detection boxes. The soft-core non-maximum suppression algorithm is used to eliminate redundant detection frames and output a target detection report that satisfies any direction.
2. The method for detecting targets in arbitrary directions using remote sensing images according to claim 1, wherein: The steps for obtaining the optical flow density map are as follows: Perform radiometric correction and geometric registration on the input remote sensing image data; Based on the multi-spectral flow and optical flow pyramid structure, an improved PWC-Net network is constructed by embedding an adaptive learning mechanism and a contrastive learning loss function. The preprocessed remote sensing image data is input into the improved PWC-Net network, and multi-scale features are extracted layer by layer through the optical flow pyramid structure. The first layer uses convolution operations to generate preliminary optical flow predictions, and the second layer uses residual connections to fuse the optical flow predictions of the first layer with the features of the current layer to obtain optical flow field estimates. The optical flow density map is generated by refining the optical flow field estimation layer by layer and optimizing it in combination with multispectral optical flow.
3. The method for detecting targets in any direction in remote sensing images according to claim 1, wherein: The steps for obtaining the spatiotemporal feature map are as follows: The optical flow density map is input into a dual-branch feature extraction network. The convolutional neural network branch extracts static semantic features of shape and texture, and the temporal correlation feature branch extracts temporal correlation features of motion trajectory and change trend. The extracted static semantic features and temporal correlation features are applied with a gated attention mechanism to calculate the feature importance scores. The fully connected layers and activation functions are then connected to generate different weights, and the fused feature map is generated through weighted summation. The fused feature map is refined to generate a spatiotemporal feature map.
4. The method for detecting targets in arbitrary directions using remote sensing images according to claim 1, wherein: The steps for obtaining the triplet sample are as follows: Perform random rotation transformation and frequency domain perturbation on the spatiotemporal feature map to generate a candidate sample set; Select target region features from the candidate sample set, construct triple samples containing anchor points, positive samples, and negative samples, and optimize the triple samples through difficult sample mining method; The target area features are defined based on the features of the position, shape, texture and motion trajectory of the target in the spatiotemporal feature map.
5. The method for detecting targets in any direction in remote sensing images according to claim 4, wherein: The steps for obtaining the comprehensive similarity matrix are as follows: The triplet sample is input into the dual encoder architecture. The spatial feature encoder extracts the spatial information of the sample through a convolutional neural network and forms a spatial similarity matrix by calculating the Euclidean distance between the anchor point and the spatial feature vectors of the positive sample and the negative sample. The temporal feature encoder uses a recurrent neural network to extract the temporal information of the sample and forms a temporal similarity matrix by calculating the Euclidean distance between the temporal feature vectors of the anchor point and the positive and negative samples. The spatial similarity matrix and the temporal similarity matrix are weightedly fused to generate a comprehensive similarity matrix.
6. The method for detecting targets in any direction in remote sensing images according to claim 5, characterized in that The self-supervision loss value and the detection loss value are calculated based on the comprehensive similarity matrix, and a weighted sum is performed to generate a joint loss function, which specifically includes the following steps: The similarity values between the anchor points and the positive and negative samples in the comprehensive similarity matrix are used to calculate the contrast loss, obtain the self-supervised loss value, and calculate the detection loss value through the cross entropy loss function and the smooth loss function; Using a weighted learning algorithm, the weights of self-supervision loss and detection loss are automatically adjusted according to the importance of the task, and a weighted sum is performed to generate a joint loss function.
7. The method for detecting targets in any direction using remote sensing images according to claim 1, wherein: The steps for obtaining the detection frame set are as follows: The angle features in the joint loss function are mapped to the frequency domain using Fourier transform, and the angle-sensitive basis vectors in the spatial domain are generated through inverse transform. The angle-sensitive basis vectors are convolved with the spatiotemporal feature map at multiple scales through a convolutional neural network to generate a feature response map reflecting the target angle and position. The Softmax function is used to activate the feature response map to generate a probability distribution, and the target detection box set is obtained through parameter decoding and non-maximum suppression methods.
8. The method for detecting targets in any direction in remote sensing images according to claim 7, wherein: The method uses a soft-core non-maximum suppression algorithm to eliminate redundant detection frames and output a target detection report that satisfies any direction, specifically including the following steps: Use the Soft-NMS algorithm to eliminate redundant detection boxes on the detection box set; The position, size, rotation angle, confidence score and category information of the target in the remote sensing image are extracted from the detection frame set after Soft-NMS processing, and compiled into the final detection report.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for detecting targets in any direction in remote sensing images according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting targets in any direction in remote sensing images according to any one of claims 1 to 8 are implemented.