Directed Object Detection Method for Remote Sensing Images Based on Dual-Domain Feature Fusion
Through the adaptive selection and fusion of airspace and frequency domain features, the problem of insufficient feature expression in remote sensing image object detection is solved, and higher detection accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202510403166.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-01
AI Technical Summary
In the existing remote sensing image object detection methods, the expression of target features is limited only through single spatial domain feature extraction, resulting in poor detection and positioning accuracy and the complex background and weak features of remote sensing images cannot be fully utilized.
Using a method based on dual-domain feature fusion, combining spatial and frequency domain adaptive selection and dual-domain feature interactive fusion, local and context information is extracted through the spatial and frequency domain adaptive selection module and frequency domain adaptive selection module, and complementary features are obtained in the dual-domain feature interaction module to enhance the expression of target information.
It effectively enhances the integration of global context information and local information, improves the accuracy and target discernibility of remote sensing target detection, and makes up for the shortcomings of single spatial domain characteristics.
Smart Images

Figure CN119919819B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing target detection, and particularly relates to a method for directed target detection of remote sensing images based on dual-domain feature fusion. Background Technique
[0002] At present, excellent performance has been achieved for general target detection algorithms in natural scenes, and their robustness and generalization ability have been well verified. However, due to certain differences in the target characteristics between remote sensing images and natural scene images, such as remote sensing images having a wide field of view, various target directions, uneven distribution, and weak feature expression.
[0003] Therefore, a large number of researchers have made improvements and designs for the different characteristics of remote sensing images based on migration algorithms. High-performance remote sensing object detectors usually rely on the RCNN framework, which consists of a region proposal network and a detection head. The RPN proposes high-quality regions of interest from the backbone feature map, and the region neural network detection head is responsible for object classification and bounding box regression. In recent years, several variants of the RCNN (Region-based Convolutional Neural Networks) framework have been proposed. Such methods have improved the performance of directed target detection of remote sensing images by improving and designing various modules based on convolutional neural networks. However, the feature extraction only through a single spatial domain limits the expression of target features and cannot exert greater potential in the remote sensing target detection task.
[0004] Regarding the problem of complex and variable backgrounds and weak target feature expression in remote sensing images, the current method of only extracting target features through the spatial domain is difficult to fully utilize the information of features, resulting in poor target detection and positioning accuracy. Summary of the Invention
[0005] In view of this, the present invention aims to provide a method for directed target detection of remote sensing images based on dual-domain feature fusion. By the adaptive selection of the spatial domain and the frequency domain and the feature interaction and fusion between the two domains, the complementary features between different domains are fully utilized, effectively enhancing the fusion of global context information and local information and making up for the lack of target information.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows:
[0007] A method for directed target detection of remote sensing images based on dual-domain feature fusion, comprising:
[0008] S1: Obtain a remote sensing image dataset, and preprocess the remote sensing image dataset to obtain a training set;
[0009] S2: Construct a remote sensing object detection network, where the remote sensing object detection network includes:
[0010] A feature extraction branch that extracts initial features of different sizes from the input remote sensing image;
[0011] A dual-domain feature fusion branch that extracts corresponding spatial domain features and frequency domain features from the initial features of different sizes; fuses the corresponding spatial domain features and frequency domain features of the same size to obtain corresponding fusion features;
[0012] An object detection branch that determines the corresponding category and location of the remote sensing image based on different fusion features;
[0013] S3: Use the training set obtained in step S1 to train the remote sensing object detection network in step S2 to obtain a remote sensing object detection model;
[0014] S4: Input the remote sensing image to be detected into the remote sensing object detection model in step S3, and output the corresponding detection category and location.
[0015] Furthermore, the dual-domain feature fusion branch in step S2 includes: a spatial domain adaptive selection module that extracts corresponding spatial domain features from the initial features; a frequency domain adaptive selection module that extracts corresponding frequency domain features from the initial features; a dual-domain feature interaction module that fuses the corresponding spatial domain features and frequency domain features of the same size to obtain corresponding preliminary fusion features; a feature pyramid module that extracts multi-scale information from the preliminary fusion features to obtain fusion features.
[0016] Furthermore, the spatial domain adaptive selection module includes a feature extraction sub-module, a pooling sub-module, a multi-scale extraction sub-module, and a feature output sub-module; perform a convolution operation on the input feature to obtain a preprocessed feature; input the preprocessed feature into the feature extraction sub-module to obtain a first feature and a second feature; input the first feature and the preprocessed feature into the pooling sub-module to obtain a pooled feature; input the pooled feature into the multi-scale extraction sub-module for multi-scale information extraction to obtain multi-scale features; input the multi-scale features, the preprocessed feature, the second feature, and the input feature into the feature output sub-module together to obtain spatial domain features.
[0017] Furthermore, in the feature extraction sub-module: perform convolution operations of at least 2 different scales on the preprocessed feature simultaneously, and add the corresponding elements of the features output by the convolution operations to obtain the first feature, and the feature output by the convolution operation with the largest scale is the second feature.
[0018] Furthermore, in the pooling sub-module: concatenate the first feature and the preprocessed feature; perform average pooling and max pooling on the concatenated feature simultaneously, and then perform pointwise convolution on the features after the two poolings to obtain a pooled feature.
[0019] Further, in the multi-scale extraction sub-module: the pooled features are subjected to channel separation, the separated features are respectively subjected to convolution operations of different scales, and then the convolved features are concatenated to obtain multi-scale features.
[0020] Further, in the feature output sub-module: a sigmoid activation operation is performed on the multi-scale features, the features after the sigmoid activation operation are respectively multiplied by the corresponding elements of the pre-processed features and the second features, and then the corresponding elements of the two features obtained by multiplication are added to obtain the spatial domain weight; the spatial domain weight is multiplied by the corresponding elements of the input features to obtain the spatial domain features.
[0021] Further, in the frequency domain adaptive selection module: global feature pooling and fast Fourier transform are respectively performed on the input features; the features obtained by global feature pooling are subjected to multiple convolution operations and then subjected to sigmoid activation operation to obtain the frequency domain weight; after the frequency domain weight is multiplied by the corresponding elements of the static filter, an adaptive filter is obtained; the features obtained by fast Fourier transform are operated by the adaptive filter, and then the obtained features are subjected to inverse fast Fourier transform to obtain the frequency domain features.
[0022] Further, in the dual-domain feature interaction module: attention feature extraction is respectively performed on the spatial domain features and the frequency domain features. The spatial domain features correspondingly obtain the second query matrix, the first key matrix and the first value matrix, and the frequency domain features correspondingly obtain the first query matrix, the second key matrix and the second value matrix; after the first query matrix, the first key matrix and the first value matrix are multiplied, and then after Softmax operation, convolution operation is performed to obtain the third feature; after the second query matrix, the second key matrix and the second value matrix are multiplied, and then after Softmax operation, convolution operation is performed to obtain the fourth feature; after the third feature and the fourth feature are concatenated, the obtained concatenated features are respectively subjected to average pooling and max pooling, and the corresponding elements of the features obtained by the two poolings are added and then subjected to convolution and sigmoid activation operations in sequence to obtain the fusion weight; the fusion weight is multiplied by the corresponding elements of the concatenated features to obtain the fusion features.
[0023] Further, the object detection branch in step S2 includes multiple detection heads, and the number of detection heads is the same as the number of initial features extracted by the feature extraction branch. Each detection head respectively processes the fusion features of the corresponding size to obtain the category and position corresponding to the remote sensing image, and then combines the categories corresponding to the fusion features of different sizes and all positions to determine the category and position corresponding to the remote sensing image; each detection head includes a category detection sub-module and a position detection sub-module; the fusion features of the corresponding size are respectively input into the category detection sub-module and the position detection sub-module. The category detection sub-module determines the category corresponding to the current fusion feature, and the position detection sub-module determines the position corresponding to the current fusion feature.
[0024] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0025] (1) In the method for directed object detection of remote sensing images based on dual-domain feature fusion according to the present invention, through the adaptive selection of the spatial domain and the frequency domain and the feature interaction and fusion between the two domains, the complementary features between different domains are fully utilized, effectively enhancing the fusion of global context information and local information and making up for the lack of target information;
[0026] (2) In the method for directed object detection of remote sensing images based on dual-domain feature fusion according to the present invention, by using the proposed spatial domain adaptive selection module and frequency domain adaptive selection module, the network model adaptively extracts local and context information according to the characteristics of the target in the spatial domain and the frequency domain, improving the discriminability of the target; by using the proposed dual-domain feature interaction module, complementary features are obtained through the fusion between the spatial domain and frequency domain features, making up for the deficiencies of single spatial domain features and enhancing the information expression ability of the target. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0028] Figure 1 is a schematic flowchart of the method for directed object detection of remote sensing images based on dual-domain feature fusion according to the embodiment of the present invention;
[0029] Figure 2 is a schematic diagram of the remote sensing object detection network according to the embodiment of the present invention;
[0030] Figure 3 is a schematic diagram of the spatial domain adaptive selection module according to the embodiment of the present invention;
[0031] Figure 4 is a schematic diagram of the frequency domain adaptive selection module according to the embodiment of the present invention;
[0032] Figure 5 is a schematic diagram of the dual-domain feature interaction module according to the embodiment of the present invention;
[0033] Figure 6 is a schematic diagram of the detection head according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0034] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0035] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0036] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0037] As Figures 1 to 2 shown, the method for directed object detection of remote sensing images based on dual-domain feature fusion according to the embodiment of the present invention includes:
[0038] S1: Obtain a remote sensing image data set, and preprocess the remote sensing image data set to obtain a training set.
[0039] In a certain embodiment, the remote sensing image data set selects the HRSC2016 data set publicly available at https: / / opendatalab.org.cn / OpenDataLab / HRSC2016, and the DIOR-R data set publicly available at https: / / opendatalab.org.cn / OpenDataLab / DIOR. The preprocessing consists of performing random flipping and random rotation operations on the remote sensing image data set, with probabilities of 0.75 and 0.25 respectively.
[0040] S2: Construct a remote sensing object detection network. Among them, the remote sensing object detection network includes a feature extraction branch, a dual-domain feature fusion branch, and an object detection branch. The feature extraction branch extracts initial features of different sizes from the input remote sensing image; the dual-domain feature fusion branch extracts corresponding spatial domain features and frequency domain features from the initial features of different sizes, and performs feature fusion on the corresponding spatial domain features and frequency domain features of the corresponding size to obtain corresponding fusion features; the object detection branch determines the corresponding category and location of the remote sensing image according to different fusion features.
[0041] In a specific embodiment, the feature extraction branch uses the ResNet50 network (the paper "Deep Residual Learning for Image Recognition" published on the arXiv platform) to extract initial features. The sizes of the extracted initial features are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the size of the input remote sensing image respectively. The input of the dual-domain feature fusion branch is composed of four feature map scales of different sizes.
[0042] In some embodiments, the dual-domain feature fusion branch includes a spatial domain adaptive selection module, a frequency domain adaptive selection module, a dual-domain feature interaction module, and a feature pyramid module. Among them, the spatial domain adaptive selection module extracts the corresponding spatial domain features from the initial features; the frequency domain adaptive selection module extracts the corresponding frequency domain features from the initial features; the dual-domain feature interaction module fuses the spatial domain features and frequency domain features of the corresponding size to obtain the corresponding preliminary fusion features; the feature pyramid module extracts the multi-scale information in the preliminary fusion features to obtain the fusion features. In a certain embodiment, the feature pyramid module adopts the FPN feature pyramid structure disclosed in the paper "Feature Pyramid Networks for Object Detection" from the arXiv platform.
[0043] In some embodiments, the structure of the spatial domain adaptive selection module is as Figure 3 shown, including a feature extraction sub-module, a pooling sub-module, a multi-scale extraction sub-module, and a feature output sub-module. First, a convolution operation is performed on the input features to obtain preprocessed features. The preprocessed features are input into the feature extraction sub-module to obtain the first feature and the second feature. The first feature and the preprocessed features are input into the pooling sub-module to obtain the pooled features. The pooled features are input into the multi-scale extraction sub-module for multi-scale information extraction to obtain multi-scale features. The multi-scale features, the preprocessed features, the second feature, and the input features are jointly input into the feature output sub-module to obtain the spatial domain features.
[0044] Specifically, in the feature extraction sub-module, the preprocessed features are simultaneously subjected to convolution operations of no less than 2 different scales, and the corresponding elements of the features output by the convolution operations are added together to obtain the first feature. The feature output by the convolution operation with the largest scale is the second feature. In the pooling sub-module, the first feature and the preprocessed feature are concatenated, and the concatenated feature is simultaneously subjected to average pooling and max pooling. The features after the two poolings are concatenated in channels, and then a pointwise convolution operation is performed on the concatenated feature to obtain the pooled feature. In the multi-scale extraction sub-module, the pooled feature is separated in channels, the separated features are respectively subjected to convolution operations of different scales, and then the convolved features are concatenated to obtain the multi-scale feature. In the feature output sub-module, a sigmoid activation operation is performed on the multi-scale feature, the features after the sigmoid activation operation are respectively multiplied by the corresponding elements of the preprocessed feature and the second feature, and then the corresponding elements of the two features obtained by the multiplication are added together to obtain the spatial domain weight; the spatial domain weight is multiplied by the corresponding elements of the input feature to obtain the spatial domain feature.
[0045] In a specific embodiment, as Figure 3 shown, other convolution operations except the pointwise convolution operation are depthwise separable convolution operations, and a depthwise separable convolution operation with a convolution kernel size of 3×3 is performed on the input feature to obtain the preprocessed feature. The feature extraction sub-module includes 4 depthwise separable convolution operations of different scales, and the convolution kernel sizes correspond to 5×5, 7×7, 9×9, and 11×11 respectively. At this time, it can be understood that the second feature is the feature obtained by the depthwise separable convolution operation with a convolution kernel size of 11×11. The multi-scale extraction sub-module includes 3 depthwise separable convolution operations of different scales, and the convolution kernel sizes correspond to 3×3, 1×11, and 11×1 respectively; the channel separation is channel equalization, that is, a feature with a channel number of C is divided into 4 features with a channel number of C / 4 each, and 3 of the features are respectively subjected to 3 depthwise separable convolution operations of different scales, and the output features of the 3 depthwise separable convolution operations of different scales are concatenated in channels with the remaining one feature to obtain the multi-scale feature.
[0046] In some embodiments, the frequency domain adaptive selection module, as Figure 4 shown, respectively performs global feature pooling and two-dimensional (2D) fast Fourier transform on the input feature; the feature obtained by global feature pooling is subjected to multiple convolution operations and then subjected to sigmoid activation operation to obtain the frequency domain weight; after the frequency domain weight is multiplied by the corresponding elements of the static filter, an adaptive filter is obtained; the feature obtained by the fast Fourier transform is operated by the adaptive filter, and the obtained feature is subjected to inverse fast Fourier transform to obtain the frequency domain feature.
[0047] It can be understood that the two-dimensional fast Fourier transform of the input feature is performed as follows:
[0048] ;
[0049] Among them, x[h, w] represents the input feature, h and w are the vertical and horizontal coordinates of each element in the input feature respectively, H and W represent the height and width of the input feature respectively, and X[h', w'] represents the feature obtained by the fast Fourier transform, and h' and w' represent the vertical and horizontal coordinates of each element in the feature obtained by the fast Fourier transform respectively.
[0050] After the feature obtained by the two-dimensional fast Fourier transform is operated by the adaptive filter, the obtained feature is subjected to the inverse fast Fourier transform as follows:
[0051] ;
[0052] Among them, X'[h', w'] represents the feature obtained by the operation of the adaptive filter, and x'[h, w] represents the frequency domain feature.
[0053] In a specific embodiment, global average pooling is used to complete global feature pooling, and the feature output by global feature pooling is continuously subjected to two two-dimensional (2D) convolution operations, and the ReLU activation operation is performed on the feature after the first two-dimensional convolution operation. It can be understood that the feature after the second two-dimensional convolution operation is subjected to the sigmoid activation operation to obtain the frequency domain weight. The static filter is a coefficient matrix composed of random numbers. It can be understood that the adaptive filter is a coefficient matrix with frequency domain weights, and the essence of the filtering operation of the adaptive filter is to perform a convolution operation on the feature obtained by the fast Fourier transform using the coefficient matrix with frequency domain weights.
[0054] In some embodiments, the dual-domain feature interaction module is as Figure 5 shown, and the attention feature extraction is respectively performed on the spatial domain feature and the frequency domain feature. The spatial domain feature correspondingly obtains the second query matrix, the first key matrix, and the first value matrix, and the frequency domain feature correspondingly obtains the first query matrix, the second key matrix, and the second value matrix; after multiplying the first query matrix, the first key matrix, and the first value matrix, and then performing the Softmax operation and then the convolution operation, the third feature is obtained; after multiplying the second query matrix, the second key matrix, and the second value matrix and then performing the convolution operation, the fourth feature is obtained; after splicing the third feature and the fourth feature, the obtained spliced feature is respectively subjected to average pooling and max pooling, and the corresponding elements of the features obtained by the two poolings are added and then sequentially subjected to convolution and sigmoid activation operations to obtain the fusion weight; the fusion weight is multiplied by the corresponding elements of the spliced feature to obtain the fusion feature.
[0055] It can be understood that attention feature extraction is respectively performed on the spatial domain features and the frequency domain features. The second query matrix Q2, the first key matrix K1, and the first value matrix V1 are correspondingly obtained for the spatial domain features, and the first query matrix Q1, the second key matrix K2, and the second value matrix V2 are correspondingly obtained for the frequency domain features.
[0056] In a specific embodiment, layer normalization operations are respectively performed on the spatial domain features and the frequency domain features first, and then attention feature extraction is performed. After multiplying the first query matrix Q1, the first key matrix K1, and the first value matrix V1, a Softmax operation is performed, and then a two-dimensional convolution operation is performed to obtain a third feature, as shown in the following formula:
[0057] ;
[0058] where F represents the third feature, C, M, and N respectively represent the number of channels, height, and width of the spatial domain features, and Conv2d represents the two-dimensional convolution operation.
[0059] After multiplying the second query matrix Q2, the second key matrix K1, and the second value matrix V2, a Softmax operation is performed, and then a two-dimensional convolution operation is performed to obtain a fourth feature, as shown in the above formula, which will not be elaborated here.
[0060] After concatenating the third feature and the fourth feature, average pooling and max pooling are respectively performed on the obtained concatenated feature. The corresponding elements of the features obtained by the two poolings are added, and then a one-dimensional convolution and a sigmoid activation operation are sequentially performed to obtain a fusion weight; the fusion weight is multiplied by the corresponding elements of the concatenated feature to obtain a fusion feature.
[0061] In some embodiments, the object detection branch in step S2 includes multiple detection heads, and the number of detection heads is the same as the number of initial features extracted by the feature extraction branch. Each detection head respectively processes the fusion feature of the corresponding size to obtain the category and position corresponding to the remote sensing image, and then combines the categories and all positions corresponding to the fusion features of different sizes to determine the category and position corresponding to the remote sensing image. The position specifically includes the horizontal and vertical coordinates of the center point of the object, the length and width of the object, and the rotation angle of the object. Each detection head includes a category detection sub-module and a position detection sub-module; the fusion features of the corresponding size are respectively input into the category detection sub-module and the position detection sub-module. The category detection sub-module determines the category corresponding to the current fusion feature, and the position detection sub-module determines the position corresponding to the current fusion feature.
[0062] In one embodiment, the number of detection heads is 4, corresponding to features of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 sizes respectively. Each detection head includes a class detection sub-module and a position detection sub-module: In the class detection sub-module, there are 5 cascaded 3×3 two-dimensional convolution operations. After the first four convolution operations, the feature size remains the same as the input feature. After the last convolution operation, the feature size is H×W×K (K is the number of classes); in the position detection sub-module, there are also 5 3×3 two-dimensional convolution operations. After the first four convolution operations, the feature size remains the same as the input feature, but after the last convolution operation, the feature size is H×W×5. If there are detection probabilities obtained by the 4 detection heads within the preset confidence range, the corresponding class and its position are retained, and then redundant target boxes are removed through the NMS (Non-Maximum Suppression) processing operation. At this time, the class in the remaining target boxes is the class corresponding to the remote sensing image, and the position where the content in the remaining target boxes is located is the position corresponding to the remote sensing image. The setting of the confidence is adjusted according to the actual situation and is set to 0.25 in one embodiment.
[0063] S3: Use the training set obtained in step S1 to train the remote sensing target detection network in step S2 to obtain a remote sensing target detection model.
[0064] In one embodiment, the SGD optimizer is adopted, and the weight decay is 0.0001; the initial learning rate is 0.005, which decreases by 10% in the 8th training cycle and the 11th training cycle; for the HRSC2016 dataset, Epoch is set to 36 rounds, and batch_size is set to 2; for the DIOR-R dataset, Epoch is set to 12 rounds; batch_size is set to 2; the position loss function is the Smooth L1 loss function, and the class loss function is the cross-entropy loss function.
[0065] S4: Input the remote sensing image to be detected into the remote sensing target detection model in step S3, and output the corresponding detection class and position.
[0066] In one embodiment, map50 (when the intersection over union threshold is 0.5, the proportion of correct detection results when the overlap degree between the detection box and the ground truth box reaches a certain level, that is, the accuracy of the model detection results) is used as the evaluation index. The target detection accuracy of the remote sensing images to be detected in the HRSC2016 dataset is 90.7%, and the target detection accuracy of the remote sensing images to be detected in the DIOR-R dataset is 65.7%. It can be seen that the target detection method provided by the present invention has good accuracy.
[0067] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in the disclosure of the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.
[0068] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A directed object detection method for remote sensing images based on dual-domain feature fusion, characterized in that Including: S1: Obtain a remote sensing image dataset, and preprocess the remote sensing image dataset to obtain a training set; S2: Construct a remote sensing object detection network; wherein, the remote sensing object detection network includes: A feature extraction branch for extracting initial features of different sizes from the input remote sensing image; A dual-domain feature fusion branch for extracting corresponding spatial domain features and frequency domain features from the initial features of different sizes; performing feature fusion on the corresponding spatial domain features and frequency domain features of the same size to obtain corresponding fusion features; the dual-domain feature fusion branch includes: A spatial domain adaptive selection module for extracting corresponding spatial domain features from the initial features; the spatial domain adaptive selection module includes a feature extraction sub-module, a pooling sub-module, a multi-scale extraction sub-module, and a feature output sub-module; performing a convolution operation on the input feature to obtain a preprocessed feature; inputting the preprocessed feature into the feature extraction sub-module to obtain a first feature and a second feature; inputting the first feature and the preprocessed feature into the pooling sub-module to obtain a pooled feature; inputting the pooled feature into the multi-scale extraction sub-module for multi-scale information extraction to obtain multi-scale features; inputting the multi-scale features, the preprocessed feature, the second feature, and the input feature into the feature output sub-module to obtain the spatial domain feature; A frequency domain adaptive selection module for extracting corresponding frequency domain features from the initial features; A dual-domain feature interaction module for performing feature fusion on the corresponding spatial domain features and frequency domain features of the same size to obtain corresponding preliminary fusion features; A feature pyramid module for extracting multi-scale information in the preliminary fusion features to obtain the fusion features; An object detection branch for judging the category and position corresponding to the remote sensing image according to different fusion features; S3: Use the training set obtained in step S1 to train the remote sensing object detection network in step S2 to obtain a remote sensing object detection model; S4: Input the remote sensing image to be detected into the remote sensing object detection model in step S3, and output the corresponding detection category and position.
2. The method for directed object detection of remote sensing images based on dual-domain feature fusion according to claim 1, wherein, In the feature extraction sub-module: Perform convolution operations of different scales on the preprocessed feature at least 2 times, and add the corresponding elements of the features output by the convolution operations to obtain the first feature, and the feature output by the convolution operation with the largest scale is the second feature.
3. The method for directed target detection of remote sensing images based on dual-domain feature fusion according to claim 1, wherein In the pooling sub-module: Concatenate the first feature and the preprocessed feature; perform average pooling and max pooling on the concatenated feature at the same time, and then perform pointwise convolution on the features after the two poolings to obtain the pooled feature.
4. The method for directed object detection of remote sensing images based on dual-domain feature fusion according to claim 1, wherein In the multi-scale extraction sub-module: Separate the channels of the pooled feature, perform convolution operations of different scales on the separated features respectively, and then concatenate the convolved features to obtain the multi-scale features.
5. The method for directed object detection of remote sensing images based on dual-domain feature fusion according to claim 1, wherein In the feature output sub-module: Perform sigmoid activation operation on the multi-scale features, multiply the features after sigmoid activation operation with the corresponding elements of the preprocessed features and the second features respectively, and then add the corresponding elements of the two multiplied features to obtain the spatial domain weight; multiply the spatial domain weight with the corresponding elements of the input features to obtain the spatial domain features.
6. The method for directed target detection of remote sensing images based on dual-domain feature fusion according to claim 1, wherein In the frequency domain adaptive selection module: Perform global feature pooling and fast Fourier transform on the input features respectively; The features obtained by the global feature pooling are subjected to multiple convolution operations and then sigmoid activation operation to obtain the frequency domain weight; multiply the frequency domain weight with the corresponding elements of the static filter to obtain the adaptive filter; The features obtained by the fast Fourier transform are operated by the adaptive filter, and then the obtained features are subjected to inverse fast Fourier transform to obtain the frequency domain features.
7. The method for directed object detection of remote sensing images based on dual-domain feature fusion according to claim 1, characterized in that, In the dual-domain feature interaction module: Perform attention feature extraction on the spatial domain features and the frequency domain features respectively. The spatial domain features correspondingly obtain the second query matrix, the first key matrix and the first value matrix, and the frequency domain features correspondingly obtain the first query matrix, the second key matrix and the second value matrix; Multiply the first query matrix, the first key matrix and the first value matrix, then perform Softmax operation and then convolution operation to obtain the third feature; Multiply the second query matrix, the second key matrix and the second value matrix, then perform Softmax operation and then convolution operation to obtain the fourth feature; After splicing the third feature and the fourth feature, perform average pooling and max pooling on the obtained spliced features respectively. The corresponding elements of the features obtained by the two poolings are added and then subjected to convolution and sigmoid activation operations in sequence to obtain the fusion weight; multiply the fusion weight with the corresponding elements of the spliced features to obtain the fusion features.
8. The method for directed target detection of remote sensing images based on dual-domain feature fusion according to claim 1, wherein The object detection branch in step S2 includes multiple detection heads, and the number of detection heads is the same as the number of initial features extracted by the feature extraction branch. Each detection head processes the fusion features of the corresponding size to obtain the category and position corresponding to the remote sensing image, and then combines the categories and all positions corresponding to the fusion features of different sizes to determine the category and position corresponding to the remote sensing image; Each detection head includes a category detection sub-module and a position detection sub-module; The fusion features of the corresponding size are respectively input into the category detection sub-module and the position detection sub-module. The category detection sub-module determines the category corresponding to the current fusion feature, and the position detection sub-module determines the position corresponding to the current fusion feature.
Citation Information
Patent Citations
Optical remote sensing image target detection system and method based on frequency domain feature mining
CN117593640A
Face forgery detection method and system based on multi-modal collaborative learning
CN118397681A
Method for detecting small target in aerial image of unmanned aerial vehicle
CN118762168A
Cited By
A synthetic image detection method based on dual-domain feature fusion enhancement
CN122597863A