Remote sensing image target detection method based on adaptive sampling and dynamic neural network
By employing adaptive sampling and dynamic neural network methods, a target detection model for remote sensing images is constructed, which solves the problem of randomness in target distribution and orientation scale in remote sensing images, thereby improving detection accuracy and recall.
Patent Information
- Application Number
- CN202511015038.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies fail to effectively handle the randomness of target distribution and orientation scale in remote sensing images, resulting in low detection accuracy and insufficient recall.
An adaptive sampling and dynamic neural network approach is adopted. By selecting adaptive sampling points and interacting with the channel and position information of the dynamic neural network, a remote sensing image target detection model is constructed. The model includes a backbone network, a query initialization module, an adaptive sampling point selection module, an adaptive multi-scale fusion module, and a dynamic neural network. The detection model is optimized to improve the target positioning accuracy.
It significantly improves the accuracy and recall of target detection in remote sensing images, overcomes the detection difficulties caused by the randomness of target distribution and orientation, and achieves higher detection precision and recall.
Smart Images

Figure CN120976780A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, and further relates to remote sensing image target detection technology, specifically a remote sensing image target detection method based on adaptive sampling and dynamic neural networks. It can be used to detect targets such as aircraft and vehicles in remote sensing images taken by satellites. Background Technology
[0002] Remote sensing imagery, a novel type of satellite observation data from the sky to the ground, is obtained by optical sensors observing a specific area from a satellite. Compared to natural terrestrial images, satellite images encompass a wider observation area, thus containing more ground information. They have wide applications in fields such as water quality monitoring, resource surveys, and land area monitoring. Therefore, developing target detection technology for remote sensing images can provide stronger technical support for remote sensing monitoring, enabling more effective applications in agriculture, military, and other fields.
[0003] However, due to the randomness of target distribution, scale, and orientation in remote sensing images, target detection in remote sensing images is more difficult than in natural images. Therefore, proposing a target detection algorithm for remote sensing images is of great research value in order to effectively address the above problems.
[0004] The Beijing Satellite Information Engineering Institute disclosed a remote sensing image target detection method in its patent application "Remote Sensing Image Target Detection Method" (Patent Application No.: CN202310403716.5, Publication No.: CN 116524368 A). This method first uses a convolutional neural network to extract multi-scale features from the image. Guided by a mask, it extracts features from the target foreground region and generates rotated candidate boxes representing areas in the original image where targets may exist. Then, it uses RoIAlign alignment to extract features from the candidate box regions. Finally, these features are fed into a detection head composed of Smooth-L1 regression loss and corner margin classification loss for regression localization and classification, respectively. The drawback of this method is that while it is applicable to remote sensing image target detection, it does not address the randomness of target distribution, which can easily lead to missed detections and limits the recall rate.
[0005] Gao et al. proposed an object detection method in their paper "Adamixer: A Fast-Converging Query-Based Object Detector" (CVPR 2022). This method utilizes learnable content queries and bounding box queries to sample features from the feature map, and then performs classification and regression predictions on the sampled features. Simultaneously, they proposed a network that interacts with the sampled features in terms of channel and location information, further improving detection performance. However, a drawback of this method is that it only demonstrates good performance in natural images. When applied to remote sensing images, the randomness of target distribution and orientation makes it difficult to distinguish foreground from background, resulting in lower detection accuracy due to the inability to correctly locate targets. Summary of the Invention
[0006] The purpose of this invention is to address the above-mentioned shortcomings by proposing a remote sensing image target detection method based on adaptive sampling and dynamic neural networks. This method solves the problem that existing technologies do not fully consider the randomness of target distribution and orientation scale in remote sensing images, resulting in low detection accuracy.
[0007] The approach of this invention is as follows: First, features are extracted from images in the training set. Then, corresponding features are extracted from the feature maps based on predicted location information as content queries to alleviate the problem of random target distribution. To alleviate the problem of random scale and orientation, the predicted location offset from the content query is combined with location angle information to align with the target's orientation and scale, extracting more refined features. Finally, the location information is further updated through a dynamic information interaction network of channel dimensions and location. This invention can effectively model the random distribution and orientation / scale of targets in remote sensing images, thereby significantly improving the accuracy of detection results.
[0008] To achieve the above objectives, the technical solution of the present invention includes the following:
[0009] (1) Select remote sensing images and preprocess them to construct a training sample set and a test sample set; the images in the training sample set are labeled with information about the target of interest, which includes category labeling and location information labeling;
[0010] (2) A remote sensing image target detection model is constructed based on adaptive sampling and dynamic neural network, including a backbone network, a query initialization module, an adaptive sampling point selection module, an adaptive multi-scale fusion module, and a dynamic neural network. The specific implementation is as follows:
[0011] (2.1) A convolutional neural network ResNet-50 is used as the backbone network, which consists of multiple convolutional layers-ReLU activation-normalization layers and residual connections are used; features of each training set image are extracted through the backbone network to obtain a multi-scale feature pyramid.
[0012] (2.2) Construct a query initialization module to generate an initial candidate box without angle information for each pixel of each feature in each layer of the feature pyramid. Then, use one convolutional layer and two convolutional layers to predict the classification score and offset of each initial candidate box. Decode the initial candidate box according to the offset to obtain a refined candidate box. Select the E candidate boxes with the highest scores from the refined candidate boxes according to the classification score as reference boxes, where 700≥E≥300. Then, sample the corresponding feature points in the feature pyramid according to the center point coordinates of the E reference boxes as content query.
[0013] (2.3) Design an adaptive sampling point selection module, which generates N sampling point offsets by passing the content query through a linear layer, obtains the sampling point position of each reference box based on the sampling point offset, and samples at the corresponding position of each layer of feature map to obtain multi-layer sampling features;
[0014] (2.4) Construct an adaptive multi-scale fusion module, generate corresponding weights for the features of different layers in step (5) through content query based on the multilayer perceptron network, and perform weighted fusion of the features of different layers according to the weights to obtain multi-scale fusion features;
[0015] (2.5) The fused features are used to interact with the channel and position information through a known dynamic neural network to obtain channel filtering features and spatial filtering features. The spatial filtering features are then added to the content query in step (2.2) after passing through a linear layer. Finally, the corresponding position offset and category score of the reference box are predicted through a linear layer. The reference box is decoded using the same decoding method as in step (2.2) to obtain the final predicted box.
[0016] (3) Train the remote sensing image target detection model constructed in step (2):
[0017] The training sample set is divided into groups of N images for model training, where N≥1. The training model is optimized by calculating the loss function and backpropagation algorithm to obtain the trained detection model.
[0018] (4) Perform inference on each image in the test sample set individually using the trained detection model, and verify it using the predicted bounding box and category score obtained in step (2.5) to obtain the final detection model;
[0019] (5) Input the remote sensing image to be tested into the final detection model and use the model to obtain the target detection result.
[0020] Compared with the prior art, the present invention has the following advantages:
[0021] First, because the present invention performs adaptive sampling of target location and feature points on the input remote sensing image, it overcomes the problem of poor adaptability to randomly distributed targets in the prior art, which leads to a large number of missed targets. As a result, the present invention improves the ability to capture target location and improves the recall rate of detection.
[0022] Secondly, because the present invention uses a dynamic neural network to exchange information between channels and positions, it overcomes the problem of low positioning quality caused by the insensitivity of existing technologies to the scale and orientation of targets; thus, the present invention can locate targets more accurately and has better detection accuracy. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention;
[0024] Figure 2 These are partial detection results provided in an embodiment of the present invention. Detailed Implementation
[0025] The present invention will now be further described with reference to the accompanying drawings.
[0026] Example 1: Refer to Appendix Figure 1 The present invention proposes a remote sensing image target detection method based on adaptive sampling and dynamic neural networks, which specifically includes the following steps:
[0027] Step 1) Select remote sensing images and preprocess them to construct training sample sets and test sample sets; the images in the training sample set are labeled with information about the target of interest, which includes category labeling and location information labeling; the preprocessing described in this embodiment specifically involves cropping the selected remote sensing images to obtain image blocks of size 1024*1024, and dividing them into two parts, which are used as training samples and test samples respectively, with the preferred division ratio being 7:3.
[0028] Step 2) Construct a remote sensing image target detection model based on adaptive sampling and dynamic neural networks, including a backbone network, a query initialization module, an adaptive sampling point selection module, an adaptive multi-scale fusion module, and a dynamic neural network. The specific implementation is as follows:
[0029] (2.1) A convolutional neural network ResNet-50 is used as the backbone network, which consists of multiple convolutional layers-ReLU activation-normalization layers and residual connections are used; features of each training set image are extracted through the backbone network to obtain a multi-scale feature pyramid.
[0030] (2.2) Construct a query initialization module to generate an initial candidate box without angle information for each pixel of each feature in each layer of the feature pyramid. Then, use one convolutional layer and two convolutional layers to predict the classification score and offset of each initial candidate box. Decode the initial candidate box according to the offset to obtain a refined candidate box. Select the E candidate boxes with the highest scores from the refined candidate boxes according to the classification score as reference boxes, where 700≥E≥300. Then, sample the corresponding feature points in the feature pyramid according to the center point coordinates of the E reference boxes as content queries.
[0031] In this embodiment, the initial candidate box is decoded based on the offset as follows:
[0032] Assuming the initial bounding box coordinates are (x, y, w, h, θ), where (x, y) are the center point coordinates, (w, h) are the length and width, and θ is the angle with an initial value of 0; and the predicted offset is (Δx, Δy, Δw, Δh, Δθ), then the decoding formula for the initial bounding box is:
[0033] gw=w*e Δw
[0034] gh = h * e Δh
[0035] gx=Δx*w*cos(θ)-Δy*h*sin(θ)+x
[0036] hy=Δx*w*sin(θ)+Δy*h*cos(θ)+y
[0037] gθ=θ+Δθ
[0038] The refined candidate box coordinates are (gx, gy, gw, gh, gθ).
[0039] (2.3) Design an adaptive sampling point selection module to generate N sampling point offsets by passing the content query through a linear layer. Obtain the sampling point position of each reference box based on the sampling point offset, and sample at the corresponding position of each layer of feature map to obtain multi-layer sampling features.
[0040] In this embodiment, sampling points are selected in each refinement candidate box based on the sampling point offset in this step, specifically as follows:
[0041] Let the sampling point offset be (Δx, Δy), then the sampling point coordinates (x, y) can be obtained according to the following formula:
[0042] dx = Δx * gw,
[0043] dy = Δy * gh,
[0044]
[0045] Where dx and dy represent the offset of the sampling points after scaling to the original image size.
[0046] (2.4) Construct an adaptive multi-scale fusion module, which uses content query to generate corresponding weights for the features of different layers in step 5) through a multilayer perceptron network, and then performs weighted fusion of the features of different layers according to the weights to obtain multi-scale fused features; the implementation is as follows:
[0047] Suppose the multilayer feature map is F = {F1, F2, ..., F...} k ,…,F l ,}, where the i-th layer features F i The scale size is (H) i W i ), query the content q∈R C The weights W = {W1, W2, ..., W...} of each layer are obtained through a multilayer perceptron. k ,…,W l For each layer of feature map, bilinear interpolation is applied to sample the feature points. Let F′ be the sampled i-th layer feature map. i Its size is (n q (n, C), where C represents the channel dimension, and n... q This represents the number of candidate bounding boxes for each image; the multi-scale fusion features are obtained according to the following formula:
[0048]
[0049] (2.5) The fused features are used to interact with the channel and position information through a known dynamic neural network to obtain channel filtering features and spatial filtering features. The spatial filtering features are then added to the content query in step (2.2) after passing through a linear layer. Finally, the corresponding position offset and category score of the reference box are predicted through a linear layer. The reference box is decoded using the same decoding method as in step (2.2) to obtain the final predicted box.
[0050] The interaction process described above in this embodiment is specifically represented as follows:
[0051] F″=F′·W c ∈(n q (,N,C),
[0052] F ~ =W s ·F″∈(n q ,n o C),
[0053] Among them, W c ∈R c×c For channel interaction weights, This represents the location interaction weight; after the interaction, the final number of sampling points changes from N to n. o ;F″、F ~ These represent the channel filtering characteristics and the spatial filtering characteristics, respectively.
[0054] Step 3) Train the remote sensing image target detection model constructed in Step 2):
[0055] The training sample set is divided into groups of N images for model training, where N≥1. The training model is optimized by calculating the loss function and backpropagation algorithm to obtain the trained detection model.
[0056] The loss function described in this embodiment consists of two parts. The first part of the loss is used to optimize the module and the backbone network, and is calculated using the classification score and offset of the initial candidate box obtained by the query initialization module in step (2.2). The second part of the loss is used to optimize the adaptive sampling point selection module, the adaptive multi-scale fusion module and the dynamic neural network, and to further optimize the query initialization module and the backbone network. The loss is calculated using the corresponding position offset and category score of the reference box predicted in step (2.5).
[0057] The calculation of the first and second parts of the loss mentioned above both include: using the classification scores and category labels corresponding to the E reference boxes to calculate the Focal Loss based on binary cross-entropy, and using the reference boxes and location information labels to calculate the intersection-union ratio loss, adaptive localization loss (RIoU Loss), and mean absolute error loss (L1 Loss).
[0058] Step 4) Perform inference on each image in the test sample set individually using the trained detection model, and verify it using the predicted bounding box and category score obtained in step (2.5) to obtain the final detection model;
[0059] Step 5) Input the remote sensing image to be tested into the final detection model and use the model to obtain the target detection result.
[0060] Example 2: The overall implementation steps of the remote sensing image detection method proposed in this example are the same as in Example 1. The parameter settings are given below, and the implementation process of the present invention is further described in detail with specific examples:
[0061] Step 1. Generate training and testing sample sets;
[0062] Since most remote sensing images have high resolution and cannot be directly sent to the network, the selected remote sensing images are cropped. At least 1900 images are extracted from the selected remote sensing images to form a training set and at least 800 images are extracted to form a test set, and then cropped into image patches of size 1024*1024.
[0063] Step 2. Group the training sample set into groups of 4 images each, and extract features from each group of images using a convolutional neural network to obtain a multi-scale feature pyramid.
[0064] Step 3. Generate an initial candidate box without angle information for each pixel of each feature layer in the feature pyramid, and stitch the multi-scale feature pyramids together in the size dimension to obtain a large feature map containing multi-scale information. At the same time, stitch the candidate boxes together in the same way.
[0065] Step 4. Predict the category and offset using the classification head and regression head respectively, decode the initial bounding box based on the offset to obtain refined candidate boxes, and select the n with the highest classification score. q (500) candidate boxes are used as reference boxes, and feature points at the corresponding positions of the center coordinates of the candidate boxes are sampled in the large feature map as content queries.
[0066] The steps for decoding the initial bounding box into refined candidate bounding boxes are as follows:
[0067] Assuming the initial bounding box coordinates are (x, y, w, h, θ), where (x, y) are the center point coordinates, (w, h) are the length and width, and θ is the angle, initially set to 0, and the predicted offset is (Δx, Δy, Δw, Δh, Δθ), then the decoding formula for the initial bounding box is:
[0068] gw=w*e Δw
[0069] gh = h * e Δh
[0070] gx=Δx*w*cos(θ)-Δy*h*sin(θ)+x
[0071] gy=Δx*w*sin(θ)+Δy*h*cos(θ)+y
[0072] gθ=θ+Δθ
[0073] The refined candidate box coordinates are (gx, gy, gw, gh, gθ).
[0074] Step 5. Generate N sampling point offsets by passing the content query through a linear layer. In each refined candidate box, select N sampling points based on the sampling point offsets, and sample feature points at the corresponding positions in the feature map. The number of feature points is then n. q ×N;
[0075] The steps for selecting the sampling points are as follows:
[0076] For the coordinate vector (gx, gy, gw, gh, gθ) of the candidate box, let the offset of one of the sampling points be (Δx, Δy), then the formula for selecting the sampling point is:
[0077] dx=Δx*gw
[0078] dy=Δy*gh
[0079]
[0080] Step 6. Generate different weights for features from different layers through content query, and fuse the features from different layers according to the weights to obtain a multi-scale fused feature map.
[0081] The detailed steps for fusing the multi-layer features are as follows:
[0082] Suppose the multilayer feature map is F = {F1, F2, ..., F...} k ,…,F l For the i-th layer feature F i Its scale is (H) i W i ), query the content q∈R C The weights W = {W1, W2, ..., W} of each layer are obtained through linear layers. k ,…,W l For each feature map layer, bilinear interpolation is applied to sample feature points based on the sampling points generated in step 5. The sampled feature map layer i is denoted as F′. i Its size is (n q Given a vector group (x, y), where C represents the channel dimension, the feature fusion for each point (x, y) is as follows:
[0083]
[0084] Step 7. The fused features are processed through a dynamic neural network to interact with channel and location information, and finally the corresponding location offset is obtained. The same decoding method as in Step 4 is used to decode the final prediction box, predict its category, and calculate the loss function.
[0085] The dynamic neural network and loss function are as follows:
[0086] For the fusion feature F′ generated in step 6, let its dimension be (n q (,N,C), where n q n represents the number of candidate bounding boxes for each image. p Let N represent the number of sampled feature points mentioned in step 5, and C represent the dimension of its feature channels. The dynamic neural network first generates channel interaction weights W through content querying. c ∈RC×C Interaction weights with location The interaction process is as follows:
[0087] F = F′·W c ∈(n q ,N,N)
[0088] F = W s ·F∈(n q ,n o C)
[0089] After the above formula interaction, the final number of sampling points changes from N to n. o .
[0090] The features F obtained above are processed through a linear layer to predict the classification score and position offset. Based on the position offset, the refined candidate boxes obtained in step 4 are decoded using the same decoding method to obtain the predicted position information. Focal loss is calculated using the classification score and category label, with a weight of 2. Intersection over Union (IoU) loss and L1 loss are calculated using the predicted position information and position label, with weights of 5 and 2, respectively.
[0091] The effects of this invention will be further illustrated below with simulation experiments:
[0092] 1. Simulation experimental conditions:
[0093] The hardware platform for the simulation experiment consisted of an Intel i7-7820X CPU with a clock speed of 3.5GHz and 128GB of RAM. The graphics card used was a GeForce RTX 3090 Ti.
[0094] The software platform for the simulation experiment is: Linux 22.04 operating system and Python 3.8.
[0095] The input images used in the simulation experiment are from the Dota-v1.0 remote sensing dataset, which includes 15 categories such as airplanes, ships, storage tanks, baseball fields, tennis courts, basketball courts, ground runways, ports, bridges, large vehicles, small cars, helicopters, roundabouts, football fields, and swimming pools. The image format is .png.
[0096] Thus obtained Figure 2 The experimental results.
[0097] from Figure 2 As can be seen from the detection results, the method of the present invention can detect the presence of targets in remote sensing images with relatively high accuracy and can identify their categories.
[0098] In summary, this invention enhances the ability to capture and locate targets in remote sensing images through adaptive sampling and dynamic neural networks, thereby improving the accuracy of target detection and effectively alleviating the problems of random target distribution, angle, and orientation in remote sensing images.
[0099] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0100] The above simulation analysis proves the correctness and effectiveness of the method proposed in this invention.
[0101] The parts of this invention not described in detail are common knowledge to those skilled in the art.
[0102] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Obviously, those skilled in the art, after understanding the content and principle of the present invention, may make various modifications and changes in form and detail without departing from the principle and structure of the present invention. However, these modifications and changes based on the concept of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A remote sensing image target detection method based on adaptive sampling and dynamic neural networks, characterized in that, Includes the following steps: (1) Select remote sensing images and preprocess them to construct a training sample set and a test sample set; the images in the training sample set are labeled with information about the target of interest, which includes category labeling and location information labeling; (2) A remote sensing image target detection model is constructed based on adaptive sampling and dynamic neural network, including a backbone network, a query initialization module, an adaptive sampling point selection module, an adaptive multi-scale fusion module, and a dynamic neural network. The specific implementation is as follows: (2.1) A convolutional neural network ResNet-50 is used as the backbone network, which consists of multiple convolutional layers-ReLU activation-normalization layers and residual connections are used; features of each training set image are extracted through the backbone network to obtain a multi-scale feature pyramid. (2.2) Construct a query initialization module to generate an initial candidate box without angle information for each pixel of each feature in each layer of the feature pyramid. Then, use one convolutional layer and two convolutional layers to predict the classification score and offset of each initial candidate box. Decode the initial candidate box according to the offset to obtain a refined candidate box. Select the E candidate boxes with the highest scores from the refined candidate boxes according to the classification score as reference boxes, where 700≥E≥300. Then, sample the corresponding feature points in the feature pyramid according to the center point coordinates of the E reference boxes as content query. (2.3) Design an adaptive sampling point selection module, which generates N sampling point offsets by passing the content query through a linear layer, obtains the sampling point position of each reference box based on the sampling point offset, and samples at the corresponding position of each layer of feature map to obtain multi-layer sampling features; (2.4) Construct an adaptive multi-scale fusion module, generate corresponding weights for the features of different layers in step (5) through content query based on the multilayer perceptron network, and perform weighted fusion of the features of different layers according to the weights to obtain multi-scale fusion features; (2.5) The fused features are used to interact with the channel and position information through a known dynamic neural network to obtain channel filtering features and spatial filtering features. The spatial filtering features are then added to the content query in step (2.2) after passing through a linear layer. Finally, the corresponding position offset and category score of the reference box are predicted through a linear layer. The reference box is decoded using the same decoding method as in step (2.2) to obtain the final predicted box. (3) Train the remote sensing image target detection model constructed in step (2): The training sample set is divided into groups of N images for model training, where N≥1. The training model is optimized by calculating the loss function and backpropagation algorithm to obtain the trained detection model. (4) Perform inference on each image in the test sample set individually using the trained detection model, and verify it using the predicted bounding box and category score obtained in step (2.5) to obtain the final detection model; (5) Input the remote sensing image to be tested into the final detection model and use the model to obtain the target detection result.
2. The remote sensing image target detection method according to claim 1, characterized in that: The preprocessing described in step (1) specifically involves cropping the selected remote sensing image to obtain an image patch of size 1024*1024, and dividing it into two parts, which are used as training samples and test samples respectively.
3. The remote sensing image target detection method according to claim 1, characterized in that: The decoding of the initial candidate box based on the offset in step (2.2) is implemented as follows: Assuming the initial bounding box coordinates are (x, y, w, h, θ), where (x, y) are the center point coordinates, (w, h) are the length and width, and θ is the angle with an initial value of 0; and the predicted offset is (Δx, Δy, Δw, Δh, Δθ), then the decoding formula for the initial bounding box is: w=w*e Δw gh=h*e Δh gx=Δx*w*cos(θ)-Δy*h*sin(θ)+x gy=Δx*w*sin(θ)+Δy*h*cos(θ)+y gθ=θ+Δθ The refined candidate box coordinates are (gx, gy, gw, gh, gθ).
4. The remote sensing image target detection method according to claim 3, characterized in that: In step (2.3), sampling points are selected in each refinement candidate box based on the sampling point offset, specifically as follows: Let the sampling point offset be (Δx, Δy), then the sampling point coordinates (x, y) can be obtained according to the following formula: dx = Δx * gw, dy = Δy * gh, Where dx and dy represent the offset of the sampling points after scaling to the original image size.
5. The remote sensing image target detection method according to claim 1, characterized in that: In step (2.4), features from different layers are fused according to their weights, as follows: Suppose the multilayer feature map is F = {F1, F2, ..., F...} k ,…,F l ,}, where the i-th layer features F i The scale size is (H) i W i ), query the content q∈R C The weights W = {W1, W2, ..., W...} of each layer are obtained through a multilayer perceptron. k ,…,W l For each layer of feature map, bilinear interpolation is applied to sample the feature points. Let F′ be the sampled i-th layer feature map. i Its size is (n q (n, C), where C represents the channel dimension, and n... q This represents the number of candidate bounding boxes for each image; the multi-scale fusion features are obtained according to the following formula:
6. The remote sensing image target detection method according to claim 5, characterized in that: In step (2.5), the fused features are used to interact with channel and location information through a dynamic neural network. The interaction process is as follows: F″=F′·W c ∈(n q ,N,C), F ~ =W s ·F″∈(n q ,n o ,C), Among them, W c ∈R C×C For channel interaction weights, This represents the location interaction weight; after the interaction, the final number of sampling points changes from N to n. o ;F″、F ~ These represent the channel filtering characteristics and the spatial filtering characteristics, respectively.
7. The remote sensing image target detection method according to claim 1, characterized in that: The loss function in step (3) consists of two parts. The first part of the loss is used to optimize the module and the backbone network, and is calculated using the classification scores and offsets of the initial candidate boxes obtained by the query initialization module in step (2.2). The second part of the loss is used to optimize the adaptive sampling point selection module, the adaptive multi-scale fusion module and the dynamic neural network, and to further optimize the query initialization module and the backbone network. The loss is calculated using the corresponding position offsets and category scores of the reference boxes predicted in step (2.5).
8. The remote sensing image target detection method according to claim 7, characterized in that: The calculation of the first part of the loss and the second part of the loss both include: using the classification scores and category labels corresponding to the E reference boxes to calculate the Focal Loss based on binary cross-entropy, and using the reference boxes and location information labels to calculate the intersection-union ratio loss, adaptive localization loss (RIoU Loss), and mean absolute value error loss (L1 Loss).
Citation Information
Patent Citations
Remote sensing image target detection method
CN116524368A
Remote sensing image target detection methods
CN116524368B