Remote sensing target detection method and device based on diffusion model
By introducing variable variance scheduling and common feature interaction modules into the diffusion model, the problem of instability in the diffusion model to noise sensitivity and detection results is solved, and more efficient and accurate rotation object detection is achieved.
Patent Information
- Application Number
- CN202510250559.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-27
AI Technical Summary
The existing diffusion model-based object detection method uses fixed variance scheduling when dealing with object detection tasks of different complexity and diversity, resulting in limited model performance and overly sensitive to the generated random noise, resulting in unstable detection results.
Variable variance scheduling method and designing common feature interaction modules are used to reduce the sensitivity of the diffusion model to random noise, thereby improving the stability of training and prediction.
The stability and efficiency of the model in the rotation object detection task is improved, and more efficient and accurate detection performance is achieved.
Smart Images

Figure CN120219708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to a remote sensing target detection method and device based on a diffusion model. Background Art
[0002] The purpose of target detection is to predict a set of bounding boxes and related class labels of target objects in an image. As a basic visual recognition task, it has become the cornerstone of many related recognition scenarios, such as instance segmentation, pose estimation, action recognition, target tracking, and visual relationship detection. Due to the different appearances, shapes, and poses of various objects, combined with the interference of factors such as illumination and occlusion during imaging, target detection has always been one of the most challenging problems in the field of computer vision. Different target detection schemes have their own application scenarios and great significance in fields such as public safety, military system applications, and civilian commercial software.
[0003] Detection paradigm based on empirical priors: Most detectors complete the detection task by performing box regression and classification on empirically designed candidate objects, such as sliding windows, region proposals, anchor boxes, and Oriented RepPoints. This approach is conceptually intuitive, relatively simple to implement, has relatively high computational efficiency, and has good interpretability. However, it has poor generalization ability for new or unseen object categories. These methods are usually sensitive to the rotation of the target and require the design of multiple detectors for different angles. They rely on manually designed features, which require a large amount of experience and domain knowledge.
[0004] Detection paradigm based on learnable object queries: such as DETR, Sparse r-cnn, etc. Eliminate some manually designed components of empirical priors and establish end-to-end detection. It has better generalization ability for new or unseen object categories. However, this method requires a large amount of labeled data to train the model, is sensitive to data quality and diversity, has a complex model structure, may lead to overfitting and difficult parameter tuning, and has poor interpretability.
[0005] The variance scheduling used in existing target detection methods based on diffusion models is usually fixed. When dealing with target detection tasks of different complexities and diversities, it may not be able to adapt to various situations, resulting in limited model performance. Moreover, it may be too sensitive to the generated random noise, leading to unstable detection results. The original model also repeatedly uses the detection decoder in each sampling step, that is, trains different detection decoder weights according to the noise boxes generated at different times, and then selects the corresponding detection decoder according to the time during inference, which greatly increases the number of model parameters, increases redundant calculations, and reduces the inference efficiency of the model. Summary of the Invention
[0006] In view of the above problems, the present invention proposes a remote sensing target detection method and device based on a diffusion model, which improves the method of applying the existing diffusion model to rotating target detection. By using a variable variance scheduling method and designing a common feature interaction module, the original diffusion model reduces its sensitivity to random noise, thereby improving the stability of training and prediction.
[0007] In a first aspect, the present invention provides a remote sensing target detection method based on a diffusion model, including:
[0008] Step 1: Obtain a remote sensing image dataset and preprocess the remote sensing image dataset;
[0009] Step 2: Train an image encoding module based on the preprocessed remote sensing image dataset; the image encoding module is used to extract multi-scale features of the remote sensing image to obtain a multi-scale feature map;
[0010] Step 3: Construct an initial noise box through a variable variance scheduling method, and train a detection decoding module based on the initial noise box and the multi-scale feature map extracted by the image encoding module; the detection decoding module performs noise prediction on the initial noise box to restore the target box to complete target detection;
[0011] Step 4: Combine the trained image encoding module and detection decoding module to form a remote sensing target detection model, and the remote sensing target detection model is used to perform target detection on remote sensing images.
[0012] Further, the preprocessing in Step 1 includes: cropping, randomly cropping, and randomly horizontally flipping the image.
[0013] Further, the image encoding module uses a ResNet101 convolutional neural network and a Feature Pyramid Network (FPN) to extract high-level features of the remote sensing image; wherein, the ResNet101 convolutional neural network is used to extract a multi-scale feature map of the remote sensing image, including a plurality of residual blocks; each residual block contains a 1×1 convolutional layer and a 3×3 convolutional layer; the Feature Pyramid Network (FPN) fuses the multi-scale feature map through upsampling and lateral connections.
[0014] Further, constructing an initial noise box through a variable variance scheduling method specifically includes:
[0015] Adding Gaussian noise to the true bounding box by calculating variable variance scheduling parameters to generate an initial noise box; wherein, the true bounding box is a box pre-annotated in the remote sensing image, and the variable variance scheduling parameters are calculated by the following formula:
[0016] β 1,2,...,T= sigmoid(linspace(start, end, T)) * w + b (1)
[0017] where β 1,2,...,T represents the variance scheduling parameter, start represents the starting point, end represents the ending point, T represents the time step, w represents the weight, b represents the offset, linspace represents the linspace function, and sigmoid represents the sigmoid function;
[0018] The process of adding Gaussian noise to the true bounding box is achieved through the following formula:
[0019]
[0020] where z0 represents the true bounding box, and z t represents the bounding box after diffusion of the true bounding box for t time steps, represents the cumulative variance scheduling parameter at time step t, and I represents the identity matrix, where:
[0021]
[0022] where β s represents the variance scheduling parameter at time step s.
[0023] Furthermore, the detection decoding module includes RROI Align, a feature interaction module, a noise prediction module, and a box update module;
[0024] Correspondingly, the detection decoding module is trained using the initial noise box and the multi-scale features extracted by the image encoding module, specifically including:
[0025] Step A: Input the initial noise box and the multi-scale feature maps extracted by the image encoding module into the RROI Align for feature extraction to obtain the ROI feature map;
[0026] Step B: Perform preliminary detection on the ROI feature map through the feature interaction module to obtain the initial prediction box;
[0027] Step C: Input the initial prediction box into the noise prediction module to obtain the noise prediction at the current moment;
[0028] Step D: Iteratively denoise the initial prediction box based on the noise prediction at the current moment and the number of noise steps at the current moment to obtain the target prediction box;
[0029] Step E: Calculate the matching cost between the target prediction box and the true target box. Take the target prediction box with a low matching cost as the positive sample, and the rest as negative samples. Replace the negative samples with random boxes, and keep the positive samples as the input of the RROIAlign. Repeat steps A to E N times.
[0030] Further, the feature interaction module includes a multi-head self-attention module, a feature fusion module, a classification module, and a rotated box regression module; where
[0031] The multi-head self-attention module captures the dependencies between different positions in the ROI feature map through multiple self-attention heads to generate attention weights. The multi-head self-attention module includes an average pooling layer, a fully connected layer, a ReLU layer, a fully connected layer, and a Softmax layer connected in sequence;
[0032] The feature fusion module is used to generate convolution kernels according to the attention weights respectively, sum the convolution kernels to generate a dynamic convolution, and align and fuse the ROI feature map based on the dynamic convolution to obtain a fused feature map;
[0033] The classification module is used to output the feature map after feature interaction after batch normalization and a linear layer of the fused feature map, and use the softmax function to output the feature map through a fully connected layer again, and convert it into a class probability distribution;
[0034] The rotated box regression module is used to extract regression features from the fused feature map, input the extracted features into a regression layer, and generate the parameters of the initial prediction box.
[0035] Further, the iterative denoising process in step D is as follows:
[0036] Based on the noise prediction at the current moment and the noise step t at the current moment, calculate the noise prediction with a time step of t - 1. Input the noise prediction with a time step of t - 1 and the initial prediction box into the box update model for denoising to obtain the prediction box with a time step of t - 1;
[0037] Based on the noise prediction with a time step of t - 1, calculate the noise prediction with a time step of t - 2. Input the noise prediction with a time step of t - 2 and the prediction box with a time step of t - 1 into the box update module for denoising to obtain the prediction box with a time step of t - 1;
[0038] Repeat calculating the prediction noise at the previous moment and input it into the box update model until the prediction box with a step of 0 is obtained, completing the iterative denoising process;
[0039] The calculation formula for the noise prediction with a time step of t - 1 is as follows:
[0040]
[0041] wherein, β t represents the noise variance at time step t, ∈~N(0,I) represents a random noise of the standard normal distribution, and z t represents the noise data at time step t, and z t-1 represents the noise data at time step t-1. represents the cumulative variance scheduling parameter at time step t, μ represents the parameter for controlling the degree of noise addition, and ∈ θ (z t ,t) represents the noise prediction module; σ t represents the standard deviation at time step t, and σ t The calculation formula is as follows:
[0042]
[0043] wherein, represents the cumulative variance scheduling parameter at time step t-1; the variance scheduling parameter is calculated by formula (4):
[0044]
[0045] wherein, β s represents the variance scheduling parameter at time step s.
[0046] Furthermore, calculate the matching cost between the target prediction box and the true target box, and the calculation formula is as follows:
[0047] C=λ cls *C cls +λ smooth L1 *C smooth L1 +λ giou *C giou (11)
[0048] wherein, C cls represents the focal loss between the prediction and the true class label, C smooth L1 represents the smooth L1 loss, C giou represents the GIoU loss, and λ cls 、λ smooth L1 and λ giou represent the weights corresponding to the losses.
[0049] Furthermore, the calculation matching cost formula (11) in step E is used as a multi-task loss function to optimize the remote sensing target detection model.
[0050] In a second aspect, the present invention provides a remote sensing target detection device based on a diffusion model, including:
[0051] A dataset acquisition and preprocessing module, which is used to acquire a remote sensing image dataset and preprocess the remote sensing image dataset;
[0052] An image encoding module, which is trained based on the preprocessed remote sensing image dataset; the image encoding module is used to extract multi-scale features of the remote sensing image to obtain a multi-scale feature map;
[0053] A detection decoding module, which is used to construct an initial noise box by a variable variance scheduling method, and train the detection decoding module based on the initial noise box and the multi-scale feature map extracted by the image encoding module; the detection decoding module performs noise prediction on the initial noise box to restore the target box to complete target detection;
[0054] A remote sensing target detection model, which is used to form a remote sensing target detection model by combining the trained image encoding module and detection decoding module, and the remote sensing target detection model is used to perform target detection on remote sensing images.
[0055] The beneficial effects of the present invention are as follows:
[0056] Based on the existing method of applying the diffusion model to the detection of rotating targets, the present invention uses a variable variance scheduling method to constrain the generation of noise to a certain extent according to the sample characteristics, avoiding redundant operations caused by complete randomness, making the original diffusion model less sensitive to random noise, thereby improving the stability of training and prediction. A common feature interaction module is also designed to avoid redundant iterations of the original detection decoder. The common feature interaction module is only used once in the initial stage to generate a set of initial prediction boxes, and then these prediction boxes are used as the starting points for sampling. The subsequent sampling steps are completely completed by the inverse Markov chain process, and no longer rely on the repeated use of the traditional detection decoder. The optimized method realizes higher stability and more efficient and accurate detection performance for the detection of rotating targets. Description of the Drawings
[0057] Figure 1 It is a schematic flowchart of a remote sensing target detection method based on a diffusion model provided by an embodiment of the present invention;
[0058] Figure 2 It is a schematic structural diagram of a remote sensing target detection model provided by an embodiment of the present invention;
[0059] Figure 3 It is a schematic structural diagram of a detection decoding module provided by an embodiment of the present invention. Detailed Embodiments
[0060] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0061] As Figure 1 shown, an embodiment of the present invention provides a remote sensing target detection method based on a diffusion model, including:
[0062] Step 1: Obtain a remote sensing image dataset and preprocess the remote sensing image dataset.
[0063] Specifically, the preprocessing includes: cropping, randomly cropping, and randomly horizontally flipping the image. In the embodiment of the present invention, the size of the remote sensing image is cropped to 600*384; the purpose of randomly cropping and randomly horizontally flipping the image is to enhance the training data.
[0064] Step 2: Train an image encoding module based on the preprocessed remote sensing image dataset; the image encoding module is used to extract multi-scale features of the remote sensing image to obtain a multi-scale feature map.
[0065] Step 3: Construct an initial noise box through a variable variance scheduling method, and train a detection decoding module based on the initial noise box and the multi-scale feature map extracted by the image encoding module; the detection decoding module performs noise prediction on the initial noise box to restore the target box to complete target detection.
[0066] Specifically, constructing an initial noise box through a variable variance scheduling method specifically includes: adding Gaussian noise to the true bounding box by calculating the variable variance scheduling parameter to generate an initial noise box; where the true bounding box is a box pre-annotated in the remote sensing image, and the variable variance scheduling parameter is calculated by the following formula:
[0067] β 1,2,...,T = sigmoid(linspace(start, end, T)) * w + b (1)
[0068] where β 1,2,...,TDenotes the variance scheduling parameter, which is used to control the degree of noise addition. start represents the starting point, end represents the ending point, T represents the time step, w represents the weight, b represents the offset, linspace represents the linspace function, and linspace(start, end, T) generates T equally spaced values from start to end. sigmoid represents the sigmoid function, which is used to map the linear space to the interval (0, 1) and is calculated by the following formula:
[0069]
[0070] The process of adding Gaussian noise to the true bounding box is implemented by the following formula:
[0071]
[0072] Among them, z0 represents the true bounding box, and z t represents the bounding box after the true bounding box is diffused by the time step t, represents the cumulative variance scheduling parameter at the time step t, and I represents the identity matrix, where:
[0073]
[0074] Among them, β s represents the variance scheduling parameter at the time step s.
[0075] It can be understood that by introducing four adjustable parameters start, end, ω, and b, the present invention enables the model to better adapt to different data distributions and task requirements, thereby optimizing the variance scheduling of the diffusion model in the object detection task, reducing the model training time, and improving the training efficiency.
[0076] Step 4: Combine the trained image encoding module and the detection decoding module to form a remote sensing object detection model, which is used to detect objects in remote sensing images.
[0077] The embodiment of the present invention uses the variable variance scheduling method, which to a certain extent restricts the generation of noise according to the sample characteristics, avoids the redundant operations caused by complete randomness, makes the original diffusion model less sensitive to random noise, and thus improves the stability of training and prediction.
[0078] Based on the above embodiments, the embodiment of the present invention provides the structure of the image encoding module, as Figure 2As shown in the figure, the image encoding module uses the ResNet101 convolutional neural network and the Feature Pyramid Network (FPN) to extract high-level features of remote sensing images. Among them, the ResNet101 convolutional neural network is used to extract multi-scale features of remote sensing images, including multiple residual blocks, and each residual block contains a 1×1 convolutional layer and a 3×3 convolutional layer. The Feature Pyramid Network (FPN) fuses the multi-scale feature maps through upsampling and lateral connections.
[0079] The processing process of the residual block in the ResNet101 convolutional neural network is as follows: a 1×1 convolutional layer is used for dimensionality reduction and increasing non-linearity, then a 3×3 convolutional layer is used for feature extraction, and finally a 1×1 convolutional layer is used to restore the dimension. This structure enables the network to learn deeper feature representations and avoids the problem of gradient disappearance. ResNet outputs four different scales of feature maps, corresponding to the outputs of its different residual blocks: res2, res3, res4, and res5. The number of channels of these feature maps is 64, 128, 256, and 512 respectively. Then, the Feature Pyramid Network fuses the rich semantic information of the high layer with the rich location information of the low layer. Based on ResNet101, the role of the Feature Pyramid Network (FPN) is to fuse feature maps of different scales to capture multi-scale object information. FPN starts from the output feature maps res2, res3, res4, and res5 of ResNet101, and gradually constructs a feature pyramid of different scales through upsampling and lateral connections. That is, the res5 feature map is upsampled and added to the res4 feature map to obtain the feature map of the p5 layer; p5 is then upsampled and added to res3 to obtain the feature map of the p4 layer; and so on. Finally, the feature maps of the p3, p2, and p1 layers are obtained. These feature maps have the same number of channels, 256, covering multi-scale information from the low layer to the high layer.
[0080] Based on the above embodiments, the embodiments of the present invention provide the structure of the detection and decoding module, as Figure 3 shown. The detection and decoding module includes RROI Align, a feature interaction module, a noise prediction module, and a box update module.
[0081] Specifically, the feature interaction module includes a multi-head self-attention module, a feature fusion module, a classification module, and a rotated box regression module.
[0082] The multi-head self-attention module captures the dependency relationships between different positions in the ROI feature map through multiple self-attention heads and generates attention weights. Among them, the multi-head self-attention module includes an average pooling layer, a fully connected layer, a ReLU layer, a fully connected layer, and a Softmax layer connected in sequence.
[0083] Correspondingly, the processing flow of the multi-head self-attention module is as follows: The RoI features are used as the input to the self-attention calculation module. First, they are processed by average pooling to reduce the spatial dimension and extract important features. Then, the pooled features are linearly transformed through a fully connected layer, and subsequently, the ReLU activation function is applied to introduce non-linearity:
[0084] ReLU(x) = max(0, x) (5)
[0085] If the input x is greater than 0, the output is x. If the input x is less than or equal to 0, the output is 0.
[0086] Then, through the Softmax function:
[0087]
[0088] where, z i is the i-th element in the vector z, representing the unnormalized attention score of a single feature or input. is the exponential function of z i used to convert the attention score to a positive number because the Softmax function requires positive inputs. is the sum of the exponential operations of all attention scores, used for normalization to ensure that the sum of all attention weights is 1. N is the total number of elements in the vector z. In this way, the Softmax function helps the model allocate different attentions among different features or inputs, enabling the model to focus more on the information that is more important for the current task.
[0089] The feature fusion module is used to generate convolution kernels according to the attention weights respectively, sum the convolution kernels to generate a dynamic convolution, and align and fuse the RoI feature map based on the dynamic convolution to obtain the fused feature map.
[0090] The processing process of the feature fusion module is as follows: The input RoI features are subjected to a convolution operation to capture the interaction information between features, obtaining the final feature representation for subsequent classification and regression tasks:
[0091] utput = Dynamic Conv(RoI features, Dynamic Kernel) (7)
[0092] where Dynamic Kernel is generated by weighting different convolution kernels conv1, conv2,..., conv k with the weight parameters π1, π2,..., π k obtained from the attention calculation, used to reflect the different importance of each convolution kernel:
[0093]
[0094] A classification module, which is used to output the feature map after feature interaction by passing the fused feature map through batch normalization and a linear layer, and then output the feature map through a fully connected layer using the softmax function, and convert it into a class probability distribution;
[0095] A rotated bounding box regression module, which is used to extract regression features from the fused feature map and input the extracted features into a regression layer to generate the parameters of the initial prediction box. Among them, the parameters are [x, y, w, h, θ], where (x, y) represents the coordinates of the initial prediction box, w and h respectively represent the width and height of the initial prediction box, and θ represents the rotation angle of the initial prediction box.
[0096] Correspondingly, the detection decoding module is trained using the initial noise box and the multi-scale features extracted by the image encoding module, specifically including:
[0097] Step A: Input the initial noise box and the multi-scale feature map extracted by the image encoding module into RROI Align for feature extraction to obtain an ROI feature map;
[0098] Step B: Perform preliminary detection on the ROI feature map through a feature interaction module to obtain an initial prediction box;
[0099] Step C: Input the initial prediction box into a noise prediction module to obtain the noise prediction at the current moment;
[0100] Step D: Iteratively denoise the initial prediction box based on the noise prediction at the current moment and the noise step at the current moment to obtain a target prediction box;
[0101] Specifically, the iterative denoising process in Step D is as follows:
[0102] Calculate the noise prediction with a time step of t - 1 based on the noise prediction at the current moment and the noise step t at the current moment, input the noise prediction with a time step of t - 1 and the initial prediction box into a box update model for denoising to obtain a prediction box with a time step of t - 1;
[0103] Calculate the noise prediction with a time step of t - 2 based on the noise prediction with a time step of t - 1, input the noise prediction with a time step of t - 2 and the prediction box with a time step of t - 1 into a box update module for denoising to obtain a prediction box with a time step of t - 1;
[0104] Repeat calculating the prediction noise of the previous moment and inputting it into the box update model until a prediction box with a step of 0 is obtained, completing the iterative denoising process.
[0105] The calculation formula for the noise prediction with a time step of t - 1 is as follows:
[0106]
[0107] Among them, β t represents the noise variance at time step t, ∈~N(0,I) represents a random noise of the standard normal distribution, and z t represents the noise data at time step t, and z t-1 represents the noise data at time step t-1. represents the cumulative variance scheduling parameter at time step t, μ represents the parameter controlling the degree of noise addition, and ∈ θ (z t ,t) represents the noise prediction module; σ t represents the standard deviation at time step t, and σ t The calculation formula is as follows:
[0108]
[0109] Among them, represents the cumulative variance scheduling parameter at time step t-1; the variance scheduling parameter is calculated by formula (4).
[0110] Furthermore, the sampling result z t-1 can be stored and compared with the previous sampling results. This comparison process can be used to evaluate the progress of the sampling process to ensure that while the model gradually reduces noise, the target box information is also gradually restored. When the predetermined number of sampling steps (for example, t = 0) is reached or other termination conditions are met, the entire process ends. At this time, the obtained sampling result z0 is the target box information predicted by the model. The final sampling result undergoes non-maximum suppression to generate the final object detection result.
[0111] Step E: Calculate the matching cost between the predicted target box and the true target box, use the target prediction box with a low matching cost as the positive sample, and the rest as negative samples; replace the negative samples with random boxes, retain the positive samples as the input of the RROI Align, and repeat steps A to E N times. In the embodiment of the present invention, N = 6.
[0112] Specifically, when training the model, calculate the classification loss, Smooth L1 loss, and GIoU loss between the predicted bounding box and the target bounding box, and use the dynamic K-value matching algorithm to find the best match; calculate the matching cost between the predicted target box and the true target box, and calculate the matching relationship through the Hungarian Matcher Dynamic K matcher:
[0113] C = λ cls *C cls +λ smooth L1 *C smooth L1 +λgiou *C giou (11)
[0114] Among them, C cls is the focal loss between the predicted and true class labels, C smooth L1 represents the smooth L1 loss, C giou represents the IoU loss, λ cls 、λ smooth L1 and λ giou represent the weights corresponding to the losses. The optimal transport method is used to assign multiple predictions to each ground truth. Specifically, for each ground truth, the top k predictions with the lowest matching cost are selected as its positive samples, and the other predictions are used as negative samples.
[0115] An embodiment of the present invention designs a common feature interaction module to avoid the redundant iteration of the original detection decoder. The common feature interaction module is only used once in the initial stage to generate a set of initial prediction boxes, and then these prediction boxes are used as the starting points for sampling. The subsequent sampling steps are completely completed by the inverse Markov chain process and no longer rely on the repeated use of the traditional detection decoder. It can achieve higher stability, more efficient and accurate detection performance for rotating object detection.
[0116] Based on the above embodiment, further, the calculation matching cost formula (11) in step E is used as a multi-task loss function to optimize the remote sensing object detection model.
[0117] Further, the following loss function is used for training the noise prediction module:
[0118] L train =||∈ - ∈ θ (z t , t)|| 2 (12)
[0119] Among them, ∈ represents the true noise, ∈ θ (z t , t) represents the predicted noise, L train represents the training loss, which is used to measure the difference between the predicted noise and the true noise. By minimizing this loss function, we can better learn how to predict the noise, thereby improving the quality of the sampling results.
[0120] An embodiment of the present invention also provides a remote sensing object detection device based on a diffusion model, including:
[0121] A dataset acquisition and preprocessing module, configured to acquire a remote sensing image dataset and preprocess the remote sensing image dataset;
[0122] The image encoding module is trained based on the pre - processed remote sensing image dataset; the image encoding module is used to extract multi - scale features of the remote sensing image to obtain multi - scale feature maps;
[0123] The detection decoding module is used to construct initial noise boxes through a variable variance scheduling method, and is trained based on the initial noise boxes and the multi - scale feature maps extracted by the image encoding module; the detection decoding module performs noise prediction on the initial noise boxes to recover the target boxes to complete target detection;
[0124] The remote sensing target detection model is used to form a remote sensing target detection model by combining the trained image encoding module and detection decoding module, and the remote sensing target detection model is used to perform target detection on remote sensing images.
[0125] To verify the effectiveness of the method provided in the embodiments of the present invention, the following experiments are carried out:
[0126] Evaluation metrics:
[0127] On the HRSC2016 high - resolution ship dataset, the PASCAL VOC 2007 and VOC 2012 metrics are used for comparison with other methods.
[0128] Experimental settings:
[0129] (1) 450K iterations are adopted. At 350K and 420K iterations, the learning rate is divided by 10.
[0130] (2) The detection decoder iteratively optimizes each prediction from Gaussian random boxes. The top 100 score predictions are selected. The NMS is used to process the prediction results of each sampling step together to obtain the final prediction results.
[0131] Some existing aerial remote sensing target detection methods (including CAL, CAMC, DCR - ReID) are compared with the method we proposed in terms of performance on the HRSC2016 dataset.
[0132] Table 1 Model recognition performance experiments under cross - dataset settings
[0133]
[0134]
[0135] The above experimental results confirm the effectiveness and advancement of the proposed improved detection method based on the diffusion model in the task of aerial remote sensing image target detection.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A remote sensing target detection method based on a diffusion model, characterized in that: include: Step 1: Acquire a remote sensing image dataset and preprocess the remote sensing image dataset; Step 2: Train the image encoding module based on the preprocessed remote sensing image dataset; The image encoding module is used to extract multi-scale features of remote sensing images to obtain a multi-scale feature map; Step 3: constructing an initial noise frame by a variable variance scheduling method, and training a detection decoding module based on the initial noise frame and a multi-scale feature map extracted by the image encoding module; The detection decoding module performs noise prediction on the initial noise frame and restores the target frame to complete target detection; Step 4: The trained image encoding module and detection decoding module are combined into a remote sensing target detection model, and the remote sensing target detection model is used to perform target detection on remote sensing images.
2. The remote sensing target detection method based on diffusion model according to claim 1, characterized in that: The preprocessing in step 1 includes: cropping, random cropping and random horizontal flipping of the image.
3. The remote sensing target detection method based on diffusion model according to claim 1, characterized in that: The image encoding module uses ResNet101 convolutional neural network and feature pyramid network FPN to extract high-level features of remote sensing images; wherein the ResNet101 convolutional neural network is used to extract multi-scale feature maps of remote sensing images, including multiple residual blocks; each residual block contains a 1×1 convolutional layer and a 3×3 convolutional layer; the feature pyramid network FPN fuses the multi-scale feature maps through upsampling and side connections.
4. The remote sensing target detection method based on diffusion model according to claim 1, characterized in that: The initial noise frame is constructed by a variable variance scheduling method, which includes: Gaussian noise is added to the real bounding box by calculating the variable variance scheduling parameter to generate an initial noise box; wherein the real bounding box is a pre-marked box in the remote sensing image, and the variable variance scheduling parameter is calculated by the following formula: β 1,2,...,T =sigmoid(linspace(start,end,T))*w+b (1) Among them, β 1,2,...,T represents variance scheduling parameters, start represents the starting point, end represents the ending point, T represents the time step, w represents the weight, b represents the offset, linspace represents the linspace function, and sigmoid represents the sigmoid function; The process of adding Gaussian noise to the true bounding box is implemented by the following formula: Among them, z0 represents the real bounding box, z t represents the bounding box of the real bounding box after diffusion of time step t, represents the cumulative variance scheduling parameter with a time step of t, and I represents the identity matrix, where: Among them, β s represents the variance scheduling parameter with time step s.
5. The remote sensing target detection method based on diffusion model according to claim 1, characterized in that: The detection and decoding module includes RROI Align, feature interaction module, noise prediction module and frame update module; Correspondingly, the detection and decoding module is trained using the initial noise frame and the multi-scale features extracted by the image encoding module, specifically including: Step A: inputting the initial noise frame and the multi-scale feature map extracted by the image encoding module into the RROIAlign to perform feature extraction to obtain an ROI feature map; Step B: performing preliminary detection on the ROI feature map through the feature interaction module to obtain an initial prediction frame; Step C: inputting the initial prediction frame into the noise prediction module to obtain the noise prediction at the current moment; Step D: iteratively denoising the initial prediction frame based on the noise prediction at the current moment and the number of noise steps at the current moment to obtain a target prediction frame; Step E: Calculate the matching cost between the target prediction box and the real target box, take the target prediction box with low matching cost as the positive sample, and the rest as negative samples; replace the negative samples with random boxes, retain the positive samples as the input of RROIAlign, and repeat steps A to E N times.
6. The remote sensing target detection method based on diffusion model according to claim 5, characterized in that: The feature interaction module includes a multi-head self-attention module, a feature fusion module, a classification module and a rotating frame regression module; wherein, The multi-head self-attention module captures the dependency between different positions in the ROI feature map through multiple self-attention heads and generates attention weights; the multi-head self-attention module includes an average pooling layer, a fully connected layer, a ReLU layer, a fully connected layer and a Softmax layer connected in sequence; The feature fusion module is used to generate convolution kernels according to the attention weights, sum the convolution kernels to generate dynamic convolution, align and fuse the ROI feature maps based on the dynamic convolution, and obtain a fused feature map; The classification module is used to output the fused feature map as a feature map after feature interaction after batch normalization and linear layer, and use a softmax function to output the feature map through a fully connected layer to convert it into a category probability distribution; The rotating frame regression module is used to extract regression features from the fused feature map, input the extracted features into the regression layer, and generate parameters of the initial prediction frame.
7. The remote sensing target detection method based on diffusion model according to claim 5, characterized in that: The iterative denoising process in step D is as follows: Calculating a noise prediction with a time step of t-1 based on the noise prediction at the current moment and the number of noise steps t at the current moment, inputting the noise prediction with a time step of t-1 and the initial prediction frame into the frame update model for denoising, and obtaining a prediction frame with a time step of t-1; Calculating a noise prediction with a time step of t-2 based on a noise prediction with a time step of t-1, inputting the noise prediction with a time step of t-2 and the prediction box with a time step of t-1 into the box update module for denoising, and obtaining a prediction box with a time step of t-1; Repeatedly calculate the prediction noise of the previous moment and input it into the box to update the model until a prediction box with a step length of 0 is obtained, completing the iterative denoising process; The noise prediction calculation formula with a time step of t-1 is as follows: Among them, β t represents the noise variance with a time step of t, ∈~N(0,I) represents the random noise of the standard normal distribution, z t represents the noise data with time step t, z t-1 represents the noise data with time step t-1, represents the cumulative variance scheduling parameter with a time step of t, μ represents the parameter controlling the degree of noise addition, ∈ θ (z t ,t) represents the noise prediction module; σ t represents the standard deviation of time step t, σ t The calculation formula is as follows: in, represents the cumulative variance scheduling parameter with a time step of t-1; the variance scheduling parameter is calculated by formula (4): Among them, β s Represents the variance scheduling parameter with time step s.
8. The remote sensing target detection method based on diffusion model according to claim 5, characterized in that: The matching cost between the target prediction box and the real target box is calculated using the following formula: C=λ cls *C cls +λ smooth L1 *C smooth L1 +λ giou *C giou (11) Among them, C cls represents the focal loss between the predicted and true class labels, C smooth L1 represents the smooth L1 loss, C giou represents the GIoU loss, λ cls , smooth L1 and λ giou Represents the weight corresponding to the loss.
9. The remote sensing target detection method based on diffusion model according to claim 8, characterized in that: The calculation matching cost formula (11) in the step E is used as a multi-task loss function to optimize the remote sensing target detection model.
10. A remote sensing target detection device based on a diffusion model, characterized in that: include: A data set acquisition and preprocessing module, used to acquire a remote sensing image data set and preprocess the remote sensing image data set; An image coding module is trained based on the preprocessed remote sensing image dataset; The image encoding module is used to extract multi-scale features of remote sensing images to obtain a multi-scale feature map; A detection decoding module, used for constructing an initial noise frame by a variable variance scheduling method, and training the detection decoding module based on the initial noise frame and the multi-scale feature map extracted by the image encoding module; The detection decoding module performs noise prediction on the initial noise frame and restores the target frame to complete target detection; The remote sensing target detection model is used to form a remote sensing target detection model by combining a trained image encoding module and a detection decoding module. The remote sensing target detection model is used to perform target detection on remote sensing images.
Citation Information
Cited By
Diffusion model-based three-dimensional point cloud target detection method and device
CN116863426A
Three-dimensional point cloud target detection method and device based on diffusion model
CN116863426B
Video target detection method and system based on multi-scale perception diffusion
CN121190746A