Remote sensing image target detection method based on adaptive downsampling and scale enhancement detection head

By adopting adaptive downsampling and scale-enhanced detection head methods in remote sensing image object detection, the problems of low detection accuracy of small objects and large calculation amount of detection head are solved, and efficient small objects detection and fast inference are achieved.

CN120032249AActive Publication Date: 2025-05-23ICLOUDSHIELD SECURITY TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510099535.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art is prone to loss of feature information when detecting small objects in remote sensing images, resulting in low detection accuracy and large parameter amount and calculation amount of detection head, resulting in long training time and long inference time.

Method used

The remote sensing image object detection method based on adaptive downsampling and scale-enhanced detection head is adopted. The dynamic feature extraction module flexibly adjusts the receptive field. The adaptive downsampling module emphasizes the importance of different regions in the feature map, and improves the sensitivity to small targets through scale-enhanced detection head.

Benefits of technology

The feature information of small targets in remote sensing images is effectively retained, the detection accuracy is improved, the parameter amount and calculation amount of the model are reduced, the training time is shortened, and the inference speed is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032249A_ABST
    Figure CN120032249A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image target detection method based on adaptive downsampling and a scale enhancement detection head. The method comprises the following steps: acquiring a remote sensing image data set, and preprocessing; building a remote sensing image target detection model, and adjusting a receptive field of a feature extraction network by using a dynamic feature extraction module in a feature extraction stage to obtain surrounding information of a small target; in the feature fusion stage, a self-adaptive down-sampling module is adopted, and key features of small targets are reserved; in the prediction stage, a scale enhancement detection head is used for providing optimal feature representation for classification and regression tasks through shared convolution, the detection precision of a small target is improved, and the parameter quantity is reduced; constructing a model loss function; iteratively training the model by using the training set until the model converges; and inputting the images of the test set into the optimal detection model to obtain a detection result. According to the method, the detection precision of the tiny target of the remote sensing image is effectively improved, and the parameter quantity of the model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a remote sensing image target detection method based on adaptive downsampling and scale enhancement detection head, and belongs to the field of computer vision target detection. Background Art

[0002] In the field of computer vision, downsampling operations play a key role. Max pooling is one of the commonly used downsampling methods. It slides a predefined window of a fixed size in the feature map and selects the maximum value in the window as the output. Strided convolution is also a common downsampling method. During the convolution process, a step size greater than 1 is used to change the sliding mode of the convolution kernel. However, for small targets in remote sensing images, since these targets are usually small in size and have fewer features, key feature information is easily lost during the downsampling process, which affects the detection performance. To solve this problem, many researchers are committed to optimizing the downsampling operation to retain more texture and detail information in the feature map, thereby improving the detection accuracy of small targets. Lu et al. proposed a plug-and-play robust feature downsampling method. The method combines three different downsampling operations: convolution, slicing, and max pooling, to extract multiple feature maps to create a feature map containing a complementary feature set.

[0003] In the field of object detection, the design of the detection head is key. Many researchers have improved the detection head to improve detection accuracy. For example, Wu et al. proposed a dual-head method, which includes a fully connected head focused on classification and a convolutional head for bounding box regression. Li et al. proposed an attention-based decoupling that allows classification and regression tasks to focus on the required features. The hybrid attention module assigns different weights to classification and regression in different branches, thereby effectively improving detection performance. Zhang et al. combined an asymmetric decoupled lightweight detector head designed with deep separable convolutions to achieve a significant reduction in computational complexity and a significant improvement in inference speed.

[0004] As the convolutional neural network deepens, the receptive field can theoretically increase linearly. However, the effective receptive field is only a part of the theoretical receptive field. To this end, many researchers design convolutional neural networks with large receptive fields. For example, Azad et al. introduced the concept of deformable large kernel attention, using large kernel convolution to fully understand the context, and benefiting from the deformable convolution, the sampling network can be flexibly distorted, so that the model can adapt to different data patterns.

[0005] The above methods optimize the target detection task from different angles, but ignore the importance of different features in the downsampling process, and the detection head does not fully consider the feature representation and detection requirements of dense small targets, resulting in low detection accuracy of small targets. At the same time, the large number of parameters and calculations of the detection head lead to a long training process and significantly increase the inference time. In addition, small targets are difficult to identify based on appearance alone and need to be identified with the help of the surrounding environment. Therefore, a method that focuses on small target detection is urgently needed to solve the above problems. Summary of the invention

[0006] In view of the deficiencies in the above-mentioned existing methods, the present invention provides a remote sensing image target detection method based on adaptive downsampling and scale enhancement detection head.

[0007] The present invention is implemented by the following scheme:

[0008] Step 1: Obtain remote sensing image target detection dataset;

[0009] Step 2: Preprocess the remote sensing image dataset;

[0010] Step 3: Establish a remote sensing image target detection model;

[0011] Step 4: Use the training set of the data set to train the model, construct a loss function to update the parameters of the model, and obtain the optimal model;

[0012] Step 5: Use the optimal model to detect the test set images of the data set to obtain the final test results.

[0013] Furthermore, in step 2, the AI-TOD remote sensing dataset is cropped into 800×800 pixel blocks with an overlap of 200 pixels; if the original image is smaller than 800×800, it is padded with zeros; in addition, the label file is converted into txt format.

[0014] Furthermore, the step 3 establishes a remote sensing image target detection model: in the feature extraction stage, a dynamic feature extraction module is used to flexibly adjust the receptive field of the feature extraction network to obtain contextual information of tiny targets; in the feature fusion stage, an adaptive downsampling module is used to emphasize the importance of different regions in the feature map and retain key information features to achieve efficient and adaptive downsampling; in the prediction stage, a scale-enhanced detection head with a decoupled design is used to provide the best feature representation for classification and regression tasks and reduce the number of model parameters through shared convolution; a high-resolution feature map is input into the detection head to obtain a feature map specifically used to predict tiny targets, thereby improving sensitivity to tiny targets.

[0015] Furthermore, the dynamic feature extraction module divides the channel dimension into two branches through 1×1 convolution, one of which is directly passed to the output, and the other branch dynamically adjusts the effective receptive field through processing by the dynamic perception unit; then, the features of the two branches are spliced; finally, the number of channels is adjusted through 1×1 convolution to obtain the output.

[0016] Furthermore, the dynamic perception unit is composed of an expansion re-parameter module, a channel attention and a feed-forward network.

[0017] Furthermore, the expansion re-parameter module uses multiple expansion convolutions to enhance the effect of a large-kernel standard convolution. The expansion re-parameter module structure is divided into two types: a and b. In the training stage, a is composed of three parallel branches, and the kernel sizes in the convolution layer are {5, 3, 3}, and the void rates are {1, 2, 3}. b is composed of five parallel branches, and the convolution kernels are {5, 7, 3, 3, 3}, and the void rates are {1, 2, 3, 4, 5}. After the convolution layer, their outputs are dimensionally spliced ​​after the normalization layer. In the inference stage, the batch normalization layer is fused into the convolution layer using the structural re-parameterization method to improve the inference speed. The size of the fused convolution kernel is related to the size of the expanded convolution kernel in each branch, and the largest convolution kernel is taken as the size of the fused convolution kernel.

[0018] Furthermore, the channel attention allows the model to emphasize useful features by weighting the channels, thereby improving the quality of features generated by the model. First, input feature map X∈R C×H×W Through global average pooling, space compression is performed to obtain a new feature Z∈R C×1×1 ; Then, through a gating mechanism consisting of two fully connected layers and an activation function, weights are generated for each feature channel. Finally, the weights are applied to the original feature map through channel-by-channel weighting.

[0019] Furthermore, the feedforward network improves the nonlinearity of feature representation through two layers of linear changes, activation functions and normalization operations, while extracting more advanced and comprehensive features. First, the input features are mapped to a higher-dimensional space through a layer of linear transformation; next, after the GELU activation function, nonlinear elements are introduced to enable the model to capture more complex feature representations; in addition, global response normalization is used for channel normalization to adjust the contribution of each channel feature; finally, a layer of linear transformation is used to map the high-dimensional space features to the original dimension; at the same time, batch normalization is used to normalize the data, which helps to reduce covariate shift during training.

[0020] Furthermore, the adaptive downsampling module is composed of a weight branch and a downsampling branch. First, the spatial weight of each pixel position in the input feature map is calculated by average pooling and convolution operations, so that the model can adaptively assign weights to the input features, and the obtained feature weights need to be rearranged so as to be element-by-element multiplied with the downsampled feature map. The specific feature rearrangement operation refers to extracting each 2×2 area of ​​the feature map as a separate set of information. Each 2×2 area corresponds to 4 elements, which come from 4 pixels in the area (upper left, upper right, lower left, and lower right). The four values ​​of each small area are probabilistically distributed to obtain the attention of each position in the small area, which helps the model focus on important small areas in the feature map and extract small target features more accurately. Secondly, the downsampling branch uses grouped convolution to divide the channels of the input feature map into multiple groups, and each group is convolved separately to reduce the spatial dimension and maintain the correlation between channels, while reducing the number of parameters and computational cost of the model. Finally, the adaptive weight is element-by-element multiplied with the rearranged downsampled feature map, and the results in each group are weighted summed. This selective aggregation of information ensures that key features can receive higher weights, thereby achieving effective and efficient downsampling operations.

[0021] Furthermore, the scale enhancement detection head adopts the design concept of the decoupling head and is composed of four decoupling heads. First, each decoupling head is responsible for learning to extract specific scale and semantic information. By processing the feature map processed by the neck network, it is possible to predict small, medium and large scale targets. At the same time, the high-resolution feature map is input into the detection head to obtain a feature map specifically used to predict tiny targets. Secondly, shared convolution is used to extract features and reduce the number of parameters of the model. The features after shared convolution are used to predict the position of the bounding box and the probability of the target category, respectively. These predicted values ​​will be used to generate the final detection results. Finally, a scaling operation is performed on the feature map in the regression branch to match the scale of the detection result.

[0022] Furthermore, the loss function consists of two parts: classification loss and regression loss; the classification loss adopts the binary cross entropy loss function; the regression loss adopts the Distribution Focal loss and the NWD loss function.

[0023] Beneficial effects of the present invention:

[0024] The present invention proposes a remote sensing image target detection method based on adaptive downsampling and scale enhancement detection head, the model includes a feature extraction network, a feature fusion network and a prediction network; the present invention constructs a dynamic feature extraction module, flexibly adjusts the receptive field of the feature extraction network to obtain the surrounding environment information of small targets; the present invention designs an adaptive downsampling module, emphasizes the importance of different areas in the feature map, retains key information, so as to achieve efficient and adaptive downsampling; the present invention proposes a scale enhancement detection head, inputs a high-resolution feature map into the detection head, improves the sensitivity to tiny targets, and reduces the number of parameters of the model by extracting features through shared convolution. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0026] Figure 1 is an overall flow chart of an embodiment of the present invention;

[0027] Figure 2 This is a schematic diagram of the structure of a dynamic feature extraction module in an embodiment of the present invention;

[0028] Figure 3 This is a schematic diagram of the expansion parameter module structure in an embodiment of the present invention;

[0029] Figure 4 Schematic diagram of the structure of channel attention in an embodiment of the present invention;

[0030] Figure 5 Schematic diagram of the structure of a feedforward network in an embodiment of the present invention;

[0031] Figure 6 Schematic diagram of the structure of an adaptive downsampling module in an embodiment of the present invention;

[0032] Figure 7 Schematic diagram of the structure of the scale enhancement detection head in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The technical solutions in the embodiments of the present invention will be further described below in conjunction with the accompanying drawings in the embodiments of the present invention, but are not intended to limit the present invention.

[0034] like Figure 1 As shown, the present invention provides a remote sensing image target detection method based on adaptive downsampling and scale enhancement detection head, the steps are as follows:

[0035] Step 1: Obtain remote sensing image target detection dataset;

[0036] Step 2: Preprocess the remote sensing image dataset;

[0037] Step 3: Establish a remote sensing image target detection model;

[0038] Step 4: Use the training set of the data set to train the model, construct a loss function to update the parameters of the model, and obtain the optimal model;

[0039] Step 5: Use the optimal model to detect the test set images of the data set to obtain the final test results.

[0040] Furthermore, in step 2, the AI-TOD remote sensing dataset is cropped into 800×800 pixel blocks with an overlap of 200 pixels. If the original image is smaller than 800×800, it is padded with zeros. In addition, the label file is converted into txt format.

[0041] Furthermore, the step 3 establishes a remote sensing image target detection model: in the feature extraction stage, a dynamic feature extraction module is used to flexibly adjust the receptive field of the feature extraction network to obtain contextual information of small targets and generate high-quality feature representations; channel attention is used to weight the channels so that the model emphasizes useful features, thereby improving the quality of features generated by the model; a feedforward network is used to enhance the nonlinearity of feature representation and extract more advanced and complex features; in the feature fusion stage, an adaptive downsampling module is used to emphasize the importance of different regions of features and retain key information to achieve efficient and adaptive downsampling; in the prediction stage, a scale-enhanced detection head with a decoupled design is used to extract features through shared convolution to provide the best feature representation for classification and regression tasks and reduce the number of parameters and calculations of the model; a high-resolution feature map is input into the detection head to obtain a feature map specifically used for predicting small targets; and the sensitivity to small targets is improved.

[0042] like Figure 2 As shown in the figure, the dynamic feature extraction module divides the channel dimension into two branches through 1×1 convolution, one of which is directly passed to the output and the other is processed by the dynamic perception unit. The dynamic perception unit is composed of an expansion re-parameter module, a channel attention and a feedforward network. Then, the features of the two branches are concatenated. Finally, the number of channels is adjusted through 1×1 convolution to obtain the output. The implementation definition process of this module is:

[0043] Branch 1 =Res(Conv(x))

[0044] Branch 2 =Dsu(Conv(x))

[0045] Output = Conv(concat(Branch 1 ,Branch 2 ))

[0046] In the formula, Branch 1 and Branch 2 They represent the results after being processed by the two branches respectively, and Dsu represents the dynamic perception unit.

[0047] like Figure 3 As shown, the dilated re-parameter module uses multiple dilated convolutions to enhance the effect of a large-kernel standard convolution. The dilated re-parameter module structure is divided into two types, a and b. During the training phase, a is composed of three parallel branches, and the kernel sizes in the convolution layer are {5, 3, 3}, and the void rates are {1, 2, 3}. b is composed of five parallel branches, and the convolution kernels are {5, 7, 3, 3, 3}, and the void rates are {1, 2, 3, 4, 5}. After the convolution layer, their outputs are dimensionally spliced ​​after the normalization layer. When splicing dimensions, the feature map size is made consistent by padding size. The calculation formula for padding is as follows:

[0048]

[0049] In the formula, k represents the size of the convolution kernel, and r represents the dilation rate. In the inference stage, the normalization layer is fused into the convolution layer using the structural reparameterization method to improve the inference speed. The size of the fused convolution kernel is related to the size of the expanded convolution kernel in each branch. The largest convolution kernel is taken as the size of the fused convolution kernel. The calculation formula for the expanded convolution kernel size is:

[0050] K=r*(k-1)+1

[0051] like Figure 4 As shown in the figure, the channel attention makes the model emphasize useful features by weighting the channels, thereby improving the quality of the features generated by the model. First, input the feature map X∈R C×H×W Through global average pooling, space compression is performed to obtain a new feature Z∈R C×1×1 Then, through a gating mechanism consisting of two fully connected layers and an activation function, weights are generated for each feature channel. Finally, the weights are applied to the original features by weighting each channel. The process is defined as:

[0052]

[0053] Where GAP and FC represent global average pooling and fully connected layer operations respectively. Represents a weighted operation.

[0054] like Figure 5As shown in the figure, the feedforward network is operated through two layers of linear changes, activation functions and normalization. First, the input features are mapped to a higher dimensional space through a layer of linear transformation. Next, the channel normalization is performed through the GELU activation function and global response normalization to adjust the contribution of each channel feature. Finally, the high-dimensional space features are mapped to the original dimension through a layer of linear transformation. The process is defined as:

[0055] Output FF =Linear(GRN(GELU(Linear(x))))

[0056] In the formula, Linear represents linear change, and GRN represents the global response normalization operation.

[0057] like Figure 6 As shown, the adaptive downsampling module consists of a weight branch and a downsampling branch. First, the spatial weight of each pixel position in the input feature map is calculated by average pooling and convolution operations, and rearranged so as to perform element-by-element multiplication with the downsampled feature map. The specific feature rearrangement operation refers to extracting each 2×2 area of ​​the feature map as a separate set of information. Each 2×2 area corresponds to 4 elements, which come from 4 pixels in the area (upper left, upper right, lower left, and lower right). The four values ​​of each small area are probabilistically distributed to obtain the attention of each position in the small area, which helps the model focus on important small areas in the feature map and extract small target features more accurately. Secondly, the downsampling branch uses grouped convolution to divide the channels of the input feature map into multiple groups, and each group is convolved separately to reduce the spatial dimension and maintain the correlation between channels, while reducing the number of parameters and computational cost of the model. Finally, the adaptive weight is multiplied element-by-element with the rearranged downsampled feature map, and the results in each group are weighted summed.

[0058] like Figure 7 As shown in the figure, the scale enhancement detection head adopts the design concept of the decoupling head and consists of four decoupling heads. First, the number of channels is adjusted by 1×1 convolution. Next, the feature map is passed through two shared 3×3 convolutions to extract the relevant features of the bounding box and category prediction, thereby obtaining richer features and reducing the number of parameters of the model. Subsequently, the feature map is convolved to convert it into the prediction of the bounding box and category probability required by the model, where the regression branch performs a scaling operation on the feature map to improve the robustness of the model. Finally, the scaled feature map is spliced ​​with the feature map processed by the classification convolution layer in the channel dimension to obtain the final result.

[0059] Furthermore, the loss function consists of two parts: classification loss and regression loss. Classification loss represents the difference between the predicted category and the true label, helping the model to correctly classify the target; while regression loss mainly measures the error between the predicted box and the true box, helping the model to accurately locate the position of the target.

[0060] The classification loss uses binary cross entropy loss, and multiple binary classifications are superimposed to achieve multi-label classification. The definition of classification loss is as follows:

[0061] L cls = -LlogP-(1-L)log(1-P)

[0062] Where L represents label confidence and P represents prediction confidence.

[0063] In regression loss, first use Distribution Focal Loss to quickly focus the model's predicted position on the value near the label. The calculation formula is as follows:

[0064] DFL(p(i),p(i+1))=-[(y i+1 -y)log(p(i))+(yy i )·log(p(i+1))]

[0065] where p(i) and p(i+1) are given by y i ,y i+1 , y is determined, where y is the label value. Then, in order to reduce the sensitivity of IoU to small target position deviations, NWD loss is used to further refine the target position. This loss function first models the bounding box as a 2D Gaussian distribution, and then calculates the similarity using the Gaussian distribution corresponding to the NWD metric. The calculation formula is as follows:

[0066] L NWD =1-NWD(N p ,N g )

[0067] Among them, N p Represents the Gaussian distribution model of the prediction box, N g A Gaussian distribution model representing the ground-truth box.

[0068] The above is a specific embodiment of the present invention. It should be noted that the present invention is not limited to the above specific implementation. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A remote sensing image target detection method based on adaptive downsampling and scale enhancement detection head, characterized in that: The method comprises: Step 1: Obtain remote sensing image target detection dataset; Step 2: Preprocess the remote sensing image dataset; Step 3: Establish a remote sensing image target detection model; In the feature extraction stage, a dynamic feature extraction module is used to flexibly adjust the receptive field of the feature extraction network to obtain contextual information of small targets; the dynamic feature extraction module divides the channel dimension into two branches through 1×1 convolution, one of which is directly passed to the output, and the other branch is processed by the dynamic perception unit to dynamically adjust the effective receptive field; then, the features of the two branches are spliced; finally, the number of channels is adjusted through 1×1 convolution to obtain the output; In the feature fusion stage, an adaptive downsampling module is used to emphasize the importance of different regions in the feature map and retain the key information features of small targets to achieve efficient and adaptive downsampling; the adaptive downsampling module consists of a weight branch and a downsampling branch; the weight branch requires feature rearrangement to extract each 2×2 region of the feature map as a separate set of information; each 2×2 region corresponds to 4 elements, which come from 4 pixels in the region (upper left, upper right, lower left, and lower right); the four values ​​of each small region are probabilistically distributed to obtain the attention of each pixel in the small region, which helps the model focus on important areas in the feature map and extract small target features more accurately; In the prediction stage, the scale-enhanced detection head adopts the design concept of the decoupled head. It consists of four decoupled heads that can predict the detection heads of tiny, small, medium, and large targets. At the same time, shared convolution is used to extract the features of classification and regression tasks to predict the bounding box position and the probability of the target category, so as to reduce the number of model parameters. Step 4: Construct the target detection loss function and perform iterative training; Step 5: Use the trained weights for the test set and output the detection results.

2. The method for remote sensing image target detection based on adaptive downsampling and scale enhancement detection head according to claim 1, characterized in that: The AI-TOD remote sensing dataset was cropped into 800 × 800 pixel blocks with an overlap of 200 pixels; If the original image is smaller than 800×800, fill it with zeros; Also, convert the label file to txt format.

3. The method for remote sensing image target detection based on adaptive downsampling and scale enhancement detection head according to claim 1, characterized in that: The dynamic perception unit is composed of an expansion re-parameter module, a channel attention and a feed-forward network.

4. The method for remote sensing image target detection based on adaptive downsampling and scale enhancement detection head according to claim 1, characterized in that: The dilation re-parameter module uses multiple dilated convolutions to enhance the effect of a large kernel standard convolution. The module uses different dilation rates during the training phase to capture features of different scales, increase the receptive field, and reduce the overlap of feature maps. In the inference stage, the normalization layer is fused into the convolution layer using the structural reparameterization method, thereby improving the inference speed; the channel attention enables the model to emphasize useful features by weighting the channels, thereby improving the quality of features generated by the model; the feedforward network retains the spatial structural information of the features through two layers of linear changes, activation functions and normalization operations.

5. The method for remote sensing image target detection based on adaptive downsampling and scale enhancement detection head according to claim 1, characterized in that: The adaptive downsampling module allows the model to learn the optimal weight distribution through weight branches. Specifically, each 2×2 area corresponds to 4 elements in the area (for example, upper left corner, upper right corner, lower left, and lower right); these four values ​​are probabilistically distributed to obtain attention in each area, which helps the model focus on important detail information in the feature map and extract small target features more accurately; secondly, the downsampling branch uses grouped convolution to reduce the spatial dimension and maintain the correlation between channels, while reducing the number of parameters and computational cost of the model; finally, the adaptive weights are weighted with the downsampled feature map; this selectively aggregates information to ensure that key features can get higher weights, thereby achieving effective and efficient downsampling operations.

6. The method for remote sensing image target detection based on adaptive downsampling and scale enhancement detection head according to claim 1, characterized in that: The scale enhancement detection head inputs the high-resolution feature map into the detection head to obtain a feature map specifically used to predict tiny targets. Secondly, in order to reduce the number of model parameters, features are extracted through shared convolution. The features after shared convolution are used to predict the position of the bounding box and the probability of the target category respectively. These predicted values ​​will be used to generate the final detection results.

7. The method for remote sensing image target detection based on adaptive downsampling and scale enhancement detection head according to claim 1, characterized in that: The loss function consists of two parts Composition: classification loss and regression loss; The classification loss uses the binary cross entropy loss function; the regression loss uses the Distribution Focal loss and NWD loss functions.

Citation Information

Patent Citations

  • Remote sensing image target detection method based on progressive feature smoothing and scale adaptive expansion convolution

    CN117612029A

  • Multi-scale remote sensing image target detection method based on enhanced small target feature extraction

    CN117809200A

  • Remote sensing image aircraft detection method based on multiple adaptive mechanisms

    CN117853944A

  • Method and system for detecting small target ship in SAR remote sensing image

    CN119131607A

Cited By

  • Re-parameterization unmanned aerial vehicle target detection method based on multi-core fusion and omnidirectional connection

    CN121121561A