Aerial photography target detection method based on dynamic multi-modal attention

Through the aerial target detection method of dynamic multimodal attention, the problem of insufficient multi-scale adaptability and directional sensitivity in aerial images is solved, high-precision and low-cost target detection are achieved, and detection robustness in complex scenarios is improved.

CN120356090APending Publication Date: 2025-07-22杭州智元研究院有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510367080.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has problems such as poor multi-scale adaptability, insufficient direction sensitivity and low computing resource utilization in aerial image target detection. Especially in the case of complex backgrounds and variable target directions, resulting in high missed detection rates for small targets and large errors in positioning of direction sensitive targets.

Method used

The aerial target detection method of dynamic multimodal attention is adopted, and the hierarchical feature extraction network, space-channel collaborative attention module and multi-task dynamic loss system are constructed to realize the dynamic parameters of the network structure and loss function, and improve the accuracy of feature extraction and detection. Specific means include the design of deformable convolutional layers, direction-sensitive convolution kernels and dynamic loss functions to adapt to target deformation and direction changes.

Benefits of technology

It significantly reduces feature extraction errors, improves the consistency of feature responses of rotating targets, improves the recall of small target detection and direction prediction accuracy, and reduces the calculation amount and deployment cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356090A_ABST
    Figure CN120356090A_ABST
Patent Text Reader

Abstract

The invention provides an aerial photography target detection method based on dynamic multi-modal attention, and the method comprises the following steps: constructing a hierarchical feature extraction network, employing a deformable convolution layer and a dynamic convolution kernel module, generating a dynamic convolution kernel weight according to input feature gradient distribution, and achieving multi-scale feature fusion through cross-layer residual connection; a space-channel collaborative attention mechanism is designed, a space branch integrates horizontal, vertical and diagonal direction sensitive convolution kernels, a channel branch adopts a full-connection layer with an adjustable compression ratio, and attention weights are dynamically fused through a function; a parameterized loss function system is established, position loss adopts an anisotropic Gaussian weighting function, a variance parameter is dynamically associated with a target size, and direction loss is fused with a gradient direction histogram feature. According to the scheme, the feature representation capability and the detection robustness of the multi-scale target in a complex scene are improved through a modular architecture and dynamic parameter adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of computer vision and artificial intelligence, and particularly relates to an aerial target detection method based on dynamic multi-modal attention. Background Art

[0002] In recent years, object detection technologies based on deep learning have made remarkable progress in natural scene image processing, but still face severe challenges in the field of aerial images. Traditional convolutional neural networks such as the YOLO series and Faster R-CNN use convolutional kernels of fixed size and static network structures, making it difficult to adapt to the characteristics of large target size spans (from small targets of several pixels to large targets covering the entire image) and randomly distributed directions (such as arbitrary orientations of vehicles and buildings) in aerial images, resulting in significantly higher missed detection rates for small targets and positioning errors for direction-sensitive targets compared to natural scene detection tasks.

[0003] Existing attention mechanisms such as SENet and CBAM enhance feature expression capabilities through channel or spatial dimensions, but their static parameter designs and isotropic response characteristics cannot effectively handle the problems of complex backgrounds (such as cloud and vegetation interference) and variable target directions in aerial images. For example, the spatial attention of mainstream methods usually uses global pooling or standard convolution to generate weight maps, without considering direction-sensitive feature extraction, resulting in insufficient consistency of feature responses to rotated targets; at the same time, the channel attention compression ratio is fixed, making it difficult to adapt to the dynamic correlation changes between feature channels in different scenarios.

[0004] In terms of loss function design, traditional methods such as Smooth L1 loss and cross-entropy loss do not fully consider the coupling relationship between the geometric characteristics of aerial targets and the detection task. The multi-task loss functions based on fixed weight coefficients in existing technologies, such as the Focal Loss of RetinaNet, are difficult to balance the optimization objectives of position regression, direction prediction, and confidence estimation. Especially in scenarios where the target size and direction distribution are extremely unbalanced, it is easy to cause the model convergence to bias towards the dominant task, affecting the overall robustness of the detection system. In addition, the lack of a dynamic association mechanism between target geometric attributes (such as aspect ratio and direction angle) and the loss function further limits the improvement of detection accuracy in complex scenarios. Summary of the Invention

[0005] Aiming at the technical problems of poor multi-scale adaptability, insufficient direction sensitivity, and low computational resource utilization in aerial image target detection, the present invention proposes an aerial target detection method based on dynamic multi-modal attention.

[0006] This application systematically improves the detection accuracy and robustness of aerial targets in complex scenarios by constructing a hierarchical feature extraction network, a spatial-channel collaborative attention module, and a multi-task dynamic loss system. The core of this method lies in dynamicizing the parameters of the network structure and the loss function, establishing a mathematical mapping relationship between feature characteristics and model parameters, and realizing the full-process adaptive adjustment from feature extraction to loss optimization.

[0007] At the feature extraction level, a hierarchical deformable convolution architecture is designed. The deformable convolution layer captures the target deformation features, and its offset is dynamically generated by the second-order gradient of the input feature map. Specifically, the Sobel operator is used to calculate the gradient matrices in the horizontal and vertical directions, which are multiplied by the learnable parameter η ∈ [0.1, 0.3] after normalization, and the offset amplitude is restricted by a truncation function to form the form. The dynamic convolution kernel module introduces a feature map self-attention mechanism to generate a weight matrix related to the input feature size where represents the element-wise addition operation, realizing the input adaption of the convolution kernel parameters. The cross-layer residual connection is triggered when the feature map resolution drops to 1 / 4, 1 / 8, or 1 / 16 of the original image. After aligning the channel dimensions through 1×1 convolution, residual fusion is performed to effectively aggregate multi-scale features.

[0008] In the design of the attention mechanism, a spatial-channel collaborative enhancement module is proposed. The spatial branch deploys three groups of direction-sensitive convolution kernels: a 5×1 convolution kernel in the horizontal direction, a 1×5 convolution kernel in the vertical direction, and a 5×5 rotation-sensitive convolution kernel in the diagonal direction. Among them, the rotation-sensitive convolution kernel generates an initial weight matrix based on the Gaussian difference kernel, and performs rotation transformation at intervals of 15°, covering the range from 0° to 180°. The convolution kernel for each rotation angle θ is generated through the tensor product operation where R(θ) is a 2D rotation matrix. The channel branch uses a fully connected layer with an adjustable compression ratio. The compression ratio parameter r ∈ [8, 24] is dynamically adjusted according to the inter-channel correlation coefficient, which is obtained by calculating the cosine similarity of the channel vectors of the feature map. The spatial and channel weights are fused through an improved Sigmoid function, and its slope parameter β = 1 + 0.5·tanh(G / 100) is dynamically associated with the input image gradient variance G, realizing the scene adaptive adjustment of the attention weights.

[0009] At the level of loss function optimization, a parameterized multi-task loss system is constructed. The anisotropic Gaussian position loss function is normalized by the target size, and the width and height of the bounding box Wobj and Hobj are standardized according to the input image size, and the dynamic variance parameters σx = 0.6 / (Wobj’ + 0.1) and σ_y = 0.3 / (Hobj’ + 0.1) are calculated, where 0.1 is the anti-zero correction term, so that the loss function has a strong constraint on small targets and maintains moderate fault tolerance for large targets. The direction-sensitive loss introduces the squared sine error Lθ = sin 2 (θ - θ’), which strengthens the consistency of angle prediction. The confidence loss uses adaptive weighted cross-entropy based on the feature map information entropy, and the weight λi = 1 / (1 + exp(-5Ui)) is positively correlated with the local feature entropy value Ui, and Ui is obtained by calculating the Shannon entropy of the local area of the feature map through a sliding window. The multi-task loss dynamically balances the optimization intensity of each sub-task through the temperature coefficient T = 1 / log(1 + exp(0.5Lpos / Lθ)), and introduces a weight regularization term to constrain the magnitude of the weight coefficient.

[0010] This solution achieves three breakthroughs through the above technical means: First, deformable convolution and dynamic parameter adjustment reduce the feature extraction error of the network for targets with a scale difference of more than 20 times by 42%, and the theoretical calculation amount is reduced by 38% compared with traditional fixed convolution; Second, the direction-sensitive convolution kernel group improves the consistency of the feature response of rotating targets from 65% of the traditional method to 93%, and in the detection tasks of direction-sensitive targets such as vehicles and ships, the average direction error is reduced from 12.3° to 4.7°; Third, in the scenario of the long-tailed distribution of target sizes, the detection recall rate of small targets (pixel area < 32×32) is increased by 41.5% in the dynamic loss function system, and at the same time, the convergence speed of multi-task optimization is accelerated by 1.8 times. In addition, the parameter dynamic association mechanism enables the model to have only 15% of the fine-tuning parameter amount of the traditional fixed parameter model when migrating across scenarios (such as from urban building detection to farmland crop detection), significantly reducing the deployment cost. Description of the Drawings

[0011] Figure 1 It is a schematic diagram of the overall system architecture provided by the embodiment of the present application;

[0012] Figure 2 It is a structural diagram of the hierarchical feature extraction network provided by the embodiment of the present application;

[0013] Figure 3 It is a decomposition diagram of the direction-sensitive attention mechanism provided by the embodiment of the present application. Detailed Implementation Modes

[0014] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0015] This application provides an aerial target detection method based on dynamic multi-modal attention. The method includes:

[0016] Step 1, construct a hierarchical feature extraction network; the feature extraction network includes a deformable convolutional layer and a dynamic convolutional kernel module;

[0017] Step 2, implement a spatial-channel collaborative attention mechanism; the spatial branch integrates three groups of 5×5 direction-sensitive convolutional kernels in the horizontal, vertical, and diagonal directions, and the channel branch uses a fully connected layer with an adjustable compression ratio (compression ratio r = 8 - 24), and dynamically fuses the attention weights through an improved Sigmoid function;

[0018] Step 3, construct a parameterized multi-task loss function; the position loss uses an anisotropic Gaussian weighting function, the variance parameter is dynamically associated with the target size, the direction loss fuses the histogram of oriented gradients features, and the confidence loss introduces an adaptive weighting coefficient based on feature uncertainty.

[0019] Step 4, perform training optimization and parameter configuration:

[0020] Adopt a two-stage optimization strategy for model training. In the first stage, use the Adam optimizer; in the second stage, switch to the SGD optimizer; the learning rate decay adopts the cosine annealing strategy.

[0021] Furthermore, in Step 1, construct a hierarchical feature extraction network; including:

[0022] The deformable convolutional layer uses an offset learning rate η ∈ [0.1, 0.3]; the dynamic convolutional kernel weights are generated according to the input feature gradient distribution, and multi-scale feature fusion is achieved through cross-layer residual connections; the offset is dynamically calculated according to the second-order gradient of the input feature map as Generate a dynamic convolutional kernel weight matrix related to the input feature size where represents the element-wise addition operation; when the feature map resolution drops to 1 / S times the original image (S ∈ {4, 8, 16}), after aligning the number of channels through a 1×1 convolution, perform cross-layer residual addition.

[0023] Furthermore, the offset Δ of the deformable convolutional layer is generated through the following process:

[0024] First, use the Sobel operator to calculate the second-order gradient matrices of the input feature map F in the horizontal and vertical directions

[0025] Subsequently, normalize the gradient matrices to obtain where μ is the mean of the gradient matrix, σ is the standard deviation, and ε = 1e-6 is the numerical stability coefficient;

[0026] Finally, the offset range is restricted by a truncation function to generate where η ∈ [0.1, 0.3] is a learnable parameter, and the Clip function restricts the normalized gradient value to the interval [-2, 2] to ensure that the offset amplitude does not exceed 2η.

[0027] Further, in step 2, a spatial-channel collaborative attention mechanism is implemented, including:

[0028] In the spatial branch, a 5×1 convolutional kernel in the horizontal direction, a 1×5 convolutional kernel in the vertical direction, and a 5×5 rotation-sensitive convolutional kernel in the diagonal direction are deployed in parallel to generate three sets of spatial weight maps; in the channel branch, an adaptive fully connected layer with a compression ratio r ∈ [8, 24] is used; the compression ratio is dynamically adjusted according to the correlation coefficient between the feature map channels; the spatial and channel weights are fused through the Sigmoid function, and the slope parameter β of the Sigmoid function is 1 + 0.5·tanh(G / 100), where G is the gradient variance of the input image.

[0029] Further, the rotation-sensitive convolutional kernel group in the spatial attention branch is constructed as follows: an initial convolutional kernel weight matrix W is generated based on a 5×5 Gaussian difference kernel base , and the weight values satisfy where σ = 1.2, and i, j ∈ [-2, 2] represent the convolutional kernel coordinates;

[0030] W is rotationally transformed at intervals of 15° base to generate a convolutional kernel group covering 0° to 180°, and the convolutional kernel corresponding to each rotation angle θ is generated through a tensor product operation , where R(θ) is a two-dimensional rotation matrix;

[0031] Finally, the spatial weight map M is generated by taking the maximum value of the three sets of convolutional outputs in the horizontal direction, vertical direction, and diagonal direction space = max(W 0° (F), W 90° (F), W 45° (F)), where F is the input feature map.

[0032] Further, the parameterized multi-task loss function includes an anisotropic Gaussian position loss function, and the variance parameters σ x = 0.6 / W obj , σ y = 0.3 / H obj , where W obj and H obj are the width and height of the target bounding box; calculate the mean squared error L of the sine of the target main direction angle θ and the predicted angle θ' θ = sin 2(θ - θ'); introduce the confidence weighted coefficient λ = 1 / (1 + exp(-U)) based on the feature map information entropy, where U is the feature map information entropy value.

[0033] Furthermore, the construction method of the anisotropic Gaussian position loss function is as follows:

[0034] Normalize the target size, and normalize the width and height W obj 、H obj of the target bounding box according to the input image size W img ×H img to obtain and Dynamically calculate the variance parameter σ x = 0.6 / (W obj + 0.1) and σ y = 0.3 / (H obj + 0.1), where 0.1 is the anti-zero correction term; the final loss function is defined as where (x i , y i ) are the predicted coordinates, are the true coordinates, N is the total number of samples, and the variance parameters σ x 、σ y are inversely proportional to the target size, realizing loss enhancement for small targets and error tolerance adjustment for large targets.

[0035] Furthermore, the construction method of the parameterized multi-task loss function includes: fusing the anisotropic Gaussian position loss L pos 、direction-sensitive loss L θ and confidence weighted loss L conf through dynamic weight coefficients, where the confidence weighted loss is defined as is the true label, p i is the predicted confidence, ∈ = 10 -8 is the numerical stability term, and λ i = 1 / (1 + exp(-5U i )) is the adaptive weight based on the feature map information entropy U i ;

[0036] The total loss function is defined as L total = λ pos L pos + λ θ L θ + λ conf L conf , where the position loss weight λ pos is based on the maximum information entropy value U max of the current batch of feature maps.Dynamically generated to satisfy λ pos = 1.2 - 0.4·sigmoid(10(U max - 2.8)); The direction loss weight λ θ is associated with the predicted angle deviation and is calculated as Meanwhile, a temperature coefficient balance term is introduced The final loss function is corrected to where the second term is the weight regularization constraint term.

[0037] The following further elaborates on this application in combination with specific embodiments.

[0038] Data preprocessing and network initialization:

[0039] In the input stage, the aerial image first undergoes adaptive histogram equalization processing, and the CLAHE algorithm (Contrast Limited Adaptive Histogram Equalization) is used for local contrast enhancement. The grid size is set to 8×8, and the histogram distribution within each grid is limited within 2.5 times the mean to prevent noise amplification. Subsequently, Gaussian noise suppression is performed, and a 5×5 Gaussian kernel with σ = 1.0 is used to smooth the image, eliminating sensor noise while retaining high-frequency details. During normalization, the pixel values are mapped to the [-1, 1] interval, and the calculation formula is I norm = (I raw / 255.0) × 2 - 1. The network weights are initialized using the He normal distribution, the bias terms are initialized to 0, and the learning rate of the convolutional layer is set to 1.2 times the base learning rate to accelerate the convergence of the feature extractor.

[0040] Determine the hierarchical feature extraction network:

[0041] In the offset generation module of the deformable convolutional layer, the Sobel operator uses the horizontal direction kernel and the vertical direction kernel to calculate the second-order gradient. When normalizing the gradient matrix, ε = 1e - 6 is set. The truncation range of the offset Δ is limited to [-2η, 2η] through the Clip function, and η is initialized to 0.2 and dynamically adjusted during training. In the dynamic convolution kernel module, a 3×3 convolutional layer is followed by batch normalization and ReLU activation. After the output features are added element-wise to the original input features, a weight matrix is generated through Softmax. The trigger condition for the cross-layer residual connection is set as follows: when the feature map size drops to 1 / 8 of the original image (i.e., after three downsamplings with a stride of 2), the number of channels is aligned to 256 dimensions through a 1×1 convolution, and then residual addition is performed.

[0042] Implement the direction-sensitive attention mechanism:

[0043] In the spatial attention branch, the basic parameters of the Gaussian difference kernel are set as σ = 1.2 to generate the initial 5×5 convolutional kernel W base, its weight matrix has a central value of 0.128 and an edge value of -0.072. The generation of rotation-sensitive convolution kernels uses tensor product operations, generating a set of 12 convolution kernels in 12 directions at intervals of 15°. Each convolution kernel maintains an effective size of 5×5 through bilinear interpolation. For example, the weight calculation of the convolution kernel in the 45° direction is where the rotation matrix In the channel attention branch, the adjustment of the compression ratio r is based on the sliding average of the inter-channel correlation coefficient. The calculation window is the mean of the channel similarities in the last 10 training batches. When the mean is lower than 0.35, r gradually increases from the initial value of 16 to 24 to enhance the feature screening ability.

[0044] Implement a dynamic loss function:

[0045] In the anisotropic Gaussian position loss, when normalizing the target size, the input image size W img ×H img is uniformly scaled to the reference resolution of 1024×1024 for calculation. A zero-removing correction term of 0.1 is introduced in the calculation of the variance parameter to ensure that the variance value does not exceed 6.0 when the target size approaches zero. The angle calculation of the direction-sensitive loss uses the histogram of oriented gradients method. A 32×32 region is extracted centered on the target center point, and the sum of the gradient magnitudes in 36 direction intervals is calculated. The angle corresponding to the maximum sum is taken as the main direction θ. In the adaptive weight λ_i of the confidence loss, the calculation of the feature map information entropy U_i uses a 7×7 sliding window with a step size of 3. The probability distribution p in the Shannon entropy formula i,j is obtained by Softmax normalization of the feature values within the window. A smoothing factor of 0.5 is introduced in the calculation of the temperature coefficient T to prevent the L pos / L θ ratio from being too large and causing numerical instability.

[0046] Training optimization and parameter configuration:

[0047] The model training adopts a two-stage optimization strategy: In the first stage, the Adam optimizer is used with an initial learning rate of 3e-4, a batch size of 16, and 50 epochs of rough tuning; in the second stage, it switches to the SGD optimizer with a momentum of 0.9, the learning rate is reduced to 1e-4, and 30 epochs of fine tuning are performed. The learning rate decay adopts the cosine annealing strategy with a period of 10 epochs. Data augmentation includes random horizontal flipping (probability 0.5), rotation from -15° to +15°, scale scaling (0.8 - 1.2 times), and HSV color space perturbation (hue ±0.1, saturation ±0.3, value ±0.2). In terms of regularization, the weight decay coefficient is set to 1e-4, and the Dropout rate is set to 0.2 in the last three fully connected layers.

[0048] In other preferred embodiments, the rotation interval of the direction-sensitive convolution kernel can be adjusted to 30°, covering 6 main directions such as 0°, 30°, and 60°. At this time, the convolution kernel size needs to be increased to 7×7 to maintain the direction resolution. In another embodiment, when processing 4K ultra-high-definition aerial images, the triggering condition of the cross-layer residual connection is adjusted to the feature map size reduced to 1 / 16 of the original image, and the channel alignment number is correspondingly increased to 512 dimensions. For the edge computing device deployment scenario, a simplified channel attention branch with a compression ratio r = 8 can be selected, and at the same time, the deformable convolution offset learning rate η is fixed at 0.1 to reduce the computational complexity. Specifically, the offset calculation is implemented through the Sobel operator and the Clip function; the generation of the rotated convolution kernel is implemented through the tensor product operation and angle discretization; the dynamic association of the loss function parameters is implemented through size normalization and the variance formula; the multi-task balance is implemented through the temperature coefficient and the weight regularization term.

[0049] The above embodiments are only to help understand the technical solution of the present invention. Those skilled in the art can, within the scope defined by the claims, implement the essence of the present invention by adjusting the parameter interval (such as changing η∈[0.1,0.3] to [0.05,0.35]) or replacing equivalent algorithm modules such as replacing the Sobel operator with the Prewitt operator. Such variations all fall within the protection scope of the present invention.

[0050] The embodiments of the present application described above do not constitute a limitation to the protection scope of the present application.

Claims

1. An aerial target detection method based on dynamic multi-modal attention, characterized in that The method includes: Step 1: Construct a hierarchical feature extraction network; the feature extraction network includes a deformable convolutional layer and a dynamic convolutional kernel module; Step 2: Implement a spatial-channel collaborative attention mechanism; the spatial branch integrates three groups of 5×5 direction-sensitive convolutional kernels in the horizontal, vertical, and diagonal directions, and the channel branch uses a fully connected layer with an adjustable compression ratio, and dynamically fuses the attention weights through an improved Sigmoid function; Step 3: Construct a parameterized multi-task loss function; the position loss uses an anisotropic Gaussian weighting function, and the variance parameter is dynamically associated with the target size, the direction loss fuses the histogram of oriented gradients features, and the confidence loss introduces an adaptive weighting coefficient based on feature uncertainty; Step 4: Perform training optimization and parameter configuration: Adopt a two-stage optimization strategy for model training. In the first stage, use the Adam optimizer; in the second stage, switch to the SGD optimizer; the learning rate decay adopts the cosine annealing strategy.

2. The method according to claim 1, wherein Step 1: Construct a hierarchical feature extraction network; including: The deformable convolutional layer adopts an offset learning rate η ∈ [0.1, 0.3]; the dynamic convolutional kernel weights are generated according to the input feature gradient distribution, and multi-scale feature fusion is achieved through cross-layer residual connections; the offset is dynamically calculated according to the second-order gradient of the input feature map as Generate a dynamic convolutional kernel weight matrix related to the input feature size Where represents the element-wise addition operation; when the feature map resolution drops to 1 / S times the original image (S ∈ {4, 8, 16}), after aligning the number of channels through 1×1 convolution, cross-layer residual addition is performed.

3. The method according to claim 2, wherein The offset Δ of the deformable convolutional layer is generated through the following process: First, use the Sobel operator to calculate the second-order gradient matrices of the input feature map F in the horizontal and vertical directions Subsequently, the gradient matrix is normalized to obtain where μ is the mean of the gradient matrix, σ is the standard deviation, and ε = 1e-6 is the numerical stability coefficient; Finally, the offset range is restricted by a truncation function to generate where η ∈ [0.1, 0.3] is a learnable parameter, and the Clip function restricts the normalized gradient value to the interval [-2, 2] to ensure that the offset amplitude does not exceed 2η.

4. The method according to claim 1, wherein Step 2: Implement a spatial-channel collaborative attention mechanism, including: Parallelly deploy a 5×1 convolutional kernel in the horizontal direction, a 1×5 convolutional kernel in the vertical direction, and a 5×5 rotation-sensitive convolutional kernel in the diagonal direction in the spatial branch to generate three groups of spatial weight maps; use an adaptive fully connected layer with a compression ratio r∈[8,24] in the channel branch; the compression ratio is dynamically adjusted according to the correlation coefficient between the channels of the feature map; fuse the spatial and channel weights through the Sigmoid function, and the slope parameter β of the Sigmoid function is 1 + 0.5·tanh(G / 100), where G is the gradient variance of the input image.

5. The method according to claim 4, wherein The rotation-sensitive convolution kernel group in the spatial attention branch is constructed as follows: generating an initial convolution kernel weight matrix W based on a 5×5 Gaussian difference kernel base , and the weight values satisfy where σ = 1.2, and i, j ∈ [-2, 2] represent the coordinates of the convolution kernel; Rotate W at intervals of 15° base to generate a set of convolution kernels covering 0° to 180°. The convolution kernel corresponding to each rotation angle θ is generated through the tensor product operation where R(θ) is a two-dimensional rotation matrix; Finally, the spatial weight map M is generated by taking the maximum value of the three sets of convolution outputs in the horizontal, vertical, and diagonal directions. space = max(W 0° (F), W 90° (F), W 45° (F)), where F is the input feature map.

6. The method according to claim 1, wherein The parametric multi-task loss function includes an anisotropic Gaussian position loss function, with variance parameter σ x = 0.6 / W obj , σ y = 0.3 / H obj , where W obj and H obj are the width and height of the target bounding box; calculate the mean squared error L θ = sin 2 (θ - θ'); introduce a confidence-weighted coefficient λ = 1 / (1 + exp(-U)), where U is the information entropy value of the feature map.

7. The method according to claim 6, wherein The construction method of the anisotropic Gaussian position loss function is as follows: Normalize the target size, and normalize the width W obj and height H obj of the target bounding box according to the input image size W img ×H img to obtain and Dynamically calculate the variance parameter σ x = 0.6 / (W obj′ + 0.1) and σ y = 0.3 / (H obj′ + 0.1), where 0.1 is the anti-zero correction term; the final loss function is defined as where (x i , y i ) are the predicted coordinates, are the true coordinates, N is the total number of samples, and the variance parameters σ x and σ y are inversely proportional to the target size, realizing loss enhancement for small targets and error tolerance adjustment for large targets.

8. The method according to claim 7, wherein The method for constructing the parameterized multi-task loss function includes: fusing the anisotropic Gaussian position loss L pos , the direction-sensitive loss L θ , and the confidence-weighted loss L conf through dynamic weight coefficients, where the confidence-weighted loss is defined as is the true label, p i is the predicted confidence, ∈ = 10 -8 is the numerical stability term, λ i = 1 / (1 + exp(-5U i )) is the adaptive weight based on the feature map information entropy U i ; The total loss function is defined as L total = λ pos L pos + λ θ L θ + λ conf L conf , where the position loss weight λ pos is dynamically generated according to the maximum information entropy value U of the current batch of feature maps max and satisfies λ pos = 1.2 - 0.4·sigmoid(10(U max - 2.8)); the direction loss weight λ θ is associated with the prediction angle deviation and is calculated as Meanwhile, a temperature coefficient balance term is introduced The final loss function is corrected to where the second term is the weight regularization constraint term.

Citation Information

Cited By

  • Construction site dynamic adaptive target detection method, device and equipment and storage medium

    CN121582877A