Rotating target detection method based on frequency division characteristic refinement
Through the four-stage frequency division feature refinement and sample coordinated tuning loss, the problem of insufficient feature extraction and inconsistency in the rotation object detection of remote sensing images is solved, and high-precision rotation object detection is achieved.
Patent Information
- Application Number
- CN202510351993.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-01
AI Technical Summary
The existing remote sensing image rotation object detection technology is insufficient in complex scenarios, resulting in a decrease in detection accuracy. Especially targets with large aspect ratios are susceptible to angle deviations, and classification and regression are inconsistent, which affects detection performance.
The four-stage frequency-dividing feature refinement method is adopted, combining position-aware attention blocks, self-focused Transformer blocks and context Transformer blocks to enhance feature extraction capabilities, and design sample coordinated tuning losses to optimize the model training process.
It improves the accuracy and robustness of rotary object detection, solves the problems of insufficient feature extraction and inconsistent classification regression in complex backgrounds, and improves the detection accuracy.
Smart Images

Figure CN120236064A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image rotated object detection, aiming to classify and locate objects in remote sensing images. Specifically, it is a rotated object detection method based on refined frequency-divided features. Background Art
[0002] Rotated object detection is one of the most basic tasks in remote sensing image processing. Its purpose is to identify the categories of single or multiple specific objects in a given remote sensing image and perform precise positioning. Compared with general object detection, rotated object detection can provide more accurate direction and scale information of the object, playing an important role in fields such as intelligence reconnaissance, disaster monitoring, and urban planning. At the same time, it is also an important basis for tasks such as object tracking and image segmentation.
[0003] Currently, most existing rotated object detection technologies are implemented based on deep learning. The rotated object detection technologies based on deep learning are mainly divided into two categories: convolutional-based rotated object detection and Transformer-based rotated object detection. The convolutional-based rotated object detection method extracts features using convolution by presetting anchor points and prior boxes or directly predicting key points, effectively capturing the spatial local information in the image. The Transformer-based rotated object detection method introduces self-attention to enhance the feature representation ability, expands the image receptive field, models the global relationship, and realizes the fusion of long-distance information.
[0004] Although significant progress has been made in remote sensing image rotated object detection, it still faces challenges. Pure convolutional methods only extract local features, while although the Transformer can model global dependencies, it lacks the ability to extract local information, affecting the detection accuracy. The remote sensing image scene is complex, the object scales are diverse and unevenly distributed, and densely arranged vehicles, ships, etc. are vulnerable to occlusion and interference with each other, increasing the detection difficulty. In addition, for objects with a large aspect ratio such as bridges and ports, the angle deviation will cause a sudden drop in the Intersection over Union (IoU), and the independent training of classification and regression leads to inconsistency between classification and regression, affecting the detection performance. Therefore, there is a need for a new rotated object detection method to solve the above problems. Summary of the Invention
[0005] The present invention aims to solve the problem of the decline in detection accuracy in remote sensing image rotated object detection due to insufficient feature extraction, high scene complexity, and uneven object distribution. Especially for objects with a large aspect ratio, their angle deviation is likely to cause a drop in IoU, as well as the problem of inconsistency between classification and regression, affecting the quality evaluation of the prediction box.
[0006] To solve the above problems, the present invention provides a rotated object detection method based on refined frequency-divided features. The method includes the following steps:
[0007] 1) Extract the initial feature X from the remote sensing image of size H×W×3 as the input of the frequency division module. The feature extraction part consists of two 3×3 depthwise separable convolutions;
[0008] 2) The frequency division module processes the input feature to obtain high-frequency features, low-frequency features, and complete features respectively;
[0009] 3) Based on the frequency division features obtained in step 2), input them into the feature interaction module to enhance the multi-level information expression of the features. This module consists of a position-aware attention block, a self-focusing Transformer block, and a context Transformer block, which extract local features, global features, and context features from high-frequency features, low-frequency features, and complete features respectively. Then, these three are concatenated by channels and fused using 1×1 convolution;
[0010] 4) The frequency division module and the feature interaction module in steps 2) and 3) are repeatedly executed four times to complete the four-stage feature refinement process. Then, use the Feature Pyramid Network (FPN) to extract multi-scale features, and through the Region Proposal Network (RPN) and Region of Interest (RoI) alignment, the final detection features are obtained;
[0011] 5) Finally, input the detection features into the detection head, and complete the category discrimination and precise positioning of the target through the classification branch and the regression branch respectively, so as to obtain the final detection result. The classification branch uses two fully connected layers, the regression branch uses two convolutional layers, and a sample collaborative tuning loss is designed to train the model.
[0012] The advantages of the present invention are as follows: First, the present invention proposes a four-stage frequency division feature refinement process, and designs a position-aware attention block, a self-focusing Transformer block, and a context Transformer block, which solves the problems of insufficient feature extraction and loss of context information in complex background environments. Second, the present invention proposes a sample collaborative tuning loss, which solves the problem of independent calculation of the classification branch and the regression branch. Brief Description of the Drawings
[0013] Figure 1 is the flow chart of the rotation target detection method based on frequency division feature refinement of the present invention.
[0014] Figure 2 is the model diagram of the rotation target detection method based on frequency division feature refinement of the present invention.
[0015] Figure 3 is the position-aware attention block.
[0016] Figure 4 is a self-focusing Transformer block.
[0017] Figure 5 is a context Transformer block. Detailed implementation manners
[0018] The present invention provides a rotational object detection method based on frequency division feature refinement. This method first extracts initial features from the input image; secondly, performs a four-stage frequency division feature refinement process from rough to fine on the extracted initial features, and each stage includes a frequency division module and a feature interaction module. The feature interaction module consists of a position-aware attention block, a self-focusing Transformer block, and a context Transformer block, which extract local features, global features, and context features respectively to fully capture multi-level feature information; then, splice the obtained local features, global features, and context features according to the channel dimension, and use convolution to obtain detection features; finally, select the region of interest, and use the fully connected layer and the convolutional layer to perform classification and regression respectively to obtain the final object category and accurate position, and complete the detection task. This model optimizes the features using the sample collaborative tuning loss. As Figure 1 shown, the present invention includes the following steps:
[0019] 1) Extract initial feature X from a remote sensing image with a size of H×W×3 as the input of the frequency division module. The feature extraction part consists of two 3×3 depthwise separable convolutions;
[0020] 2) The frequency division module processes the input features to obtain high-frequency features, low-frequency features, and complete features respectively.
[0021] 3) Based on the frequency division features obtained in step 2), input them into the feature interaction module, which consists of a position-aware attention block, a self-focusing Transformer block, and a context Transformer block, and extract local features, global features, and context features from the high-frequency features, low-frequency features, and complete features respectively. Then splice these three according to the channel and use 1×1 convolution for fusion.
[0022] 4) Repeat the frequency division module in step 2) and the feature interaction module in step 3) four times to complete the four-stage feature refinement process. Then use the Feature Pyramid Network (FPN) to extract multi-scale features, and through the Region Proposal Network (RPN) and Region of Interest (RoI) alignment, obtain the final detection features.
[0023] 5) Finally, input the detected features into the detection head, and complete the category discrimination and precise positioning of the target through the classification branch and the regression branch respectively, so as to obtain the final detection result. The classification branch uses two fully connected layers, the regression branch uses two convolutional layers, and a sample collaborative tuning loss is designed to train the model.
[0024] Further, the frequency division module in step 2) is specifically:
[0025] 2.1) Divide the input features along the channel dimension into high-frequency features and low-frequency features At the same time, take the complete input feature as the third branch; then, feature X I , X II and X III respectively go through depth convolution, pointwise convolution, batch normalization processing and Hardswish activation function to obtain high-frequency features, low-frequency features and complete features: X1, X2, X3.
[0026] X i = H(BN(PwC(DwC(X j )))), i = 1, 2, 3, j = I, II, III (1)
[0027] Among them, DwC(*) represents depth convolution, PwC(*) represents pointwise convolution, BN(*) represents batch normalization layer, and H(*) represents Hardswish activation function.
[0028] 2.2) Convolution is used to extract high-frequency features, Transformer is used to extract low-frequency features, and since the acquisition of the context relationship between features does not purely use global features and the receptive field range is larger than that of ordinary convolution, the complete input feature is used to extract context information. Therefore, X1, X2, X3 are respectively used as the inputs of the position-aware attention block, self-focusing Transformer block and context Transformer block.
[0029] Further, the feature interaction module in step 3) is specifically:
[0030] 3.1) Extract local features. Design a position-aware attention block, and use one-dimensional convolution to obtain attention maps of horizontal and vertical positions to enhance the local position information in high-frequency features, thereby improving the model's ability to capture spatial details.
[0031] 3.1.1) First, for the input high-frequency feature X1, perform average pooling along the horizontal and vertical directions respectively on each channel to obtain the feature f hand f w :
[0032]
[0033] 3.1.2) Then, perform one-dimensional convolution on the features f h and f w respectively to enhance the position information in the horizontal and vertical directions, and then use group normalization and the Sigmoid function in sequence to obtain the position attention maps p h and p w .
[0034] p h = σ(GN(Conv1d(f h ))) (4)
[0035] p w = σ(GN(Conv1d(f w ))) (5)
[0036] where σ(*) represents the Sigmoid activation function, GN(*) represents group normalization, Conv1d(*) represents one-dimensional convolution, and the convolution kernel size is 7.
[0037] 3.1.3) Finally, use the horizontal position attention map p h and the vertical position attention map p w as weights, multiply them with the input feature X1 to enhance the significant information in the horizontal and vertical directions respectively, and obtain the local feature Y L :
[0038] Y L = X1 * p h * p w (6)
[0039] where X1 is the input high-frequency feature, p h is the horizontal position attention map, and p w is the vertical position attention map.
[0040] 3.2) Extract global features. Propose a self-focus attention mechanism, design a self-focus Transformer block, and adaptively extract global features from the input low-frequency features, which helps the model perform more accurate detection in complex scenarios.
[0041] The self-focus Transformer block first performs layer normalization on the input features, then extracts features through self-focus attention, adds the obtained result to the input features to form a residual connection, and obtains intermediate features. These intermediate features are then processed by layer normalization and a feed-forward neural network, and the output result is added to the intermediate features to obtain the extracted global feature YG Among them, the self - attention process is as follows:
[0042] 3.2.1) First, project the low - frequency feature X2 that has undergone layer normalization onto the query, key, and value:
[0043] Q = X2W Q (7)
[0044] K = X2W K (8)
[0045] V = X2W V (9)
[0046] 3.2.2) Second, calculate the cosine similarity between Q and K, and use the Softmax function for normalization to obtain the saliency score map A QK :
[0047]
[0048] Among them, ||*|| represents the Euclidean norm of the feature.
[0049] 3.2.3) Then, use the saliency score map A QK to update K and V respectively, and obtain the updated and
[0050]
[0051] 3.2.4) Finally, perform self - attention calculation on Q and the updated and :
[0052]
[0053] where d k is the dimension of the matrix, used as a scaling factor to prevent the dot - product value from being too large, resulting in gradient vanishing or explosion.
[0054] 3.3) Extract context features. Introduce the context attention mechanism and design the context Transformer block for super - pixel feature aggregation of pixel information to extract context features.
[0055] The context Transformer block first performs layer normalization on the input features, then extracts features through context attention. The resulting result is added to the input features to form a residual connection, obtaining intermediate features. These intermediate features are then processed through batch normalization and a feed - forward neural network, and the output result is added to the intermediate features to obtain the extracted context feature Y C. Among them, the context attention process is as follows:
[0056] 3.3.1) First, the feature map is respectively input into two branches for processing. One branch divides X3 into windows of size d and performs average pooling on each window to generate superpixel features The other branch reconstructs X3 to obtain pixel features Among them,
[0057] 3.3.2) The core of the context attention mechanism lies in the iterative update of the superpixel features. The update process for the t-th iteration is as follows:
[0058] First, the superpixel features updated in the (t - 1)-th iteration are unfolded and reshaped into an appropriate shape to ensure that when calculating the affinity matrix M t , each token of the pixel feature f p only needs to be calculated with the surrounding superpixel tokens, thus reducing redundant operations. Calculate the affinity matrix M t as follows:
[0059]
[0060] Among them, Unfold(*) represents using convolution to extract the local receptive field around each superpixel, and Softmax(*) aims to obtain a normalized affinity distribution.
[0061] Then, calculate the sum of the affinity matrix for normalizing the token features and reconstruct them back to the original shape through convolution:
[0062]
[0063] Among them, Fold(*) represents using convolution to reconstruct the features.
[0064] Finally, multiply the affinity matrix M t and the pixel feature f p element-wise, and then divide by the corresponding elements of for normalization to obtain the updated superpixel features
[0065]
[0066] f m = Fold(M t × f p ) (17)
[0067] Among them, ∈ is a numerical stability constant to avoid numerical division by zero errors and reduce floating - point errors.
[0068] Fold(*) represents using convolution to reconstruct features.
[0069] 3.3.3) After n - times of iterative updates, the final super - pixel feature f s is obtained, and then self - attention calculation is performed to obtain the feature and up - sampling is performed based on the affinity matrix M obtained from the last iteration, mapping it to the size of pixel - level tokens.
[0070]
[0071] Among them, d is the dimension of the K matrix, used as a scaling factor, and M represents the affinity matrix obtained from the last iteration.
[0072] 3.4) The local features Y L extracted from the three parallel branches based on 3.1), 3.2) and 3.3), G the global features Y C and the context features Y
[0073] are concatenated by channel and fused using 1×1 convolution.
[0074] Furthermore, the sample collaborative tuning loss in step 5) is specifically:
[0075] 5.1) To align the classification scores and regression scores, this paper introduces positive / negative sample weight factors and designs a sample collaborative tuning loss. pos The positive sample weight factor ω pos reflects the importance of positive samples in classification and localization. Obviously, ω pos should be positively correlated with the classification score and the IoU value, and at the same time, the exponential function is used to widen the weight difference. Therefore, ω
[0076] ω pos = e α(s×τ) (20)
[0077] Among them, s is the classification score, τ is the IoU value, and α is a hyperparameter.
[0078] The negative sample weight factor ω negIt consists of two parts: the probability of negative samples and the importance of negative samples. When the IoU value of the anchor box is less than the threshold, it is determined as a negative sample. Thus, it can be seen that IoU is the only factor determining the probability of the anchor box being a negative sample. Since the smaller the IoU, the more we hope it is a negative sample, and the closer the IoU value is to 1, the more we hope it is a positive sample, a monotonically decreasing function is used to represent the probability of the anchor box being a negative sample. At the same time, in the inference stage, the prediction of negative samples does not affect the recall rate but affects the accuracy, that is, the classification score of negative samples is as small as possible. Therefore, ω neg is set to:
[0079] ω neg = s × (-log(τ)) (21)
[0080] where s is the classification score and τ is the IoU value.
[0081] 5.2) The sample collaborative tuning loss of this model is the weighted sum of the classification loss and the regression loss :
[0082]
[0083] where λ is the balance factor, N is the number of predicted boxes, s is the classification score, and b and b' are the positions of the predicted box and the ground truth box respectively.
[0084] The present invention has wide applications in the field of rotating object detection, such as intelligence reconnaissance, disaster monitoring, urban planning, etc. The following refers to the attached Figure 1 to describe the present invention in detail.
[0085] (1) Extract features from the remote sensing image as the initial features, which are used as the input of the frequency division module.
[0086] (2) The frequency division module processes the input features to obtain high-frequency features, low-frequency features, and complete features respectively;
[0087] (3) Use the high-frequency features, low-frequency features, and complete features as the input of the feature interaction module to extract local features, global features, and context features respectively.
[0088] (3.1) Use the position-aware attention block for the high-frequency features to extract local features.
[0089] (3.2) Use the self-focusing Transformer block for the low-frequency features to extract global features.
[0090] (3.3) Use the context Transformer block for the complete features to extract context information.
[0091] (3.4)Fuse local features, global features, and context features.
[0092] (4) Repeat steps (2) and (3) four times for multi-level information extraction from coarse to fine to obtain four-stage refined features.
[0093] (5) Use the Feature Pyramid Network (FPN) to obtain multi-scale feature representations.
[0094] (6) Align with the Region Proposal Network (RPN) and Region of Interest (RoI) to obtain the final detection features.
[0095] (7) Enter the detection head for classification and regression operations, and use the sample co-tuning loss to train the model.
[0096] The present invention provides a rotation target detection method based on frequency division feature refinement, which is applicable to rotation target detection tasks, has high detection accuracy and good robustness. Experiments show that this method can quickly and effectively perform rotation target detection.
Claims
1. A rotating target detection method based on frequency division feature refinement, characterized in that: For a given remote sensing image, perform the following operations: Step 1), extract the initial feature X from the remote sensing image of size H×W×3 as the input of the frequency division module; the feature extraction part consists of two 3×3 depth-separable convolutions; Step 2), the frequency division module processes the input features to obtain high-frequency features, low-frequency features, and complete features respectively; Step 3), based on the frequency-divided features obtained in step 2), the feature interaction module is input to enhance the multi-level information expression of the features; the feature interaction module is composed of a position-aware attention block, a self-focusing Transformer block, and a contextual Transformer block, which respectively extract local features, global features, and contextual features from high-frequency features, low-frequency features, and complete features; Then concatenate these three channels and fuse them using 1×1 convolution. Step 4), the frequency division module and feature interaction module of step 2) and step 3) are repeatedly executed four times to complete the four-stage feature refinement process; Then use the feature pyramid network FPN to extract multi-scale features, and align them with the region of interest RoI through the candidate region generation network RPN to obtain the final detection features; Step 5), finally, the detection features are input into the detection head, and the classification branch and regression branch are used to complete the category discrimination and precise positioning of the target, thereby obtaining the final detection result; The classification branch uses two fully connected layers, the regression branch uses two convolutional layers, and a sample co-tuning loss is designed to train the model.
2. The rotating target detection method based on frequency division feature refinement according to claim 1 is characterized in that: The frequency division module in step 2) is specifically: Step 2.1) Input features Divided into high-frequency features along the channel dimension and low frequency characteristics At the same time, the complete input feature As the third branch; then, feature X I ,X II and X III After deep convolution, point-by-point convolution, batch normalization and Hardswish activation function, high-frequency features, low-frequency features and complete features are obtained: X1, X2, X3; X i =H(BN(PwC(DwC(X j )))),i=1,2,3,j=I,II,III (1) Among them, DwC(*) represents depthwise convolution, PwC(*) represents point-wise convolution, BN(*) represents batch normalization layer, and H(*) represents Hardswish activation function; Step 2.2) Convolution extracts high-frequency features, Transformer extracts low-frequency features, and the complete input features are used to extract contextual information; X1, X2, and X3 are used as inputs to the position-aware attention block, the self-focused Transformer block, and the contextual Transformer block, respectively.
3. The rotating target detection method based on frequency division feature refinement according to claim 1 is characterized in that: The feature interaction module in step 3) is specifically: Step 3.1), extract local features; Design a position-aware attention block and use one-dimensional convolution to obtain attention maps of horizontal and vertical positions to enhance the local position information in high-frequency features, thereby improving the model's ability to capture spatial details. Step 3.1.1), for the input high-frequency feature X1, average pooling is performed in the horizontal and vertical directions on each channel to obtain the feature f h and f w : Step 3.1.2), respectively for feature f h and f w Perform one-dimensional convolution to enhance its horizontal and vertical position information, and then use group normalization and Sigmoid function in turn to obtain the horizontal and vertical position attention map p h and p w ; p h =σ(GN(Conv1d(f h ))) (4) p w =σ(GN(Conv1d(f w ))) (5) Among them, σ(*) represents the Sigmoid activation function, GN(*) represents group normalization, Conv1d(*) represents one-dimensional convolution, and the convolution kernel size is 7; Step 3.1.3), the horizontal position attention map p h and the vertical position attention map p w As a weight, it is multiplied with the input feature X1 to enhance its significant information in the horizontal and vertical directions respectively, and the local feature Y is obtained. L : Y L =X1*p h *p w (6) Among them, X1 is the high-frequency feature of the input, p h is the horizontal position attention map, p w is the vertical position attention map; Step 3.2), extract global features; propose a self-focusing attention mechanism and design a self-focusing Transformer block to adaptively extract global features from the input low-frequency features, which helps the model perform more accurate detection in complex scenarios; The autofocus Transformer block first performs layer normalization on the input features, then extracts features through autofocus attention, and adds the results to the input features to form a residual connection to obtain the intermediate features; the intermediate features are then processed through layer normalization and feedforward neural network, and the output results are added to the intermediate features to obtain the extracted global features Y G ; Among them, the self-focusing attention process is as follows: Step 3.2.1), project the low-frequency feature X2 after layer normalization onto the query, key and value: Q=X2W Q (7) K=X2W K (8) V=X2W V (9) Step 3.2.2), calculate the cosine similarity of Q and K, and use the Softmax function to normalize them to obtain the significance score map A QK : Among them, ||*|| represents the Euclidean norm of the feature; Step 3.2.3), using the saliency score map A QK Update K and V respectively to get the updated and Step 3.2.4), compare Q with the updated and Perform self-attention calculation: Among them, d k yes The dimension of the matrix is used as a scaling factor to prevent the dot product value from being too large, causing the gradient to disappear or explode; Step 3.3) Extract context features; introduce the context attention mechanism and design the context Transformer block to aggregate pixel information with superpixel features to extract context features; The contextual Transformer block first normalizes the input features, then extracts features through contextual attention, and adds the results to the input features to form a residual connection to obtain intermediate features; the intermediate features are then processed through batch normalization and feedforward neural network, and the output results are added to the intermediate features to obtain the extracted contextual features Y C ; Among them, the context attention process is as follows: Step 3.3.1) First, the feature map after layer normalization Input into two branches for processing respectively; one branch divides X3 into windows of size d and performs average pooling on each window to generate superpixel features The other branch reconstructs X3 to obtain pixel features in, Step 3.3.2) The core of the contextual attention mechanism is the iterative update of superpixel features. For example, the update process of the tth iteration is: First, the superpixel features after the t-1th iteration are updated Expand and reshape to the appropriate shape to ensure that the affinity matrix M is calculated t When the pixel feature f p Each token only needs to be calculated with the surrounding superpixel tokens, thereby reducing redundant operations; calculate the affinity matrix M t as follows: Among them, Unfold(*) means using convolution to extract the local receptive field around each superpixel, and Softmax(*) aims to obtain the normalized affinity distribution; Then, the sum of the affinity matrices is calculated, used to normalize the token features, and reconstructed back to the original shape through convolution: Among them, Fold(*) means using convolution to reconstruct features; Finally, the affinity matrix M t and pixel feature f p Do matrix multiplication and then Divide the corresponding elements and perform normalization to obtain the updated superpixel features f m =Fold(M t ×f p ) (17) Among them, ∈ is a numerical stability constant to avoid numerical division by zero errors and reduce floating point errors; Fold(*) means using convolution to reconstruct features; Step 3.3.3) After n iterations, the final superpixel feature f is obtained. s After that, self-attention calculation is performed to obtain the feature And upsample the affinity matrix M obtained from the last iteration and map it to the size of pixel-level tokens; Where d is the dimension of the K matrix, used as a scaling factor, and M represents the affinity matrix obtained in the last iteration; Step 3.4) extracts the local features Y based on the three parallel branches of 3.1), 3.2) and 3.3) L , global feature Y G and context feature Y C Concatenate by channel and use 1×1 convolution for fusion.
4. The rotating target detection method based on frequency division feature refinement according to claim 1 is characterized in that: The sample collaborative tuning loss in step 5) is specifically: Step 5.1), in order to align the classification score and the regression score, the positive / negative sample weight factor is introduced, and the sample co-tuning loss is designed; Positive sample weight factor ω pos Reflects the importance of positive samples in classification and positioning; ω pos Set to: oh pos =e α(s×τ) (20) Among them, s is the classification score, τ is the IoU value, and α is the hyperparameter; Negative sample weight factor ω neg It consists of two parts: the probability of negative samples and the importance of negative samples; when the IoU value of the anchor box is less than the threshold, it is judged as a negative sample. IoU is the only factor that determines the probability of the anchor box being a negative sample; a monotonically decreasing function is used to represent the probability that the anchor box is a negative sample; in the inference stage, the prediction of negative samples will not affect the recall rate but will affect the accuracy rate, that is, the classification score of negative samples should be as small as possible. neg Set to: oh neg =s×(-log(τ)) (21) Among them, s is the classification score and τ is the IoU value; Step 5.2), the sample co-tuning loss of this model is the classification loss and regression loss The weighted sum of: Among them, λ is the balance factor, N is the number of predicted boxes, s is the classification score, and b and b' are the positions of the predicted box and the true box, respectively.
Citation Information
Cited By
Image detection processing method, device and equipment
CN122049597A