Flexible object detection method based on deformable feature sensing Transform and rotation bounding box

By constructing a flexible object detection method based on deformable features perception Transformer and rotary bounding box, the problem of insufficient detection accuracy of rotary object in the prior art is solved, and high-precision detection of flexible objects with significant deformation and variable direction is achieved.

CN120510347AActive Publication Date: 2025-08-19XIAMEN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510975842.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-19
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

The existing object detection methods have accuracy bottlenecks when dealing with the detection of flexible objects in rotating objects and complex scenarios, especially in the face of significant deformation, occlusion and overlap, making it difficult to achieve high-precision detection.

Method used

Using a flexible object detection method based on deformable feature perception Transformer and rotary bounding box, a CSPDarknet53 backbone network is constructed, combined with the attention mechanism of deformable feature perception Transformer and the rotary bounding box RBB, the GIoU bounding box loss function is used for training and detection.

Benefits of technology

It realizes high-precision detection of rotating flexible objects, which is suitable for practical application requirements in complex scenarios, and improves the model's flexible object detection accuracy in significant deformation and variable direction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510347A_ABST
    Figure CN120510347A_ABST
Patent Text Reader

Abstract

The invention relates to a flexible object detection method based on deformable feature sensing Transform and a rotating bounding box, and belongs to the field of flexible object detection. The method comprises the steps that a flexible object detection model is constructed, the flexible object detection model takes CSPDarknet53 as a backbone network and fuses an attention mechanism based on deformable feature perception Transformer, and a detection head adopts a rotation bounding box (RBB) and a generalized intersection-to-union ratio (GIoU) bounding box loss function; and training a flexible object detection model by using the flexible object image data set to obtain a trained flexible object detection model, detecting a to-be-detected flexible object image in an actual scene, and outputting a flexible object detection result. According to the method, a spatial adaptive feature sampling strategy and a rotation sensing mechanism can be fully utilized, fine-grained feature representation and flexible boundary modeling capability are fused, and high-precision detection of the rotating flexible object is realized, so that the actual application requirements in a complex scene are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of flexible object detection, and in particular relates to a flexible object detection method based on a deformable feature perception Transformer and a rotating bounding box. Background Art

[0002] In the field of computer vision, the detection of flexible objects (such as seat belts and fences) faces challenges such as complex backgrounds, object deformation, and rotation. Existing target detection methods perform well in handling rotated objects and complex scenes. However, the rotation bounding box regression technology still has accuracy bottlenecks when handling rotated objects. Especially in the detection of flexible objects, due to the significant deformation of the objects, traditional horizontal bounding box or fixed angle regression methods have difficulty in accurately characterizing the target shape. In addition, problems such as occlusion, overlap, and blurring of flexible objects in complex environments make detection even more difficult. Therefore, there is an urgent need for a method that fully utilizes the spatial adaptive feature sampling strategy and the rotation perception mechanism, integrates fine-grained feature representation with flexible boundary modeling capabilities, and achieves high-precision detection of rotating flexible objects, thereby meeting the actual application needs in complex scenes. Summary of the Invention

[0003] The purpose of the present invention is to overcome the defects of the existing technology and provide a flexible object detection method based on deformable feature perception Transformer and rotating bounding box. This method is suitable for high-precision detection of flexible targets with significant deformation and changeable orientation (such as seat belts, transmission lines, and protective nets) in scenarios such as power facility inspection and industrial safety monitoring.

[0004] To achieve the above objectives, the technical solution of the present invention is: a flexible object detection method based on deformable feature perception Transformer and rotating bounding box, comprising:

[0005] Build a flexible object detection model. The flexible object detection model uses CSPDarknet53 as the backbone network and integrates the attention mechanism based on deformable feature perception Transformer. The detection head uses the rotated bounding box RBB and GIoU bounding box loss function.

[0006] A flexible object detection model is trained using a flexible object image dataset to obtain a trained flexible object detection model. The flexible object images to be detected in actual scenes are then detected, and the flexible object detection results are output.

[0007] Compared with the existing technology, the present invention has the following beneficial effects: the present invention can make full use of the spatial adaptive feature sampling strategy and the rotation perception mechanism, integrate fine-grained feature representation and flexible boundary modeling capabilities, and achieve high-precision detection of rotating flexible objects, thereby meeting the actual application needs in complex scenarios; the method of the present invention is suitable for high-precision detection of flexible targets with significant deformation and changeable directions (such as safety belts, transmission lines, and protective nets) in scenarios such as power facility inspection and industrial safety monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 This is the overall architecture of the present invention.

[0009] Figure 2 This is the 2D deformable convolution unit of the present invention.

[0010] Figure 3 is the minimum value of rectangles A and B and their containing box C.

[0011] Figure 4 Comparison of IoU and GIoU. DETAILED DESCRIPTION

[0012] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0013] The present invention provides a flexible object detection method based on a deformable feature-aware Transformer and a rotating bounding box, comprising:

[0014] Build a flexible object detection model. The flexible object detection model uses CSPDarknet53 as the backbone network and integrates the attention mechanism based on deformable feature perception Transformer. The detection head uses the rotated bounding box RBB and GIoU bounding box loss function.

[0015] A flexible object detection model is trained using a flexible object image dataset to obtain a trained flexible object detection model. The flexible object images to be detected in actual scenes are then detected, and the flexible object detection results are output.

[0016] The following is a specific implementation process of the present invention.

[0017] like Figure 1 As shown, the present invention provides a flexible object detection method based on deformable feature perception Transformer and rotating bounding box, and constructs a flexible object detection model. The flexible object detection model is specifically composed as follows:

[0018] 1. CSPDarknet53 backbone network

[0019] The flexible object detection model uses CSPDarknet53 as the backbone network architecture, combining residual structure, deep learning technology and deformable convolution mechanism to achieve efficient feature extraction. Figure 1 As shown in the figure, an improved residual connection design is used, and the training stability of the model is improved by introducing Batch Normalization and SiLU activation functions. Unlike the traditional ResNet structure, CSPDarknet53 adopts a Cross-StagePartial (CSP) structure. By dividing the feature map into two parts and processing one part, it reduces the gradient vanishing problem and accelerates the convergence of the model.

[0020] To enhance the model's ability to detect flexible objects, this design introduces a deformable convolution module into the CSPDarknet53 backbone network. This module dynamically adjusts sampling point positions, breaking the fixed receptive field limitation of traditional convolutions and enhancing the model's adaptability to deformable objects. This improvement enables the model to more accurately capture edge and deformation features when detecting flexible objects such as seat belts and fences.

[0021] The CSPDarknet53 network architecture first downsamples and compresses the input image using a Focus module (slice concatenation + 3×3 deformable convolution). The network's main structure is then constructed by stacking multiple improved residual blocks, ultimately outputting feature maps for subsequent detection tasks. Furthermore, the present invention incorporates a Spatial Pyramid Pooling (SPP) module into the middle layer of the CSPDarknet53 backbone network, enhancing multi-scale feature extraction and further improving the model's detection accuracy. By pooling feature maps at different scales, the SPP module captures spatial information at different levels of the image, thereby improving the model's generalization capabilities. The detailed structure and parameter configuration of the CSPDarknet53 backbone network are shown in Table 1.

[0022] Table 1 CSPDarknet53 backbone network parameter configuration table

[0023]

[0024] In order to improve the performance of the model in flexible object detection, the present invention adds a deformable convolution module to CBS. The deformable convolution module includes a two-dimensional deformable convolution unit, namely a 2D deformable convolution unit. Figure 2The 2D Deformable Convolution unit introduces an offset to the standard convolution, dynamically adjusting the receptive field of the convolution kernel to adapt to the position of deformed objects in the image and enhancing target perception. Furthermore, a structural reparameterization technique is introduced to improve training accuracy by adjusting the convolution kernel position during training. During inference, the offset is integrated into the standard convolution operation, allowing a single convolution kernel structure to be used, simplifying computation and improving efficiency. Furthermore, to enhance model flexibility, the 2D Deformable Convolution unit introduces a hyperparameter k to control the number of convolution kernel offset branches, improving the ability to extract complex object features. By adjusting the k value, training accuracy and inference performance can be optimized in different scenarios. Compatible with the existing CSPDarknet53 backbone network and other convolutional modules, the 2D Deformable Convolution unit significantly improves the accuracy of the model in detecting flexible objects such as rotation and deformation. Combining the 2D Deformable Convolution unit with the Transformer architecture further enhances the model's stability and efficiency when handling deformed objects and rotated bounding boxes.

[0025] 2. Attention Mechanism — Deformable Feature-Aware Transformer

[0026] In the detection of flexible objects in power scenarios, targets often have significant deformations, irregular edge contours, and occlusions, such as cables, protective nets, and safety belts. Traditional convolutional neural networks (CNNs) are limited by local fields of view and have difficulty capturing long-range features, resulting in missed or misdetected targets such as curved cables. In addition, due to the complex grid structure and multi-angle distribution of flexible protective nets, conventional detection algorithms have difficulty effectively modeling the overall structure. This paper proposes a detection method based on an attention mechanism of a deformable feature-aware Transformer, which significantly improves the detection accuracy and robustness of flexible objects through long-range feature modeling and dynamic target querying.

[0027] The Deformable Feature-Aware Transformer employs an encoder-decoder architecture, centered around a dynamic deformation modeling mechanism. This architecture achieves efficient representation of flexible objects through a spatially adaptive feature sampling strategy. During the encoding phase, this architecture extracts basic deformation features through deformable convolutional layers and generates a dynamic set of sampling points for each query location, enabling cross-scale feature fusion. Specifically, a redundant sampling strategy is established for objects such as curved cables to enhance feature integrity in occluded scenarios. During the decoding phase, this method employs a deformation-sensitive object query mechanism with a rotation prior. This method progressively refines bounding box parameters through multiple rounds of feature interaction, while simultaneously combining deep and shallow features for multi-granular feature alignment, capturing both the macroscopic outline and microstructure of the object. This design significantly improves the model's deformation adaptability and angular sensitivity to flexible objects, enabling sampling points to dynamically track the target's deformation trajectory. A rotation angle classifier is used to accurately detect orientation within a ±90° range, maintaining stable detection performance in complex scenarios.

[0028] The core innovation of the Deformable Feature-Aware Transformer lies in its multi-scale dynamic deformation modeling mechanism, which achieves efficient feature interaction within the local receptive field through a spatially adaptive sampling strategy. This architecture uses a multi-scale feature fusion design in the encoding phase, inputting the feature map output by the backbone network into the deformation-aware encoder. At each query position, a dynamic sampling module generates K spatially adaptive sampling points, which automatically adjust their spatial distribution based on the target's deformation characteristics. The entire feature aggregation process uses an improved attention score calculation method as follows:

[0029]

[0030] in: is the feature vector of the query position, K is the number of sampling points for each query position, Represents attention weight and transformation weight respectively is the feature vector at the sampling point. The deformable attention mechanism in the deformable feature-aware Transformer operates at different scales of the feature map to simultaneously capture the global outline information and local details of the target. The model generates an independent set of sampling points at each scale, improving detection accuracy through multi-scale fusion.

[0031] In the decoding phase, the model uses a dynamic target query mechanism to update the target bounding box and category through multiple autoregressive steps. In each iteration, the decoder calculates the attention weight between the target query and the feature map and generates a new bounding box prediction result. The calculation formula of the attention weight with the feature map F is:

[0032]

[0033] Where: d is the feature dimension, and Softmax ensures that the attention weights are normalized at all sampling points.

[0034] The Deformable Feature-Aware Transformer, through its unique dynamic deformation modeling mechanism, provides a new solution for flexible object detection. This architecture employs a multi-scale feature interaction strategy in the encoding phase, using spatially adaptive sampling points to achieve precise feature extraction within the local receptive field. This allows it to capture both the overall outline of the target and subtle deformation features. In the decoding phase, a dynamic query mechanism is used to iteratively optimize detection results, enabling the model to adapt to the needs of object detection in a variety of complex scenarios. This demonstrates the Transformer architecture's strong adaptability and scalability in the field of industrial vision.

[0035] The deformable feature-aware Transformer improves feature quality through the attention mechanism, assisting the detection head to generate rotated bounding boxes more accurately.

[0036] 3. Design of detection head

[0037] In the flexible object detection task, the detection head is responsible for performing target box regression and classification prediction on the feature map. However, the horizontal bounding box (HBB) and IoU (Intersection over Union) loss functions used by traditional detection heads have the following limitations when dealing with flexible objects:

[0038] First, traditional HBB restricts the target box to a horizontal rectangular box, which can only represent regular objects such as boxes and workpieces. However, in flexible object detection scenarios, such as targets such as safety belts, fences, or hanging ropes in power scenarios, their shapes often have large bends or tilts. HBB cannot accurately fit the edges of objects, resulting in the bounding box enclosing a large amount of background area and reduced detection accuracy. Secondly, the IoU loss function only calculates the intersection over union (IoU) of the predicted box and the true box, ignoring the distance information between the boxes. When the predicted box does not overlap with the true box, the IoU value is directly 0, which cannot provide an effective optimization gradient, making it difficult for the model to converge.

[0039] (1) Rotating bounding box to select flexible objects

[0040] In the flexible object detection task, this design uses a rotated bounding box (RBB) to replace the traditional horizontal bounding box (HBB) to more accurately describe flexible targets with variable orientations. This is highly related to the proposed rotated bounding box method based on long-edge representation and classification discretization.

[0041] To more accurately learn and represent the directional information of arbitrarily oriented flexible objects in images, a long-side representation (LSR) method is employed. This method, which describes a rotated bounding box (RBB), increases the bounding box's degrees of freedom, enabling the model to more accurately capture and represent the large spans, deformable shapes, and small foreground-to-background ratios of flexible objects. This method is particularly effective in detecting arbitrarily oriented flexible objects. It allows for simple and direct learning of object directional properties and is particularly well-suited for detecting flexible objects with large spans, deformable shapes, and small foreground-to-background ratios.

[0042] The long-edge representation increases the degrees of freedom (DOF) of the bounding box, enabling the model to accurately capture the morphological features of flexible objects. In this method, the bounding box angle θ is defined as the angle between the longest side and the x-axis, with an angle range of [-90°, 90°]. This differs from the OpenCV (Open Source Computer Vision Library) representation, where angles are represented counterclockwise. When the long side is parallel to the x-axis, the angle is 0°. As the rotation continues counterclockwise, the angle increases, ultimately reaching -90°. Therefore, the angle range of the latter is [-90°, 0°]. Because the long-edge representation fixes the position of the longest side, it avoids the length-width swapping problem that exists in the OpenCV representation, simplifying the computational complexity of the regression process.

[0043] A new strategy is proposed to address boundary issues: In the long-side representation, the angle θ is defined as the angle between the longest side and the x-axis, and its range is [-90°, 90°]. While this representation avoids the length-width interchange problem present in the OpenCV representation, it can lead to periodicity in the angle parameters in some cases. Specifically, when an object is rotated close to a boundary value (such as -90° or 90°), the angle parameter may suddenly change, causing problems in the numerical regression process. For example, when the angle of an object changes from 89° to 91°, it will appear numerically as a jump from 89° to -89°, which can cause instability and exploding or vanishing gradients during training.

[0044] In order to overcome the boundary problem existing in the long edge representation, the present invention proposes two innovative strategies: classification discretization and bipolar symmetric mapping.

[0045] ① Classification discretization

[0046] To address the periodicity of the angle parameters in the long-side representation and the angle mutations that may occur during the numerical regression process, a classification discretization method is used. This method converts the original angle regression task into a classification task by dividing the angle parameters into 180 equal parts, each representing an angle range of 1 degree. The advantage of this is that it avoids the mutations that occur when the angle approaches the boundary value. When the angle changes, the excess angle can be correctly classified into the adjacent angle category.

[0047] This approach prevents angle parameters from exceeding their limits and prevents gradient explosion or vanishing due to sudden angle changes during training. It also simplifies the numerical calculation process and improves the stability of model training and detection accuracy. Furthermore, this approach allows for the design of more appropriate loss functions for angle classification, further optimizing the overall performance of the model.

[0048] ② Bipolar symmetric mapping

[0049] In order to solve the periodicity and boundary mutation problems in the parameterization process of the rotation angle, the bipolar symmetric mapping method is used. This method reparameterizes the original angle parameter space to form a representation with bipolar symmetry characteristics, ensuring that the angle change maintains continuity and smooth transition throughout the entire range. For key angles close to 0°, 180°, and 360°, this mapping can be used to convert them into a more stable representation, making the numerical distance between adjacent angles more reasonable and consistent. The bipolar symmetric mapping function is expressed as follows:

[0050]

[0051] This mapping method not only solves the periodicity problem of the angle parameter but also implicitly implements the symmetry constraint of the long-side representation through absolute value operations, effectively avoiding training instability caused by sudden changes in the angle parameter. It also improves the model's ability to handle edge cases, enhancing overall detection accuracy and generalization capabilities. This makes the training process more stable and efficient.

[0052] (2) GIoU loss function design

[0053] In this invention, the GIoU loss function not only considers the overlapping area between the predicted box and the true box, but also introduces their minimum enclosing area. This comprehensive consideration enables GIoU to more comprehensively evaluate the difference between the two, rather than just focusing on the overlapping part. For example, when the predicted box and the true box do not intersect, the traditional IoU loss function cannot provide an effective optimization direction, but GIoU can still guide the predicted box to move closer to the true box by introducing the minimum enclosing area. This achieves a comprehensive measurement of the difference.

[0054] During training, the GIoU loss function provides the model with a clearer optimization objective. It not only encourages the predicted box to overlap as much as possible with the ground-truth box, but also forces the predicted box to surround the ground-truth object as tightly as possible, thereby improving positioning accuracy. This optimization approach enables the model to more precisely learn the object's position and shape, thereby improving overall detection accuracy.

[0055] In order to adjust the predicted box to be as close as possible to the ground-truth box, the GIoU loss function can effectively optimize both aspects by considering the position and size of the predicted and ground-truth boxes. It not only focuses on whether the center positions of the predicted box and the ground-truth box are aligned, but also whether their aspect ratios match, making the predicted box more accurately close to the ground-truth box in both position and size.

[0056] In the multi-task joint training framework, the GIoU loss function is weighted and summed with the loss functions of other tasks (such as angle classification loss and object confidence prediction loss) to form the final joint loss function. By optimizing the joint loss function, end-to-end optimization of multiple tasks is achieved, enabling the model to achieve good performance across different tasks. In each training iteration, the GIoU loss between the predicted box and the ground-truth box is calculated, and the gradient is calculated through backpropagation. Based on this gradient information, the model parameters are updated, adjusting the position and size of the predicted box to bring it closer to the ground-truth box. Through continuous iterative training, the model gradually learns how to accurately predict the position and size of objects.

[0057] In flexible object detection, the traditional Intersection over Union (IoU) loss function performs poorly when the predicted and ground-truth boxes do not overlap. This is because when the two boxes do not intersect, the IoU returns 0, causing the gradient to vanish and failing to provide effective optimization feedback for the model. This is particularly true in rotated object detection, where even a slight change in the rotation angle can cause the two boxes to completely disjoint. Furthermore, while RBB can capture the object's orientation, relying solely on it may not be sufficient to accurately assess the similarity between the predicted and ground-truth boxes. This is especially true for flexible objects with varying shapes and orientations, where simple RBB localization may not be accurate enough. Due to their deformable nature, flexible objects often have a wide range of aspect ratio variations. Relying solely on RBB without additional mechanisms to adapt to these variations can lead to degraded detection performance.

[0058] For two rectangles A and B and their minimum containing box C, the calculation method of IoU is as follows, Figure 3 And the following formula:

[0059]

[0060] When IoU is used directly as the objective function, when the predicted box and the ground truth bounding box do not intersect, the IoU is 0 regardless of the distance between them. This causes the back propagation gradient to be 0, making the model unable to learn.

[0061] In order to solve the above problem, the GIoU method is used, which solves the problem that IoU is 0 when the predicted box and the real box do not intersect. Unlike IoU which only focuses on the overlapping area, GIoU also considers the overlapping area as well as Figure 3 The two non-overlapping parts d1 and d2 shown in . When rectangles A and B do not intersect, the farther the distance between them, the closer the GIoU value is to -1. To better describe the degree of overlap between rectangular boxes, the distance between rectangular boxes A and B can also be determined. In addition, GIoU retains the scale invariance of IoU, and the degree of overlap between rectangular boxes A and B is independent of their spatial scale. Therefore, the value of GIoU is highly correlated with the objective function, and multiple boxes with the same degree of overlap but different bounding boxes will have different loss values. Let C be the minimum closed area containing A and B, and the calculation formula of GIoU is as follows:

[0062]

[0063] On the other hand, the present invention also considers the angle of label assignment. Since the pixels occupied by slender and flexible objects in the image are very small in the horizontal or vertical direction, such as Figure 4 As shown, IoU and GIoU have different sensitivities. Figure 4 In this figure, A represents the bounding box of the seatbelt, and B and B' represent the predicted bounding boxes. They are offset diagonally by 1 and 2 pixels, respectively. Their IoU and GIoU values are calculated as follows: IoU(A, B)=0.45, IoU(A, B')=0.19, GIoU(A, B)=0.42, and GIoU(A, B')=0.06. Compared to IoU, GIoU is more sensitive to bounding box offsets for slender and flexible objects. Therefore, using GIoU during model training increases the bounding box position offset gradient and the bounding box position fine-tuning weight.

[0064] This paper studies the traditional rotation IoU loss function for rotatable bounding boxes. The rotatable bounding box adds a rotation angle parameter to the horizontal bounding box, which enables it to better fit the boundaries of the object. The calculation formula of the rotation IoU loss function is as follows:

[0065]

[0066]

[0067] Among them, the rotatable bounding box and The angles are and , is the angle-aware intersection-over-union ratio, is the predicted rotated bounding box, is the true rotation bounding box, It is the angle-aware intersection-over-union ratio after the angle is periodically corrected by 180°. Area represents the area of the region. Predicting rotated bounding boxes and the true rotation bounding box The area of the intersection, Predicting rotated bounding boxes and the true rotation bounding box The area of the merged region (union). Rotatable bounding box With Same center point location and size, but with Same angle. As the angle difference between the two bounding boxes increases from 0° to 90°, The value of the function decreases. When the angle difference between the two rotatable bounding boxes approaches 180°, The function ignores the head and tail angles of the object, making it impossible to distinguish the head and tail of the object.

[0068] In summary, the present invention uses the GIoU bounding box regression loss function as the evaluation method, which is expressed as follows:

[0069]

[0070] in, is the prediction box, The GIoU bounding box is used to replace the original smooth L1 distance loss, thereby unifying the training objective function and the evaluation function to achieve more accurate anchor box positioning.

[0071] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A flexible object detection method based on deformable feature perception Transformer and rotating bounding box, characterized in that: include: Build a flexible object detection model. The flexible object detection model uses CSPDarknet53 as the backbone network and integrates the attention mechanism based on the deformable feature-aware Transformer. The detection head uses the rotated bounding box (RBB) and the generalized intersection-over-union (GIoU) bounding box loss function. A flexible object detection model is trained using a flexible object image dataset to obtain a trained flexible object detection model. The flexible object images to be detected in actual scenes are then detected, and the flexible object detection results are output.

2. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 1 is characterized in that: CSPDarknet53 combines an improved residual structure, deep learning technology and deformable convolution mechanism to achieve feature extraction of flexible object images; batch normalization and SiLU activation function are introduced in the improved residual structure.

3. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 2 is characterized in that: CSPDarknet53 adopts the CSP structure, including the Focus layer, DCBS layer, DCBS layer, CSP1_X layer, DCBS layer, CSP1_X layer, DCBS layer and SPP layer connected in sequence from input to output, where the DCBS layer is transformed by the CBS layer using a deformable convolution module as the convolution module.

4. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 3 is characterized in that: The deformable convolution module includes a two-dimensional deformable convolution unit. The two-dimensional deformable convolution unit introduces an offset on the basis of standard convolution to dynamically adjust the receptive field of the convolution kernel. The two-dimensional deformable convolution unit also introduces structural reparameterization technology, that is, adjusting the position of the convolution kernel during training. In the inference stage, the offset is integrated into the standard convolution operation, so that a single convolution kernel structure is used during inference.

5. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 1 is characterized in that: The deformable feature-aware Transformer adopts an encoder-decoder architecture, the core of which is the multi-scale dynamic deformation modeling mechanism, which achieves efficient representation of flexible object targets through a spatially adaptive feature sampling strategy.

6. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 5 is characterized in that: The encoder-decoder architecture uses a multi-scale feature fusion design in the encoding stage. The feature map output by the backbone network is input into the deformation-aware encoder. At each query position, a dynamic sampling module generates K spatially adaptive sampling points. These sampling points can automatically adjust their spatial distribution according to the target deformation characteristics. The entire feature aggregation process uses an improved attention score calculation method as follows: in: is the feature vector of the query position, K is the number of sampling points for each query position, denote the attention weight and transformation weight respectively, is the eigenvector at the sampling point; The encoder-decoder architecture uses a dynamic target query mechanism in the decoding stage to update the target bounding box and category through multiple autoregressive steps. In each iteration, the decoder calculates the attention weight between the target query and the feature map and generates a new bounding box prediction result. The calculation formula of the attention weight with the feature map F is: Where: d is the feature dimension, Softmax is the activation function, and it ensures that the attention weights are normalized at all sampling points.

7. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 1 is characterized in that: The long edge representation is used in the detection head to describe the rotated bounding box RBB, and a classification discretization strategy and a bipolar symmetric mapping strategy are introduced.

8. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 7 is characterized in that: In the long-side representation, the angle θ is defined as the angle between the longest side and the x-axis, and its range is [-90°, 90°]. The classification discretization strategy converts the angle regression task into a classification task by dividing the angle parameter into 180 equal parts, each representing an angle range of 1 degree. The bipolar symmetric mapping strategy reparameterizes the angle parameter space to form a representation with bipolar symmetry characteristics, ensuring that the angle changes maintain continuity and smooth transition throughout the entire range.

9. The flexible object detection method based on deformable feature perception Transformer and rotating bounding box according to claim 1 is characterized in that: The generalized intersection-over-union (GIoU) bounding box loss function is calculated as follows: in, is the prediction box, is the real frame, generalized intersection and union It is expressed as follows: C is included and The minimum closed area, IoU express and The intersection and union ratio.

Citation Information

Patent Citations

  • Detection method for flexible object in uncertain direction, electronic equipment and storage medium

    CN114241403A

  • Workpiece category and pose estimation method based on YOLOv4-tiny model

    CN115100136A

  • Complex environment power transmission line foreign matter detection method based on deep learning

    CN118506263A