End-to-end circular identification point automatic identification method for high-speed video

By adding a multi-scale direction fusion attention module and angle detection head to the YOLOv8 network model, the problem of accurate extraction of circular mark point recognition in complex scenarios is solved, and high-precision and efficient automatic recognition effect is achieved.

CN120107970APending Publication Date: 2025-06-06SHANGHAI OCEAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063542.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In complex scenarios, traditional methods have difficulty accurately and robustly identifying circular marking points in high-speed videos, especially in the presence of occlusion, light interference, and backplane reflection.

Method used

Based on the YOLOv8 network model, a multi-scale direction fusion attention module MOA and angle detection head are added, and a mixed loss function is used for training to achieve end-to-end automatic identification of circular mark points.

Benefits of technology

It improves the accuracy and efficiency of circular marking points recognition in complex scenarios, realizes automatic detection, reduces the cost of manual identification, and improves the positioning accuracy of the detection frame.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107970A_ABST
    Figure CN120107970A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end circular identification point automatic identification method for a high-speed video, and the method comprises the steps: building a detection network model which is based on a YOLOv8 network model, adding a plurality of multi-scale direction fusion attention modules MOA in a Neck neck network of the detection network model, adding an angle detection head in a Head detection head network of the detection network model, and carrying out the detection of the angle detection head. Carrying out loss calculation by adopting a mixed loss function L; respectively marking circular identification points in each image in the image sample set, including the size and the angle of each detection frame, and then training a detection network model by taking the marked images as input; and automatically identifying each circular identification point in each frame of image in the high-speed video by using the trained detection network model. By adopting the automatic identification method, the extraction precision of the mark points in a complex scene is improved, and the application range is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and in particular relates to a method for automatically identifying circular marking points in end-to-end high-speed video. Background Art

[0002] With the continuous development of high-resolution digital camera technology and advanced computing models, high-speed video measurement technology combining visual photogrammetry technology with advanced computer vision technology has received more and more attention. High-speed video measurement technology has been widely used in structural health monitoring, collapse experiments, environmental testing and other fields. In order to capture the motion state and trajectory of high-speed moving objects, it is usually necessary to set markers on the objects to be measured. Therefore, quickly and accurately identifying the markers and their center coordinates in massive image sequences has become a key issue in the development of high-speed video measurement technology.

[0003] Commonly used landmarks include petal-shaped, cross-shaped, circular landmarks, etc. Among them, the circular landmark is the most commonly used artificial landmark in video measurement. The traditional landmark contour detection process includes image binarization, edge extraction, contour fitting, and circle center extraction. For example, Ho et al. proposed a fast ellipse detector based on global geometric symmetry. Based on geometric symmetry, this method first locates candidates for ellipse and circle centers. At the same time, according to these candidate centers, all feature points in the input image are grouped into several sub-images. Then, for each sub-image, geometric symmetry is used again to obtain all ellipses and circles. Chia et al. proposed a novel contour-based method to identify object categories using line segments and ellipses. Prasad et al. proposed an ellipse detection method that uses edge curvature and relative convexity as grouping criteria, and uses 2D Hough transform and "relationship score" to rank edge contours within the group. Fornaciari et al. proposed a method that selects arcs belonging to the same ellipse at the arc level (rather than the pixel level) by relying on a selection strategy, and then estimates the ellipse center using the midpoint characteristics of parallel chords, and accumulates votes in the decomposed parameter space to estimate the remaining parameters, etc. However, in complex real environments, the presence of circular signs of various scales, occlusion, light interference, backplane reflection and other factors will reduce the robustness and accuracy of circular sign detection by conventional image processing methods. Traditional methods based on image edge features and geometric properties cannot adapt well to complex experimental scenes. Summary of the invention

[0004] The present invention provides a method for automatically identifying circular marking points in end-to-end high-speed video, which solves the technical problems of difficulty in accurately extracting marking points in complex scenes.

[0005] The present invention can be achieved through the following technical solutions:

[0006] A method for automatically identifying circular marker points in end-to-end high-speed video includes the following steps:

[0007] Step 1: Establish a detection network model. The detection network model is based on the YOLOv8 network model, adds multiple multi-scale directional fusion attention modules MOA to its Neck network, adds an angle detection head to its Head detection network, and uses a mixed loss function L for loss calculation;

[0008] The multi-scale directional fusion attention module MOA is provided with two, which are respectively connected between the second upsampling layer Upsample and the fusion layer Concat of the Neck network and between the third upsampling layer Upsample and the fusion layer Concat, and include a coordinate perception branch, a multi-scale perception branch and a cross-space information aggregation branch. The coordinate perception branch performs feature extraction along the horizontal and vertical directions in space to obtain a position attention descriptor. The multi-scale perception branch uses multiple depth-separable convolutional layers DW with different kernel sizes to extract features to obtain a global attention descriptor. The cross-space information aggregation branch performs feature aggregation on the position attention description that conforms to the global attention descriptor to obtain a spatial position feature map.

[0009] Step 2: Label the circular identification points in each image in the image sample set, including the size and angle of each detection frame, and then use the labeled images as input to train the detection network model;

[0010] Step 3: Use the trained detection network model to automatically identify each circular marker point in each frame image in the high-speed video.

[0011] Furthermore, the output image of the Backbone network is recorded as R C*H*W , the coordinate perception branch is used for the input image X∈R C*H*W Average pooling is performed along the horizontal and vertical directions of the space, and then the two images are spliced ​​together in the spatial dimension. Then, the intermediate feature map F is obtained through the Conv layer transformation, and then the feature map F is divided into two tensors represented as Fx and Fy along the horizontal and vertical directions. Then, the reshape operation is used to adjust the dimensions of the two tensors Fx and Fy, and the sigmoid gating function is applied to obtain the corresponding weights Wx and Wy. Finally, the weights Wx and Wy are used to adjust the input image X through the dot product operation to obtain the position attention descriptor Xc;

[0012] The multi-scale perception branch uses two depth-wise separable convolutional layers DW with different kernels to sequentially process the input image X∈R C*H*WPerform a convolution operation and then perform the feature map output by the two convolution operations The scale feature map A is obtained by splicing, and then the scale feature map A is subjected to global average pooling and global maximum pooling respectively. avg , A max The sigmoid gating mechanism is applied to obtain the corresponding weights, and finally the weights are used to perform weighted summation with the feature map after the depth-separable convolutional layer DW to obtain the global attention descriptor A. final .

[0013] Furthermore, the cross-spatial information aggregation branch takes the scale feature map A and the position attention descriptor Xc as input, performs adaptive global pooling operations on them respectively, obtains the corresponding channel descriptors, and then performs matrix multiplication on the channel descriptors with the corresponding scale feature map A and position attention descriptor Xc to generate two global spatial attention feature maps Y 1 , Y 2 , and then the two global spatial attention feature maps Y 1 , Y 2 Aggregate to generate spatial weights, and finally apply the generated spatial weights to the input feature map X and combine it with the global attention descriptor A final , and obtain the final spatial position feature map O.

[0014] Further, the global attention descriptor A is calculated using the following formula: final ;

[0015]

[0016] The spatial position feature map O is calculated using the following formula.

[0017]

[0018] Furthermore, the angle detection head is configured to add two channels to the original detection head of the YOLOv8 network model for outputting the angle prediction coding values ​​cos(ωθ) and sin(ωθ), and limiting them to the range of [-1,1].

[0019] Then use the following equation to calculate the angle θ,

[0020]

[0021] Among them, z represents the code value, ω represents the angular frequency, and mod represents the remainder operation. ω≤2.

[0022] Furthermore, during training, the predicted angle of the detection box is encoded using the following formula:

[0023]

[0024] Among them, θ box / θ obj ∈[0,2π) represent the predicted angle and target angle of the detection box respectively, and ω=2. Further, the hybrid loss function L is calculated using the following formula;

[0025] L acm = l smoothl1 (f p ,f g )

[0026] L box = l probiou (B(x p ,y p ,a p ,b p ,θ p ),B(x g ,y g ,a g ,b g ,θ g ))

[0027] L=λ box L box +λ acm L acm

[0028] Among them, λ box Set to 1, λ acm Set to 0.13, B(x p ,y p ,a p ,b p ,θ p ),B(x g ,y g ,a g ,b g ,θ g ) are the parameter values ​​of the detection box and the target box, respectively, f p ,f g They respectively represent the angle encoding components output by the detection network model and their corresponding true values.

[0029] The beneficial technical effects of the present invention are:

[0030] (1) Based on the object detection network YOLOv8, a near-real-time, high-precision end-to-end circular landmark recognition network Ellipse-YOLO is proposed. It solves the difficulty of accurately extracting landmark points in complex scenes such as indoor low light, outdoor overexposure, close-range measurement and long-range measurement, and realizes automatic detection of circular landmark point areas, eliminating the labor cost of manual marking.

[0031] (2) A multi-scale directional fusion attention module is proposed. First, the input features are decomposed into a position-aware branch and a multi-scale-aware branch. The position-aware branch embeds specific direction information, captures long-range dependencies, and establishes accurate position information. At the same time, the multi-scale-aware branch groups channel dimensions, evenly distributes spatial semantic features, and obtains a global receptive field. In addition, a feature interaction module Matmul Block is introduced to aggregate the information of the position-aware branch and the multi-scale-aware branch through cross-dimensional interaction through cross-spatial information aggregation branches, capture pixel-level pairwise relationships, and achieve rich feature aggregation. This method effectively alleviates the problem of information loss of landmark targets in complex backgrounds.

[0032] (3) The original detection head in the network structure of the target detection network YOLOv8 is redesigned to add an angle prediction head, and an angle enhancement coding prediction head is introduced to fit multi-angle circular or elliptical landmarks to solve the problem that the traditional target detection network cannot handle target angle prediction. In addition, the boundary discontinuity problem caused by severe mutations in angle prediction is solved by combining the encoding method of the angle smoothing function with the loss function based on Gaussian representation.

[0033] (4) Since the present invention adds the angle prediction of the border, the RobIou loss function is introduced, and Gaussian modeling is performed on the output information of the target detection network. A single Gaussian model is used to predict the uncertainty of the coordinates, width and height of the center point of the border respectively, so as to measure the reliability of the predicted detection frame. The new detection head contains the probability information of the predicted detection frame. The introduction of probability information can better understand the reliability of the predicted frame and improve the positioning accuracy of the border. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0035] Figure 2 It is a schematic diagram of the structure of the target detection network Ellipse-YOLO of the present invention;

[0036] Figure 3 It is a schematic diagram of the structure of the improved YOLOV8 network of the present invention;

[0037] Figure 4 It is a schematic diagram of the structure of the multi-scale attention module MOA of the present invention;

[0038] Figure 5 Schematic diagram of angle coding of the present invention, (a) is a schematic diagram showing the situation where the predicted angle and the target angle produce a boundary, and (b) is a functional form of the model fitting the target angle;

[0039] Figure 6 This is a schematic diagram of detection and positioning results using the automatic recognition method of the present invention. DETAILED DESCRIPTION

[0040] The specific implementation of the present invention is described in detail below with reference to the accompanying drawings and preferred embodiments.

[0041] like Figure 1 As shown, the present invention provides an automatic recognition method for circular markers in high-speed videos, which is mainly based on the YOLOv8 network model to build a detection network model Ellipse-YOLO with deep learning function. The model directly predicts the center coordinates of the ellipse corresponding to the circular marker, the major and minor axes of the ellipse, and the angle of the ellipse to identify the circular marker to be detected, and has efficient feature information acquisition capabilities and high-speed target detection capabilities. Compared with the existing model, the detection network model of the present invention significantly improves the accuracy and efficiency of circular marker recognition. The details are as follows:

[0042] Step 1: Establish a detection network model, such as Figure 2-3 As shown in the figure, the detection network model is based on the YOLOv8 network model, adds multiple multi-scale direction fusion attention modules MOA to its Neck network, adds an angle detection head to its Head detection network, and uses a mixed loss function L for loss calculation.

[0043] 1. Construct a multi-scale directional fusion attention module MOA

[0044] Due to the high sensitivity of rotation target detection to target coordinates and orientation, the relatively small size of landmarks in high-resolution data, and the complex background environment, it is easy to cause the target detection accuracy to decrease and the false detection rate to increase. Therefore, it is necessary to further enhance the network's ability to perceive directions and multi-scales. CA (Coordinate Attention) captures spatial information by performing global average pooling along the horizontal and vertical directions, embeds position information into channels, and captures long-range interactions in space. The two one-dimensional encoding branches obtain position information along their respective spatial dimensions and capture long-range interactions in different dimensional directions. However, the CA attention mechanism mainly relies on information aggregation in the horizontal and vertical directions. Although this method can retain spatial position information to a certain extent, it cannot fully capture complex nonlinear feature relationships. Especially for images with highly nonlinear and complex backgrounds, this simple aggregation of directional information may not be sufficient to describe the relationship between targets. In addition, CA has limited performance in capturing long-range dependencies. Although it can retain position information to a certain extent, it may not be as effective as the global attention mechanism for long-range feature interactions across a large spatial range. This limitation is particularly evident when dealing with tasks that require long-range dependencies, such as target detection in complex scenes.

[0045] Inspired by the parallel branch structure of CA, we redesigned the multi-scale directional attention module MOA, the structure of which is as follows: Figure 4 As shown in the figure, MOA uses three parallel routes to extract the attention descriptor of the feature map. Two of the parallel branches are located in the coordinate perception branch, and the third parallel branch is located in the multi-scale perception branch. Finally, in order to capture the dependencies between all channels and reduce the computational budget, we use the cross-spatial information aggregation branch to model the cross-channel information interaction in the channel direction. We use parallel substructures to avoid too much sequential processing and too large module depth. At the same time, based on the idea of ​​RepLKNet, we add a large core subnet to dynamically adjust the receptive field of the feature extraction backbone to more effectively process the contextual information required by the detected target.

[0046] That is to say, the multi-scale directional fusion attention module MOA is set with two, which are respectively connected between the second upsampling layer Upsample and the fusion layer Concat of the Neck network and between the third upsampling layer Upsample and the fusion layer Concat, as shown in Figure 4As shown, it includes a coordinate perception branch, a multi-scale perception branch and a cross-spatial information aggregation branch. The coordinate perception branch performs feature extraction along the horizontal and vertical directions in space to obtain a position attention descriptor. The multi-scale perception branch uses multiple depth-separable convolutional layers DW with different kernel sizes to extract features to obtain a global attention descriptor. The cross-spatial information aggregation branch performs feature aggregation on the position attention description that conforms to the global attention descriptor to obtain a spatial position feature map.

[0047] The two parallel branches in the coordinate perception branch perform one-dimensional encoding along the horizontal and vertical directions of space to capture long-distance dependencies in the horizontal direction and retain the positional relationship in the vertical direction, which helps the network to more accurately locate the object of interest, obtain the global receptive field and encode accurate position information. Finally, MOA will output the feature map enhanced by the coordinate attention and multi-scale spatial attention mechanism as the input of the subsequent feature fusion stage.

[0048] Specifically, let the output feature map of the Backbone network be Where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively. First, the input feature map X is average pooled along the horizontal and vertical directions of the space to obtain two pooled feature maps:

[0049]

[0050] Next, the two feature maps obtained are concatenated in the spatial dimension:

[0051]

[0052] Then, the concatenated feature map is transformed through a 1×1 convolutional layer to obtain the intermediate feature map F:

[0053]

[0054] Among them, d is the number of channels after dimensionality reduction, and reduction is the dimensionality reduction coefficient. Then, the feature map F is split into two tensors along the horizontal and vertical directions:

[0055] F x ,F y =Split(F,[H,W])

[0056] The dimensions of these two tensors are adjusted respectively, and the corresponding weights are generated through 1×1 convolution and Sigmoid activation function:

[0057]

[0058] Finally, using the weight Wx and W y , adjust the input feature map X through element-by-element multiplication operation to obtain the position attention descriptor X c :

[0059] X c =X⊙W x ⊙W y

[0060] In the multi-scale feature perception module, we fully utilize the idea of ​​large kernel convolution to obtain a larger effective receptive field with only a small inference delay. First, we use a large kernel convolution module based on depth wise (DW) convolution with multiple kernel sizes, i.e., depthwise separable convolution layer DW, to obtain feature maps of multiple receptive fields; since DW convolution only performs convolution in the spatial dimension, the amount of calculation and parameters are greatly reduced compared to ordinary convolution. Therefore, in this way, the dimension of the feature map is expanded and more contextual information and detailed features are captured, further enriching the information content of the feature map while reducing the performance loss caused by large convolution kernels.

[0061] Specifically, let the kernel size of the i-th depth-wise separable convolutional layer be k i And the expansion rate is d i , whose output feature is represented by A i :

[0062]

[0063] Assume that there are N = 2 DW convolutional layers, using 5×5 and 7×7 convolution kernels respectively:

[0064] A 1 =Conv 5×5 (X),A 2 =Conv 7×7,dilation=3 (A 1 )

[0065] Since DW convolution only performs convolution in the spatial dimension, the amount of calculation and parameters is greatly reduced compared to ordinary convolution. Next, the output of each DW convolution layer is modeled and dimensionally reduced through a 1×1 convolution layer to obtain the feature map

[0066]

[0067] The feature maps output by the two convolution operations are concatenated in the channel dimension to obtain the scale feature map A:

[0068]

[0069] Then, the scale feature map A is subjected to global average pooling and global maximum pooling operations respectively:

[0070]

[0071] In order to allow different spatial descriptors to interact with each other, the features after global average pooling and maximum pooling are concatenated in the channel dimension and converted into a spatial attention map with N = 2 channels through a convolutional layer. The Sigmoid gating function is then applied to obtain the weighted representation corresponding to each scale feature:

[0072]

[0073] Split ∑ along the channel dimension (i.e. the second dimension) into two independent weight graphs ∑ 1 ,∑ 2 , and then the generated weight ∑ 1 ,∑ 2 Features at different scales Perform point multiplication and weighted summation, and restore the number of channels through a 1×1 convolutional layer to obtain the final global attention descriptor A final :

[0074]

[0075] The cross-spatial information aggregation branch achieves richer feature aggregation by establishing interdependencies between channels and spatial positions. The specific steps are as follows:

[0076] First, the scale feature map A in the multi-scale perception branch is combined with the position feature descriptor X in the coordinate perception branch. c Through two-dimensional global average pooling and Softmax operation, we finally get their normalized channel descriptors:

[0077] S 1 =Softmax(AGP(A))

[0078] S 2 =Softmax(AGP(X c ))

[0079] AGP stands for Adaptive Global Pooling, which compresses the feature map to a size of 1×1. Next, matrix multiplication is performed between the channel descriptor and the corresponding feature map to generate two global spatial attention feature maps:

[0080]

[0081] The generated spatial attention feature maps are aggregated and the final spatial weight representation is generated through the Sigmoid activation function:

[0082]

[0083] Finally, the generated spatial weights are applied to the input feature map and combined with the global attention descriptor A generated by the large kernel spatial attention. final , get the final spatial position feature map:

[0084]

[0085] Among them, the two ⊙ represent the attention weighting of two different dimensions, which is the Re-weight operation. The first weighting (X⊙W) represents the spatial attention weight map, and the second weighting (X⊙W)⊙A final , emphasizing the key spatial regions in the feature map.

[0086] 2. Build the angle detection head

[0087] In traditional target detectors, we usually use the coordinates and size of the target image to generate the corresponding horizontal detection box to describe the corresponding target. However, in directional detection, due to the introduction of a new parameter, the angle, directly using the directional detection box generated by the coordinates, size, and angle to predict the corresponding target may cause some problems, such as Figure 5 As shown in (a), the figure shows that when the predicted angle is defined using the angle encoding [0,π), the predicted angle changes suddenly as the target angle changes to the angle encoding boundary. Ideally, when the target angle is 179°, the predicted angle is also 179°, but when the target angle increases slightly to 181° (greater than the angle encoding range π), the predicted angle jumps to 1°, which means that the detection network needs to fit a piecewise function such as Figure 5 As shown in (b), directly calculating the loss of the predicted angle and the target angle will lead to unstable regression performance of the detection network. This is mainly because of the discontinuity of the loss boundary generated by the angle encoding and the loss function. Therefore, we need a reasonable angle encoding scheme to correct this problem. The present invention solves this problem from two directions: angle encoding and loss function.

[0088] The specific design of the angle prediction head is as follows: by adding two channels to the original detection head part to output the angle prediction coding value, the output is limited to the range of [-1,1] through the tanh activation function, that is, the two angle prediction coding values ​​are cos(ωθ) and sin(ωθ), which is convenient for subsequent decoding and mapping.

[0089] 2.1. We introduce the ACM angle encoding module to provide a continuous, smooth, and reversible encoding function to solve the boundary problem caused by angle encoding. It transforms the angle θ reversibly through the complex exponential transformation f. The angle encoding is output instead of the angle itself at the network output, which means that the angle of the encoded detection box is the same as the angle of the target box in any angle range, that is, f(θ box )=f(θ obj ).

[0090] z=f(θ)=e jωθ =cos(ωθ)+jsin(ωθ) (1)

[0091]

[0092] Where z is the code value, ω is the angular frequency, and mod is the remainder operation. ω≤2.

[0093] According to the detection frame angle θ box Angle θ with the target frame obj The relationship is encoded as follows

[0094]

[0095] When ω=2, e -jωπ =1, we can see that formula (3) is simplified to formula (4), so we can get θ box With θ obj Consistency mapping f(θ box )=f(θ obj ).

[0096]

[0097] 2.2 ProbIou loss function

[0098] In the regression branch, most rotation object detectors predict five parameters (x 0 ,y 0 ,a,b,θ) to represent the oriented object. Since we introduced the AMC angle encoding module, we redesigned a decoupled angle encoding head, such as Figure 3 As shown, the detection network model outputs classification and regression parameters (x p ,y p ,a p ,b p ) and the angle encoding component f p That is, the predicted value is (x p ,y p ,a p ,b p ,f p), the corresponding true value is (x g ,y g ,a g ,b g , f g ), angle encoding component f g By f(θ g ). First, we calculate the ACM loss, which is calculated by p With f g The L1 loss is used to obtain the angle encoding value f p The predicted angle value θ is obtained by inverse calculation (i.e., formula (2)): p And combined with the prediction box parameters (x p ,y p ,a p ,b p ) Calculate the ProbIou loss. The hybrid loss function L is calculated as follows:

[0099] L acm = l smoothl1 (f p ,f g )

[0100] L box = l probiou (B(x p ,y p ,a p ,b p ,θ p ),B(x g ,y g ,a g ,b g ,θ g ))

[0101] L=λ box L box +λ acm L acm

[0102] Here λ box Set to 1, λ acm Set to 0.13, B(x p ,y p ,a p ,b p ,θ p ),B(x g ,y g ,a g ,b g ,θ g ) are the parameter values ​​of the detection box and the target box respectively.

[0103] Step 2: Label the circular identification points in each image in the image sample set, including the size and angle of each detection frame, and then use the labeled images as input to train the detection network model.

[0104] The present invention uses a circular landmark point dataset containing multiple scenes to comprehensively evaluate the detection network model. We collected data from various experimental video sequences, covering various experimental scenes such as structural model seismic experiments, slide rail experiments, and structural model flaw detection experiments; in terms of image acquisition, we used two high-speed cameras: CamRecordCL600×2 (optical device, made in Germany, with a resolution of 1280×1024 pixels) and BaslerACA2040-180KM (Basler, made in Germany, with a resolution of 2048×2048 pixels). The collected dataset covers three different scene conditions, including indoor, outdoor, and scenes specifically used for circular landmark point capture. The recognition landmark point we use is a white circular mark placed on a black background. The entire dataset collects a total of 1080 images containing circular landmarks. These images show circular landmarks at different scales and angles, providing rich data resources for our experiments. The detailed information of the data is shown in Table 1, and is divided into 648 training sets, 216 validation sets, and 216 test sets according to the ratio of training set: validation set: test set = 6:2:2.

[0105] Table 1 Data details

[0106]

[0107] Step 3: Use the trained detection network model to automatically identify each circular marker point in each frame image in the high-speed video.

[0108] In order to verify the feasibility of the automatic recognition method of the present invention, the detection network model is implemented in the PyTorch framework and the opencv library. The experimental platform for training and testing is a single NVIDIA GTX 1080 GPU (12GB) graphics card, AMD Ryzen 7 2700 CPU, and 48GB memory. We used as the optimizer, trained for 100 epochs, set the batch size to 16, and the initial learning rate to 0.001. When training the detector, the warm-up learning rate strategy was used, and the warm-up epoch was 5. The input size was fixed to the same size as that used in the detection process, that is, 640×640 pixels.

[0109] We use four detection rotation target detection models to detect images, namely High-rise FrameStructure Concrete (HFSC), Sliding Rail Docking (SRD), Indoor Low-light Close-up (ILC), and Three-layer Structure Seismic Model (TSSM). Figure 6 Some visualization results on the dataset are shown, where orange boxes represent missed detections and blue triangles represent false detections. The specific detection results are shown in Table 2.

[0110] In the HFSCS, SRD and TSSM scenes, compared with other models, the detection network model of the present invention accurately fits the edge of the circular marker without losing the target, and retains the geometric information of the circular marker, which is attributed to the angle correction ability of its angle encoding, while other models have target loss and misdetection. In the HFSC and TSSM scenes, FCOS lost the target in the far corner due to the large angle of the circular marker and the problem of target occlusion. Although RTMDET detected all the markers, the detection frame was too offset. In the SRD scene, due to the large difference in the size of the circular markers, the other three models all had misdetection or missed detection. In the ILC scene, due to the low scene exposure and the small target, all detection models had target loss or misdetection and multiple detection. However, the detection network model of the present invention still maintained the best performance, only missing 3 small targets and no misdetection, because the detection network model of the present invention has optimized the structure for small targets, enhanced the information reception of small targets, and is equipped with a MOA attention mechanism, which ensures the ability to obtain target features in low light backgrounds and alleviates the loss of small targets.

[0111] Table 2 Comparison of experimental results of target detection effect

[0112]

[0113] As shown in Table 2, the experimental results clearly reveal the differences in detection capabilities of each model. R3DET has an average precision (AP) of 77.8%. FCOS further improves the accuracy of the detection frame by introducing the centerness branch. RTMDET has achieved significant results and efficiency improvements in the field of circular sign detection with its lightweight network structure and efficient feature fusion mechanism. In contrast, the detection network model of the present invention achieves excellent performance in detection accuracy and speed through multi-scale directional feature perception and innovative angle encoding design. The GFLOPs parameter of the algorithm is only 26.8, but it can achieve 83.5% AP while maintaining a real-time detection speed of 41.2FPS.

[0114] Although specific embodiments of the present invention are described above, those skilled in the art should understand that these are merely examples and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is limited by the appended claims.

Claims

1. A method for automatic identification of circular markers in high-speed video end-to-end, characterized in that The following steps are involved: Step 1: Establish a detection network model. The detection network model is based on the YOLOv8 network model, adds multiple multi-scale directional fusion attention modules MOA to its Neck network, adds an angle detection head to its Head detection network, and uses a mixed loss function L for loss calculation; The multi-scale directional fusion attention module MOA is provided with two, which are respectively connected between the second upsampling layer Upsample and the fusion layer Concat of the Neck network and between the third upsampling layer Upsample and the fusion layer Concat, and include a coordinate perception branch, a multi-scale perception branch and a cross-space information aggregation branch. The coordinate perception branch performs feature extraction along the horizontal and vertical directions in space to obtain a position attention descriptor. The multi-scale perception branch uses multiple depth-separable convolutional layers DW with different kernel sizes to extract features to obtain a global attention descriptor. The cross-space information aggregation branch performs feature aggregation on the position attention description that conforms to the global attention descriptor to obtain a spatial position feature map. Step 2: Label the circular identification points in each image in the image sample set, including the size and angle of each detection frame, and then use the labeled images as input to train the detection network model; Step 3: Use the trained detection network model to automatically identify each circular marker point in each frame image in the high-speed video.

2. The method for automatically identifying circular marker points in end-to-end high-speed video according to claim 1, characterized in that: The output image of the Backbone network is recorded as R C*H*W , the coordinate perception branch is used for the input image X∈R C*H*W Average pooling is performed along the horizontal and vertical directions of the space, and then the two images are spliced ​​together in the spatial dimension. Then, the intermediate feature map F is obtained through the Conv layer transformation, and then the feature map F is divided into two tensors represented as Fx and Fy along the horizontal and vertical directions. Then, the reshape operation is used to adjust the dimensions of the two tensors Fx and Fy, and the sigmoid gating function is applied to obtain the corresponding weights Wx and Wy. Finally, the weights Wx and Wy are used to adjust the input image X through the dot product operation to obtain the position attention descriptor Xc; The multi-scale perception branch uses two depth-wise separable convolutional layers DW with different kernels to sequentially process the input image X∈R C *H*W Perform a convolution operation and then perform the feature map output by the two convolution operations The scale feature map A is obtained by splicing, and then the scale feature map A is subjected to global average pooling and global maximum pooling respectively. avg , A max The sigmoid gating mechanism is applied to obtain the corresponding weights, and finally the weights are used to perform weighted summation with the feature map after the depth-separable convolutional layer DW to obtain the global attention descriptor A. final .

3. The method for automatically identifying circular marker points in end-to-end high-speed video according to claim 2, characterized in that: The cross-spatial information aggregation branch takes the scale feature map A and the position attention descriptor Xc as input, performs adaptive global pooling operations on them respectively, obtains the corresponding channel descriptors, and then performs matrix multiplication of the channel descriptors with the corresponding scale feature map A and position attention descriptor Xc to generate two global spatial attention feature maps Y1 and Y2, and then aggregates the two global spatial attention feature maps Y1 and Y2 to generate spatial weights, and finally applies the generated spatial weights to the input feature map X, and combines them with the global attention descriptor A. final , and obtain the final spatial position feature map O.

4. The method for automatically identifying circular marker points in end-to-end high-speed video according to claim 3 is characterized in that: The global attention descriptor A is calculated using the following formula: final ; The spatial position feature map O is calculated using the following formula.

5. The method for automatic identification of circular marker points for end-to-end high-speed video according to claim 1, characterized in that: The angle detection head is set to add two channels to the original detection head of the YOLOv8 network model to output the angle prediction coding values ​​cos(ωθ) and sin(ωθ), and limit them to the range of [-1,1]. Then use the following equation to calculate the angle θ, Among them, z represents the code value, ω represents the angular frequency, and mod represents the remainder operation.

6. The method for automatically identifying circular marker points in end-to-end high-speed video according to claim 5, characterized in that: During training, the predicted angle of the detection box is encoded using the following formula: Among them, θ box / θ obj ∈[0,2π) represent the predicted angle and target angle of the detection box respectively, and ω=2 in this case.

7. The method for automatically identifying circular marker points in end-to-end high-speed video according to claim 4, characterized in that: Use the following formula to calculate the mixed loss function L; L acm =l smoothl1 (in p ,f g ) L box =l probiou (B(x p ,y p ,a p ,b p ,i p ),B(x g ,y g ,a g ,b g ,i g )) L=λ box L box +λ acm L acm Among them, λ box Set to 1, λ acm Set to 0.13, B(x p ,y p ,a p ,b p ,θ p ),B(x g ,y g ,a g ,b g ,θ g ) are the parameter values ​​of the detection box and the target box respectively, f p ,f g They respectively represent the angle encoding components output by the detection network model and their corresponding true values.