Cross-border tracking traffic early warning identification method based on multi-array camera

By deploying multi-array cameras and lidar in cross-border traffic monitoring scenarios, and using an improved Deformable DETR model and multi-head deformable attention mechanism, the shortcomings of target identification and tracking in cross-border traffic monitoring are solved, and high-precision traffic management and improvement of intelligence are achieved.

CN120126307AInactive Publication Date: 2025-06-10XIAN YUANLINGJING INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510183737.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the cross-border traffic monitoring scenarios, the existing technology has problems such as limited target recognition capabilities, poor tracking continuity, and untimely detection of abnormal behaviors, making it difficult to effectively deal with traffic violations.

Method used

By deploying multi-array cameras, lidar and auxiliary sensors to build a cross-border monitoring network, the improved Deformable DETR model is used to perform multi-scale feature extraction, multi-modal feature fusion and multi-head deformable attention mechanism optimization, combined with the multi-objective tracking framework to analyze the target motion laws, and improve the intelligence level of cross-border traffic management.

Benefits of technology

It realizes high-precision target detection and tracking in complex traffic environments, improves the intelligence level of cross-border traffic management, enhances the stability and adaptability of the system, and reduces the missed detection rate and false detection rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126307A_ABST
    Figure CN120126307A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-border tracking traffic early warning identification method based on multi-array cameras, and the method comprises the following steps: S1, deploying the multi-array cameras, a laser radar and an auxiliary sensor, and collecting original image data; s2, preprocessing the collected original image data and the point cloud data collected by the laser radar to generate a multi-modal data set; s3, constructing an improved Deformable DETR model, performing feature extraction on the multi-modal data set by using the improved Deformable DETR model, identifying the category and the position of a target, and generating a detection frame and a confidence score; s4, according to a target detection result, constructing a multi-target tracking framework, and extracting a motion trail of the target; and S5, analyzing the movement track of the target, extracting the movement characteristics of the target, and detecting an abnormal traffic event. According to the method, the improved Deformable DETR model is combined with multi-modal data fusion, accurate detection and intelligent tracking of the cross-border target are achieved, and the method has the advantages of being high in detection accuracy, high in tracking stability and accurate in anomaly recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent traffic monitoring and early warning, and particularly to a method for cross-border tracking traffic early warning recognition based on multi-array cameras. Background Art

[0002] The cross-border traffic flow is increasing day by day, and the safety and efficiency of cross-border transportation have become important issues in traffic management in various countries. In the cross-border area, the flow of vehicles and personnel is frequent, involving traffic regulations, management systems and law enforcement systems in different countries or regions, making it difficult for traditional traffic supervision means to meet the requirements of precision, real-time and intelligence. Traditional monitoring systems usually rely on fixed cameras and manual inspections, and there are problems such as limited target recognition ability, poor tracking continuity, and untimely detection of abnormal behaviors, making it difficult to effectively deal with traffic violations.

[0003] At present, the intelligent transportation system has become an important research direction in traffic management. Intelligent monitoring technologies based on computer vision and deep learning have been widely applied in traffic target detection, tracking and behavior analysis. Existing target detection methods mainly include methods based on traditional image processing technologies and end-to-end detection methods based on deep learning. In traditional methods, feature point detection algorithms such as scale-invariant feature transform, speeded-up robust features, and oriented FAST and rotated BRIEF can be used for target matching and recognition, but these methods are sensitive to conditions such as illumination, occlusion, and motion blur, and it is difficult to be stably applied in complex traffic environments. And target detection methods based on deep learning, such as YOLO, Faster R-CNN, SSD, etc., have been widely applied in the field of traffic target recognition and can achieve high-precision target detection. However, these methods still have certain limitations in cross-border scenarios, especially in aspects such as multi-camera collaborative monitoring, cross-border target tracking, and abnormal behavior detection, where there are problems of low accuracy and insufficient real-time performance.

[0004] In cross-border monitoring scenarios, multi-array cameras can provide a wider range of perspective coverage and improve the accuracy of target detection through data fusion. However, detection methods that solely rely on image data are vulnerable to factors such as occlusion, illumination changes, and target deformation in complex traffic environments, resulting in unstable target tracking. Therefore, more and more research attempts to introduce multi-modal data fusion technologies, such as combining lidar point cloud data, GPS trajectory information and other auxiliary data, to improve the stability and reliability of detection. Point cloud data has strong spatial geometric information and can provide three-dimensional position and shape information of targets, playing an important supplementary role in target detection and tracking tasks. However, how to efficiently fuse multi-modal data and improve the detection accuracy and computational efficiency is still the research focus in the current field of intelligent traffic monitoring.

[0005] Existing object detection algorithms face the following main challenges in cross-border scenarios: First, traditional object detection methods usually only focus on the image data of a single camera, making it difficult to achieve information fusion across cameras and multiple sensors, which may lead to target loss or identity mismatch in cross-border tracking. Second, most existing object tracking methods use Kalman filtering, optical flow methods, or temporal models based on long short-term memory networks, but these methods have problems with decreasing accuracy when dealing with target trajectories over long time spans. Especially when facing situations such as vehicle lane changes, occlusions, or short-term disappearances, target tracking may fail. In addition, anomaly behavior detection relies on accurate target trajectory and behavior pattern modeling, but existing methods are easily interfered by noisy data in complex traffic environments, affecting the accuracy of detection.

[0006] Therefore, how to provide a method for cross-border tracking traffic warning recognition based on multi-array cameras is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] An object of the present invention is to propose a method for cross-border tracking traffic warning recognition based on multi-array cameras. The present invention constructs a cross-border monitoring network through multi-array cameras, lidar, and auxiliary sensors, and adopts multi-scale feature extraction, deformable attention mechanism, and adaptive query vector optimization technology of the improved Deformable DETR model to enhance the robustness of object detection, and combines a multi-object tracking framework to analyze the motion laws of objects, improving the intelligent level of cross-border traffic management.

[0008] A method for cross-border tracking traffic warning recognition based on multi-array cameras according to an embodiment of the present invention includes the following steps:

[0009] S1. Deploy multi-array cameras, lidar, and auxiliary sensors to form a monitoring network in the cross-border area, calibrate the multi-array cameras, and collect original image data;

[0010] S2. Preprocess the collected original image data and the point cloud data collected by the lidar, and perform spatio-temporal alignment to generate a multi-modal data set;

[0011] S3. Construct an improved Deformable DETR model, use the improved Deformable DETR model to extract features from the multi-modal data set, identify the category and location of the object, and generate detection boxes and confidence scores;

[0012] S4. According to the results of object detection, construct a multi-object tracking framework and extract the motion trajectories of the objects;

[0013] S5. Analyze the motion trajectories of the objects, extract the moving features of the objects, and detect abnormal traffic events.

[0014] Optionally, the preprocessing includes denoising, distortion correction, illumination normalization, and color correction.

[0015] Optionally, S3 specifically includes:

[0016] S31. Construct an improved Deformable DETR model, where the improved Deformable DETR model includes a multi-scale feature extraction module, a multi-modal feature fusion module, a multi-head deformable attention module, a query optimization module, and a post-processing module. The improvement includes introducing the FPN network and the Swin Transformer network as the backbone networks of the improved Deformable DETR model, adding a multi-head deformable attention mechanism, and an adaptive query vector;

[0017] S32. Extract image data and point cloud data from the multi-modal dataset, represent the image data as a three-channel matrix I ∈ R H×W×3 , where H of the three-channel matrix represents the image height, W represents the image width, and represent the point cloud data as a three-dimensional point set The (x j , y j , z j ) of the three-dimensional point set are spatial coordinates, and r j is the reflection intensity. The point cloud data is projected and mapped to the image coordinate system;

[0018] S33. Perform feature extraction on the image data, construct a multi-scale feature pyramid, and use the FPN network combined with the Swin Transformer network for feature extraction to output the image data features;

[0019] S34. Use a three-dimensional sparse convolutional neural network to extract the point cloud data features;

[0020] S35. Perform multi-scale feature fusion through the multi-modal feature fusion module, map the features of different modalities to a unified feature space, and generate multi-modal features;

[0021] S36. Optimize the object queries based on the multi-modal features using the multi-head deformable attention mechanism;

[0022] S37. Perform object detection, output the object category, object location, and confidence score, and use the non-maximum suppression method to optimize the object detection boxes.

[0023] Optionally, S33 specifically includes:

[0024] S331. The input image data is subjected to multi-scale feature extraction using the FPN network, the input image is downsampled layer by layer, and feature pyramid layers of different resolutions are constructed. The feature pyramid layers include:

[0025] The first layer I 1 Use 1×1 convolution to extract high-resolution features The size remains H×W;

[0026] The second layer I 2 Use 3×3 convolution and pooling operations to downsample the features to H / 2×W / 2;

[0027] The third layer I 3 Use 3×3 convolution and pooling operations to downsample the features to H / 4×W / 4;

[0028] The fourth layer I 4 Use 3×3 convolution and pooling operations to downsample the features to H / 8×W / 8;

[0029] The feature map of each layer is non-linearly transformed by a convolutional neural network to form pyramid features:

[0030]

[0031] Among them, represents the feature map of the l-th layer, Conv l represents the convolutional neural network of the l-th layer, I l represents the input image data of the l-th layer, and l represents the number of layers;

[0032] S332. For the feature map of each pyramid layer Perform window partitioning, and set the window size to W s ×W s , calculate the attention weight A within the window, and calculate the weighted window feature:

[0033]

[0034] Among them, F A represents the window feature after attention weighting, retaining local spatial information, Q 1 , K 1 and V 1 represent the query, key, and value matrices respectively, W Q1 , W K1 and W V1 represent the query, key, and value weight matrices respectively, T represents the transpose operation, d k represents the dimension of the key, Softmax represents normalization, represents the feature block of the i,j-th window, and A represents the attention weight;

[0035] S333. Use window offset to enable information interaction between adjacent windows, including:

[0036] In the odd layers, the window partitioning method is the conventional W s ×W s ;

[0037] In the even layers, the entire window is translated by (W s / 2 × W s / 2) to achieve cross-window feature fusion. After the window translation, the cross-window attention is recalculated, and the attention feature F′ A :

[0038] Among them, F′ A represents the attention feature after cross-window fusion, enabling information communication between different windows;

[0039] S334. The calculated attention feature F′ A is normalized by LayerNorm and undergoes a non-linear transformation through a multi-layer perceptron, and finally the image data feature is output

[0040] Optionally, the S36 specifically includes:

[0041] S361. Initialize the target query sequence to generate an initial query vector:

[0042] q k = LayerNorm(tanh(W Q Q + W F F + b));

[0043] Among them, q k represents the initial query vector, LayerNorm represents layer normalization, tanh represents the activation function, W Q and W F represent the learning parameter matrices, Q represents the query vector set, F represents the multi-modal feature, and b represents the offset;

[0044] S362. Construct multi-head deformable attention query-key-value pairs, and calculate the deformable offsets Δx k and Δy k :

[0045] Q = W Q F, K = W K F, V = W V F;

[0046]

[0047] Among them, Q, K, and V respectively represent the query, key, and value matrices of the multi-head deformable attention, W Q 、W K and W VRespectively represent the weight matrices corresponding to the query, key, and value, Δx k and Δy k represent the deformable offset, ω m represents the normalization coefficient, W m represents the weight matrix at different scales, b m represents the bias term, tanh represents the activation function, M represents the number of scales of the deformable attention, and F represents the multimodal feature;

[0048] S363. Calculate the multi-head deformable attention weighted feature:

[0049]

[0050] Among them, F out represents the multi-head deformable attention weighted feature, H represents the number of attention heads, V h represents the value matrix of the h-th attention head, exp represents the natural exponential function, b h represents the attention bias term, q k and q j represent the k-th and j-th query vectors, represents the mapping weight of the query vector, K represents the key matrix, represents the mapping weight of the key matrix;

[0051] S364. Apply the deformable offsets Δx k and Δy k to adjust the multi-head deformable attention weighted feature query:

[0052] x′ k = x k + Δx k , y′ k = y k + Δy k ;

[0053]

[0054] Among them, x′ k and y′ k represent the target feature query positions after the deformable offset, x k and y k represent the target feature query positions before the deformable offset, and F′ out represents the adjusted multi-head deformable attention weighted feature.

[0055] Optionally, the S37 specifically includes:

[0056] S371. Based on the adjusted multi-head deformable attention weighted feature F′ out , use a multi-layer perceptron for target class classification:

[0057] C k = Softmax(W c F′ out + b c );

[0058] Among them, C k represents the target class probability distribution, Softmax represents normalization, W c represents the training classification matrix, b c represents the bias term;

[0059] S372. Calculate the target center coordinates, width, and height, and query the positions x′ k and y′ k :

[0060] B k = (x k , y k , w k , h k ) = W b F′ out + b b ;

[0061] B′ k = (x′ k , y′ k , w k , h k );

[0062] Among them, W b and b b respectively represent the bounding box regression matrix and the bias term, B k represents the target box, B′ k represents the corrected target box, w k and h k respectively represent the width and height;

[0063] S373. Calculate the confidence score:

[0064] S k = Sigmoid(W s F′ out + b s );

[0065] Among them, S k represents the confidence score, W s and b s respectively represent the confidence calculation matrix and the bias term, Sigmoid represents the activation function;

[0066] S374. Perform post - processing of object detection and optimize the object detection bounding boxes using the non - maximum suppression method:

[0067]

[0068] Among them, τ represents the non - maximum suppression threshold, τ 0 represents the initial non - maximum suppression threshold, λ represents the hyperparameter, exp represents the natural exponential function, σ represents the normalization factor, which is used to control the sensitivity of density calculation, ||(x i ,y i )-(x j ,y j )|| 2 represents the square of the Euclidean distance between object i and object j, K represents the total number of objects in the current frame, and calculate the intersection - over - union (IoU) between objects:

[0069]

[0070] Among them, IoU ij represents the intersection - over - union between object i and object j, B i represents the detection bounding box of object i, B j represents the detection bounding box of object j, ∩ represents "intersection", ∪ represents "union", if IoU ij >τ, then delete the object with lower confidence.

[0071] Optionally, the S4 specifically includes:

[0072] S41. Establish a multi - object trajectory tracking model, set the time step t and the object set K in the object set T represents the number of detected objects in the current frame, and each object T k is represented by spatio - temporal information and the feature vector , initialize the object trajectory set

[0073]

[0074] Among them, represents the spatio - temporal information of the k - th object, Γ k represents the historical trajectory data of the k - th object, T represents the length of the historical time window, and represent the center coordinates, width and height of the object's bounding box;

[0075] S42. Set the historical trajectories of two objects T i and T j and calculate the trajectory similarity:

[0076]

[0077] Among them, D traj (Γ i , Γ j ) represents the Euclidean distance of the trajectories of target T i and target T j within the time window T. Γ i represents the trajectory of target T i , Γ j represents the trajectory of target T j . represents the center coordinates of the bounding box of target T i at time t. represents the center coordinates of the bounding box of target T j at time t;

[0078] S43. Construct the target motion state transition matrix and calculate the state of the target at the next moment:

[0079]

[0080] Among them, represents the motion state vector of target k at time t, and represent the center coordinates of the bounding box of the target, represents the speed of the target, represents the motion direction angle of the target. B represents the state transition matrix, which is the mathematical model of the target motion. W represents the noise and follows a Gaussian distribution. represents the motion state vector of target k at time t + 1;

[0081] S44. Calculate the matching weight between target T i and target T j :

[0082]

[0083] Among them, R i,j represents the matching similarity score between target T i and target T j . α and β represent the weighting coefficients, which respectively control the influence of the trajectory similarity and the feature similarity. represents the cosine similarity of features between target T i and target T j . and respectively represent the feature vectors of target T i and target T j ;

[0084] S45. Calculate the optimal target matching:

[0085]

[0086] Among them, M represents the target matching matrix, which is the optimal target matching scheme, and π represents the target matching set, which contains all successfully matched target pairs (T i , T j ). denotes maximizing the similarity scores of all matched target pairs;

[0087] S45. If the matching is successful, update the trajectory of the target:

[0088]

[0089] If the target fails to match successfully, perform a target loss marking.

[0090] Optionally, the movement features include speed change, driving path, and behavior features, and the abnormal traffic events include speeding, running a red light, illegal lane change, and illegal border crossing.

[0091] The beneficial effects of the present invention are as follows:

[0092] First, in terms of target detection, the present invention adopts an improved Deformable DETR model. By introducing the FPN network and the Swin Transformer network as the backbone networks, the detection ability for targets of different scales is enhanced, enabling targets to still be accurately recognized in the cases of long distance, complex illumination, and partial occlusion environments. In addition, by introducing the multi-head deformable attention mechanism, the spatial adaptability of target detection is improved, making the detection boxes more accurate. At the same time, combined with the adaptive query vector optimization, the stability of target recognition is improved in the cross-border scenario with dense targets, reducing the missed detection rate and false detection rate.

[0093] Second, in terms of multi-target tracking, the present invention greatly improves the target identity consistency and tracking continuity in the cross-border scenario through the cross-camera target matching and trajectory prediction model. Different from the existing technology which mainly relies on single-camera tracking resulting in identity loss problems, the present invention adopts the Hungarian matching algorithm, which can still maintain the stability of the target trajectory and identity consistency when the target is temporarily occluded, the environment changes, or the camera switches, avoiding the problems of target loss or mis-identification, and improving the reliability of cross-camera target tracking.

[0094] Finally, the multi-modal data fusion ability of the present invention significantly enhances the stability and adaptability of the system. Through the fusion of lidar point cloud data and camera image data, not only the accuracy of target detection is improved, but also the reliability of target recognition and tracking can be ensured under low-light and adverse weather conditions. Compared with traditional detection systems that only rely on single-modal data, the present invention has achieved technological breakthroughs in multi-modal data alignment, feature fusion, and cross-camera target matching, enabling the system to still have high adaptability in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0096] Figure 1 is a flowchart of a method for cross-border tracking traffic warning recognition based on a multi-array camera proposed by the present invention;

[0097] Figure 2 is a structural block diagram of an improved Deformable DETR model for a method for cross-border tracking traffic warning recognition based on a multi-array camera proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0098] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0099] Refer to Figure 1 and Figure 2 , a method for cross-border tracking traffic warning recognition based on a multi-array camera, includes the following steps:

[0100] S1. Deploy multi-array cameras, lidar, and auxiliary sensors to form a monitoring network in the cross-border area, calibrate the multi-array cameras, and collect original image data;

[0101] S2. Preprocess the collected original image data and the point cloud data collected by the lidar, and perform spatio-temporal alignment to generate a multi-modal data set;

[0102] S3. Construct an improved Deformable DETR model, use the improved Deformable DETR model to extract features from the multi-modal data set, identify the category and location of the target, and generate detection boxes and confidence scores;

[0103] S4. According to the results of target detection, construct a multi-target tracking framework and extract the motion trajectories of the targets;

[0104] S5. Analyze the movement trajectory of the target, extract the movement features of the target, and detect abnormal traffic events.

[0105] In this embodiment, the preprocessing includes denoising, distortion correction, illumination normalization, and color correction.

[0106] In this embodiment, the specific steps of S3 are as follows:

[0107] S31. Construct an improved Deformable DETR model, which includes a multi-scale feature extraction module, a multi-modal feature fusion module, a multi-head deformable attention module, a query optimization module, and a post-processing module. The improvement includes introducing the FPN network and the Swin Transformer network as the backbone networks of the improved Deformable DETR model, adding a multi-head deformable attention mechanism, and an adaptive query vector.

[0108] S32. Extract image data and point cloud data from the multi-modal dataset. Represent the image data as a three-channel matrix I ∈ R H×W×3 , where H of the three-channel matrix represents the image height, W represents the image width, and represent the point cloud data as a three-dimensional point set The (x j , y j , z j ) in the three-dimensional point set are spatial coordinates, and r j is the reflection intensity. The point cloud data is projected and mapped to the image coordinate system.

[0109] S33. Extract features from the image data, construct a multi-scale feature pyramid, and use the FPN network combined with the Swin Transformer network for feature extraction to output the image data features.

[0110] S34. Use a three-dimensional sparse convolutional neural network to extract the point cloud data features.

[0111] S35. Perform multi-scale feature fusion through the multi-modal feature fusion module, map the features of different modalities to a unified feature space, and generate multi-modal features.

[0112] S36. Based on the multi-modal features, use the multi-head deformable attention mechanism to optimize the target query.

[0113] S37. Perform object detection, output the object category, object location, and confidence score, and use the non-maximum suppression method to optimize the object detection box.

[0114] In this embodiment, the specific steps of S33 are as follows:

[0115] S331. The input image data is subjected to multi-scale feature extraction using the FPN network. The input image is downsampled layer by layer to construct feature pyramid layers with different resolutions. The feature pyramid layers include:

[0116] The first layer I 1 Extract high-resolution features using 1×1 convolution The size remains H×W;

[0117] The second layer I 2 Use 3×3 convolution and pooling operations to downsample the features to H / 2×W / 2;

[0118] The third layer I 3 Use 3×3 convolution and pooling operations to downsample the features to H / 4×W / 4;

[0119] The fourth layer I 4 Use 3×3 convolution and pooling operations to downsample the features to H / 8×W / 8;

[0120] Each layer of feature map is subjected to non-linear transformation by a convolutional neural network to form pyramid features:

[0121]

[0122] Among them, represents the feature map of the l-th layer, Conv l represents the convolutional neural network of the l-th layer, I l represents the input image data of the l-th layer, and l represents the layer number;

[0123] S332. Perform window partitioning on the feature map of each pyramid layer The window size is set to W s ×W s , calculate the attention weight A within the window, and calculate the weighted window features:

[0124]

[0125] Among them, F A represents the window features after attention weighting, retaining local spatial information, Q 1 , K 1 and V 1 represent the query, key, and value matrices respectively, W Q1 , W K1 and W V1 represent the query, key, and value weight matrices respectively, T represents the transpose operation, d k represents the dimension of the key, Softmax represents normalization, represents the feature block of the i,j-th window, and A represents the attention weight;

[0126] S333. Adopt window offset to enable information interaction between adjacent windows, including:

[0127] In the odd layers, the window division method is the regular W s ×W s ;

[0128] In the even layers, the whole window is translated by (W s / 2×W s / 2) to achieve cross-window feature fusion. After the window translation, recalculate the cross-window attention and calculate the attention feature F′ A :

[0129]

[0130] Among them, F′ A represents the attention feature after cross-window fusion, enabling the information of different windows to communicate with each other;

[0131] S334. Perform LayerNorm normalization on the calculated attention feature F′ A and perform non-linear transformation through a multi-layer perceptron, and finally output the image data feature

[0132] In this embodiment, the S36 specifically includes:

[0133] S361. Initialize the target query sequence to generate an initial query vector:

[0134] q k = LayerNorm(tanh(W Q Q + W F F + b));

[0135] Among them, q k represents the initial query vector, LayerNorm represents layer normalization, tanh represents the activation function, W Q and W F represent the learning parameter matrices, Q represents the query vector set, F represents the multi-modal feature, and b represents the offset;

[0136] S362. Construct multi-head deformable attention query-key-value pairs and calculate the deformable offsets Δx k and Δy k :

[0137] Q = W Q F, K = W K F, V = W V F;

[0138]

[0139] Among them, Q, K, and V respectively represent the query, key, and value matrices of the multi-head deformable attention. W Q , W K and W V respectively represent the weight matrices corresponding to the query, key, and value. Δx k and Δy k represent the deformable offsets. ω m represents the normalization coefficient. W m represents the weight matrix of different scales. b m represents the bias term. tanh represents the activation function. M represents the number of scales of the deformable attention. F represents the multi-modal feature;

[0140] S363. Calculate the multi-head deformable attention weighted feature:

[0141]

[0142] Among them, F out represents the multi-head deformable attention weighted feature. H represents the number of attention heads. V h represents the value matrix of the h-th attention head. exp represents the natural exponential function. b h represents the attention bias term. q k and q j represent the k-th and j-th query vectors. represents the mapping weight of the query vector. K represents the key matrix. represents the mapping weight of the key matrix;

[0143] S364. Apply the deformable offsets Δx k and Δy k to adjust the multi-head deformable attention weighted feature query:

[0144] x′ k = x k + Δx k , y′ k = y k + Δy k ;

[0145]

[0146] Among them, x′ k and y′ k represent the target feature query positions after the deformable offset. x k and y k represent the target feature query positions before the deformable offset. F′ out represents the adjusted multi-head deformable attention weighted feature.

[0147] In this embodiment, S37 specifically includes:

[0148] S371. Classify the target category using a multi-layer perceptron based on the adjusted multi-head deformable attention weighted feature F': out ,

[0149] C k = Softmax(W c F' out + b c );

[0150] where C k represents the target category probability distribution, Softmax represents normalization, W c represents the training classification matrix, and b c represents the bias term;

[0151] S372. Calculate the target center coordinates, width, and height, and query the positions x' k and y' k :

[0152] B k = (x k , y k , w k , h k ) = W b F' out + b b ;

[0153] B' k = (x', k , y', k , w k , h k );

[0154] where W b and b b represent the bounding box regression matrix and the bias term respectively, B k represents the target box, B' k represents the corrected target box, and w k and h k represent the width and height respectively;

[0155] S373. Calculate the confidence score:

[0156] S k = Sigmoid(W s F' out + b s );

[0157] where S k represents the confidence score, and Ws and b s respectively represent the confidence calculation matrix and the bias term, and Sigmoid represents the activation function;

[0158] S374. After performing object detection post-processing, the non-maximum suppression method is used to optimize the object detection bounding boxes:

[0159]

[0160] where τ represents the non-maximum suppression threshold, τ 0 represents the initial non-maximum suppression threshold, λ represents the hyperparameter, exp represents the natural exponential function, σ represents the normalization factor, which is used to control the sensitivity of density calculation, ||(x i , y i ) - (x j , y j )|| 2 represents the square of the Euclidean distance between object i and object j, K represents the total number of objects in the current frame, and calculate the intersection over union between objects:

[0161]

[0162] where IoU ij represents the intersection over union between object i and object j, B i represents the detection bounding box of object i, B j represents the detection bounding box of object j, ∩ represents "intersection", ∪ represents "union", if IoU ij > τ, then delete the object with lower confidence.

[0163] In this embodiment, the S4 specifically includes:

[0164] S41. Establish a multi-object trajectory tracking model, set the time step t and the object set where K in the object set T represents the number of detected objects in the current frame, and each object T k is represented by spatio-temporal information and the feature vector Initialize the object trajectory set

[0165]

[0166] where, represents the spatio-temporal information of the k-th object, Γ k represents the historical trajectory data of the k-th object, T represents the length of the historical time window, and represent the center coordinates, width and height of the bounding box of the object;

[0167] S42. Set two targets T i and T j 's historical trajectories, and calculate the trajectory similarity:

[0168]

[0169] where D traj (Γ i , Γ j ) represents the Euclidean distance of the trajectories of target T i and target T j within the time window T, Γ i represents the trajectory of target T i , Γ j represents the trajectory of target T j . represents the center coordinates of the bounding box of target T i at time t, represents the center coordinates of the bounding box of target T j at time t;

[0170] S43. Construct the target motion state transition matrix and calculate the state of the target at the next moment:

[0171]

[0172] where represents the motion state vector of target k at time t, and represent the center coordinates of the bounding box of the target, represents the speed of the target, represents the motion direction angle of the target, B represents the state transition matrix, which is the mathematical model of the target motion, W represents the noise and follows a Gaussian distribution, represents the motion state vector of target k at time t + 1;

[0173] S44. Calculate the matching weight between target T i and target T j :

[0174]

[0175] where R i,j represents the matching similarity score between target T i and target T j , α and β represent the weighting coefficients, which respectively control the influence of trajectory similarity and feature similarity, represents the cosine similarity of features between target T i and target T j , and respectively represent the target T i and the target T j 's feature vectors;

[0176] S45. Calculate the optimal target matching:

[0177]

[0178] where M represents the target matching matrix, which is the optimal target matching scheme, π represents the target matching set, containing all successfully matched target pairs (T i , T j ), represents maximizing the similarity scores of all matched target pairs;

[0179] S45. If the matching is successful, update the trajectory of the target:

[0180]

[0181] If the target fails to match successfully, perform a target loss marking.

[0182] In this embodiment, the movement features include speed change, driving path, and behavior features, and the abnormal traffic events include speeding, running a red light, illegal lane change, and illegal crossing of the border.

[0183] Example 1:

[0184] To verify the feasibility of the present invention in implementation, the present invention is applied to an intelligent transportation management system at a cross-border port to perform intelligent identification, tracking, and early warning on cross-border vehicles, pedestrians, motorcycles, and illegal crossing targets within the port area. The scenarios cover different traffic flows, lighting conditions, target categories, and border environments to comprehensively evaluate the performance of the present invention in complex cross-border scenarios.

[0185] During the testing process, trucks, private cars, motorcycles, and pedestrians are detected respectively, and abnormal behaviors such as speeding, illegal lane change, and illegal crossing are identified. By deploying 16 high-definition AI cameras and 4 lidars, a large amount of multi-modal data is collected, and the improved Deformable DETR model proposed by the present invention is used for target recognition. This experiment mainly tests key indicators such as the target detection accuracy, target tracking stability, abnormal behavior detection rate, and system response speed of the present invention, and at the same time conducts a comparative analysis with the existing technology to verify the superiority of the present invention.

[0186] Table 1 Comparison table of experimental data

[0187] Index YOLOv5 FasterR-CNN SSD The present invention Object detection accuracy (%) 86.2 88.5 81.4 95.3 Object tracking loss rate (%) 12.3 10.8 14.6 3.7 Overspeed detection accuracy rate (%) 85.3 87.2 82.5 96.1 Illegal lane change detection accuracy rate (%) 82.7 85.5 79.3 94.8 Illegal border crossing detection accuracy rate (%) 74.1 78.6 70.2 91.3 Object detection time (seconds) 0.45 0.68 0.52 0.39 Object tracking time (seconds) 0.82 1.12 0.95 0.67 Abnormality recognition time (seconds) 1.25 1.48 1.32 0.93 Warning trigger time (seconds) 0.55 0.71 0.63 0.47 Total system response delay (seconds) 3.07 3.99 3.42 2.46

[0188] In terms of the accuracy of object detection, the detection accuracy rate of the present invention reaches 95.3%, which is improved by 9.1%, 6.8% and 13.9% compared with YOLOv5, Faster R-CNN and SSD respectively. In particular, it has significant advantages in the recognition of small targets such as pedestrians and motorcycles. This is mainly due to the improved Deformable DETR model combined with FPN and Swin Transformer, which enhances the multi-scale object detection ability and optimizes the recognition accuracy in complex scenarios.

[0189] In terms of the stability of object tracking, the object loss rate of the present invention is 3.7%, which is reduced by more than 65% compared with YOLOv5 and FasterR-CNN, improving the identity consistency and continuity of cross-border objects. The present invention can still accurately match the target and achieve stable tracking in the case of cross-camera switching, partial occlusion or short-term disappearance.

[0190] In terms of the accuracy of abnormal behavior detection, the present invention has a great improvement in detection tasks such as speeding, illegal lane change, and illegal border crossing. Especially in the detection of illegal border crossing, it reaches 91.3%, which is increased by 17.2% and 12.7% compared with YOLOv5 and Faster R-CNN respectively. This improvement mainly stems from the dual optimization of multi-modal data fusion and trajectory behavior analysis, enabling the system to more accurately distinguish normal and abnormal traffic behaviors and improving the accuracy of cross-border traffic safety monitoring.

[0191] In terms of the system response speed, the total response delay of the present invention is 2.46 seconds, which is improved by 0.61 second, 1.53 seconds and 0.96 seconds compared with YOLOv5, Faster R-CNN and SSD respectively, with an overall improvement of about 20%-38%, achieving faster traffic warnings. The present invention can quickly process object detection and tracking tasks locally and only upload key data to the cloud for storage and long-term analysis, thereby reducing data transmission delay and improving real-time performance.

[0192] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A method for cross-border tracking traffic warning recognition based on multi-array cameras, characterized in that: The steps include: S1. Deploy multi-array cameras, lidar and auxiliary sensors to form a monitoring network in cross-border areas, calibrate multi-array cameras, and collect raw image data; S2, preprocessing the collected original image data and the point cloud data collected by the lidar, and performing spatiotemporal alignment to generate a multimodal data set; S3. Build an improved Deformable DETR model, use the improved Deformable DETR model to extract features from the multimodal dataset, identify the category and location of the target, and generate a detection box and confidence score; S4. Based on the target detection results, a multi-target tracking framework is constructed to extract the target's motion trajectory; S5. Analyze the target's motion trajectory, extract the target's movement features, and detect abnormal traffic events.

2. According to claim 1, a method for cross-border tracking traffic warning recognition based on multi-array cameras is characterized in that: The preprocessing includes denoising, distortion correction, illumination normalization and color correction.

3. According to claim 1, a method for cross-border tracking traffic warning recognition based on multi-array cameras is characterized in that: The S3 specifically includes: S31, constructing an improved Deformable DETR model, wherein the improved Deformable DETR model includes a multi-scale feature extraction module, a multi-modal feature fusion module, a multi-head deformable attention module, a query optimization module and a post-processing module, wherein the improvement includes introducing an FPN network and a Swin Transformer network as the backbone network of the improved Deformable DETR model, adding a multi-head deformable attention mechanism and an adaptive query vector; S32. Extract image data and point cloud data from the multimodal dataset and represent the image data as a three-channel matrix I∈R H ×W×3 , the three-channel matrix H represents the image height, W represents the image width, and the point cloud data is represented as a three-dimensional point set The three-dimensional point set (x j ,y j ,z j ) is the spatial coordinate, r j is the reflection intensity, the point cloud data is mapped to the image coordinate system by projection; S33, extract features from image data, construct a multi-scale feature pyramid, use FPN network combined with SwinTransformer network to extract features, and output image data features; S34, extracting point cloud data features using a three-dimensional sparse convolutional neural network; S35, performing multi-scale feature fusion through a multi-modal feature fusion module, mapping features of different modalities into a unified feature space, and generating multi-modal features; S36, Based on multimodal features, a multi-head deformable attention mechanism is used to optimize target queries; S37, perform target detection, output target category, target location and confidence score, and use non-maximum suppression method to optimize the target detection frame.

4. According to claim 3, a method for cross-border tracking traffic warning recognition based on multi-array cameras is characterized in that: The S33 specifically includes: S331, the input image data is subjected to multi-scale feature extraction using the FPN network, the input image is downsampled layer by layer, and feature pyramid layers of different resolutions are constructed, the feature pyramid layers comprising: The first layer I1 uses 1×1 convolution to extract high-resolution features. The size remains H×W; The second layer I2 uses 3×3 convolution and pooling operations to downsample the features to H / 2×W / 2; The third layer I3 uses 3×3 convolution and pooling operations to downsample the features to H / 4×W / 4; The fourth layer I4 uses 3×3 convolution and pooling operations to downsample the features to H / 8×W / 8; Each layer of feature map is nonlinearly transformed by the convolutional neural network to form a pyramid feature: in, Represents the feature map of the lth layer, Conv l represents the convolutional neural network of the lth layer, I l Represents the input image data of the lth layer, where l represents the number of layers; S332, for each pyramid layer feature map Divide the window and set the window size to W s ×W s , calculate the attention weight A within the window, and calculate the weighted window features: Among them, F A represents the window feature after attention weighting, retaining the local spatial information, Q1, K1 and V1 represent the query, key and value matrices respectively, W Q1 , W K1 and W V1 denote the weight matrices of query, key, and value respectively, T denotes the transpose operation, and d k represents the dimension of the key, Softmax represents normalization, represents the feature block of the i,jth window, and A represents the attention weight; S333, using window offset to enable information exchange between adjacent windows, including: In odd layers, the window division method is the conventional W s ×W s ; On even layers, the window is shifted as a whole (W s / 2×W s / 2) to achieve cross-window feature fusion. After the window is translated, the cross-window attention is recalculated and the attention feature F′ is calculated A : Among them, F′ A Represents the attention features after cross-window fusion, so that information from different windows can be communicated; S334, the calculated attention feature F′ A Perform LayerNorm normalization and perform nonlinear transformation through a multi-layer perceptron to finally output image data features 5. According to claim 3, a method for cross-border tracking traffic warning recognition based on multi-array cameras is characterized in that: The S36 specifically includes: S361, initialize the target query sequence and generate an initial query vector: q k =LayerNorm(tanh(W Q Q+W F F+b)); Among them, q k represents the initial query vector, LayerNorm represents layer normalization, tanh represents the activation function, W Q and W F represents the learning parameter matrix, Q represents the query vector set, F represents the multimodal feature, and b represents the offset; S362: Construct multi-head deformable attention query key-value pairs and calculate deformable offset Δx k and Δy k : Q=W Q F,K=W K F,V=W V F; where Q, K, and V represent the query, key, and value matrices of multi-head deformable attention, respectively, and W Q , W K and W V Denote the weight matrices corresponding to query, key, and value, respectively, Δx k and Δy k represents the deformable offset, ω m represents the normalization coefficient, W m Represents weight matrices of different scales, b m represents the bias term, tanh represents the activation function, M represents the number of scales of deformable attention, and F represents the multimodal feature; S363. Calculate multi-head deformable attention weighted features: Among them, F out represents the multi-head deformable attention weighted feature, H represents the number of attention heads, V h represents the value matrix of the h-th attention head, exp represents the natural exponential function, b h represents the attention bias term, q k and q j represents the kth and jth query vectors, represents the mapping weight of the query vector, K represents the key matrix, represents the mapping weight of the key matrix; S364, apply deformable offset Δx k and Δy k Perform multi-head deformable attention weighted feature query adjustment: x′ k =x k +Δx k ,y′ k =y k +Δy k ; Among them, x′ k and y′ k represents the target feature query position after deformable offset, x k and k represents the target feature query position before deformable migration, F′ out Represents the adjusted multi-head deformable attention weighted features.

6. The method for cross-border tracking traffic warning recognition based on multi-array cameras according to claim 3 is characterized in that: The S37 specifically includes: S371, based on the adjusted multi-head deformable attention weighted feature F′ out , use multi-layer perceptron to classify target categories: C k =Softmax(W c F′ out +b c ); Among them, C k represents the probability distribution of the target category, Softmax represents normalization, and W c represents the training classification matrix, b c represents the bias term; S372, calculate the target center coordinates, width, height, and apply the deformable offset target feature query position x′ k and y′ k : B k =(x k ,y k ,w k ,h k )=W b F′ out +b b ; B′ k =(x′ k ,y′ k ,w k ,h k ); Among them, W b and b b denote the bounding box regression matrix and the bias term, respectively, and B k represents the target box, B′ k represents the corrected target box, w k and h k Represents width and height respectively; S373. Calculate confidence score: S k =Sigmoid(W s F′ out +b s ); Among them, S k represents the confidence score, W s and b s They represent the confidence calculation matrix and bias term respectively, and Sigmoid represents the activation function; S374, perform post-target detection processing and use non-maximum suppression method to optimize the target detection frame: Among them, τ represents the non-maximum suppression threshold, τ0 represents the initial non-maximum suppression threshold, λ represents the hyperparameter, exp represents the natural exponential function, σ represents the normalization factor, which is used to control the sensitivity of density calculation, ||(x i ,y i )-(x j ,y j )|| 2 represents the square of the Euclidean distance between target i and target j, K represents the total number of targets in the current frame, and the intersection-over-union ratio between targets is calculated: Among them, IoU ij represents the intersection-over-union ratio between target i and target j, B i represents the detection bounding box of target i, B j represents the detection bounding box of target j, ∩ represents "intersection", ∪ represents "union", if IoU ij >τ, the targets with lower confidence are removed.

7. The method for cross-border tracking traffic warning recognition based on multi-array cameras according to claim 1 is characterized in that: The S4 specifically includes: S41. Establish a multi-target trajectory tracking model, set the time step t and target set The K in the target set T represents the number of detected targets in the current frame. Each target T k By space-time information and the eigenvector Indicates that the target trajectory set is initialized in, represents the spatiotemporal information of the kth target, Γ k represents the historical trajectory data of the kth target, T represents the length of the historical time window, and Represents the center coordinates, width and height of the bounding box of the target; S42. Set two goals T i and T j The historical trajectory of , and calculate the trajectory similarity: Among them, D traj (Γ i ,Γ j ) represents the target T i and target T j The Euclidean distance of the trajectory within the time window T, Γ i Indicates the target T i The trajectory of Γ j Indicates the target T j The trajectory of Indicates the target T i The coordinates of the center of the bounding box at time t, Indicates the target T j The coordinates of the center of the bounding box at time t; S43, construct the target motion state transfer matrix, and calculate the state of the target at the next moment: in, represents the motion state vector of target k at time t, and represents the center coordinates of the target's bounding box, represents the speed of the target, represents the target's moving direction angle, B represents the state transfer matrix, which is the mathematical model of the target's movement, W represents the noise, which obeys the Gaussian distribution, represents the motion state vector of target k at time t+1; S44. Calculate target T i and target T j The matching weight between: Among them, R i,j Indicates the target T i and target T j The matching similarity score, α and β represent weighting coefficients, which control the influence of trajectory similarity and feature similarity respectively. Indicates the target T i and target T j The feature cosine similarity of and Respectively represent the target T i and target T j The eigenvector of S45. Calculate the optimal target match: Where M represents the target matching matrix, which is the optimal target matching solution, and π represents the target matching set, which contains all successfully matched target pairs (T i ,T j ), Represents maximizing the similarity score of all matching target pairs; S45. If the match is successful, update the target trajectory: If the target fails to match, the target lost flag is executed.

8. The method for cross-border tracking traffic warning recognition based on multi-array cameras according to claim 1 is characterized in that: The movement characteristics include speed changes, driving paths and behavior characteristics, and the abnormal traffic events include speeding, running red lights, illegal lane changes and illegal border crossings.

Citation Information

Cited By

  • Road scene multi-target tracking system based on visual perception of unmanned aerial vehicle

    CN120298460A

  • Object motion trajectory tracking method and system based on visual features

    CN120411174A

  • Traffic overrun early warning method and system under multi-source data fusion

    CN120808600A

  • Multi-feature trajectory prediction and cross-defense early warning method for edge and coast defense control objects

    CN121011112A

  • Target detection neural network accelerator based on compound eye array camera

    CN121415358A