A radar and video fusion dynamic space-time alignment highway monitoring method

By independently retrieving and dynamically aligning radar and video surveillance targets, and combining spatial location, motion trajectory, and appearance feature correlation matching, the problems of insufficient spatiotemporal alignment accuracy and poor target correlation robustness in existing technologies are solved, achieving high-precision and stable highway monitoring results.

CN120913409BActive Publication Date: 2026-02-10SHENZHEN ARCKE INNOVATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511415299.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-02-10
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing technologies lack sufficient spatiotemporal alignment accuracy, inaccurate time synchronization, and robustness in target association and matching in radar and video fusion monitoring, especially prone to mismatch and target identity switching in complex traffic scenarios.

Method used

By acquiring radar point cloud data and video image data in real time, the system independently retrieves radar and video surveillance targets, performs dynamic spatial alignment and multi-scale time synchronization, and performs association matching based on spatial position consistency, motion trajectory continuity and appearance feature similarity. In addition, a trajectory prediction-assisted alignment mechanism is introduced to adaptively adjust weights to optimize perception accuracy.

Benefits of technology

It significantly improves the accuracy of spatiotemporal alignment and target association, enhances robustness and stability in complex scenarios, and ensures the perception accuracy and reliability of the entire monitoring area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913409B_ABST
    Figure CN120913409B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of intelligent transportation, and relates to a radar and video fusion dynamic space-time alignment highway monitoring method. First, radar monitoring targets and video monitoring targets in a highway monitoring area are independently searched. Dynamic space alignment and multi-scale time synchronization are adopted to perform space-time alignment processing on the radar monitoring targets and the video monitoring targets. The radar monitoring targets and the video monitoring targets after space-time alignment are associated and matched. Whether the associated constraint condition is met is judged based on spatial position consistency, motion trajectory continuity and appearance feature similarity. The target association result is adaptively output according to an actual monitoring scene. The target association result is structured and packaged. A fusion target data packet is generated and output for monitoring feedback. The radar and video fusion perception accuracy and robustness are significantly improved in the whole scene, so that the accuracy and reliability of highway monitoring are effectively ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation technology and relates to a highway monitoring method that uses radar and video fusion for dynamic spatiotemporal alignment. Background Technology

[0002] As intelligent transportation evolves towards all-weather, high-precision, and all-element perception, the demand for multi-source heterogeneous sensor fusion in highway monitoring is becoming increasingly urgent. Radar and video, as two mainstream perception methods, possess unique advantages in motion parameter measurement and semantic feature recognition, respectively. Their collaborative application is considered a key path to overcome the performance bottleneck of single sensors.

[0003] In the existing technology, there are some solutions involving the fusion of radar and video to perform highway monitoring. For example, the method and device for highway event recognition based on multi-source traffic data published by Chinese Patent Publication No. CN117649632B focuses on using the fused features to train and infer the event recognition model. By constructing a feature fusion module and a spatiotemporal consistency semantic alignment module, it realizes event recognition of the monitoring video stream.

[0004] Another Chinese patent publication, CN116403179A, describes a vehicle holographic perception and risk behavior recognition system based on deep fusion of multi-source data from LiDAR and video. By setting up a multi-level processing architecture and integrating LiDAR and video information, it aims to achieve vehicle risk recognition and tracking in complex scenarios.

[0005] However, existing technologies still have the following limitations: 1. Existing technologies lack adaptability and accuracy in the spatiotemporal alignment of radar and video: On the one hand, the spatial transformation relationship depends on the static transformation matrix obtained in the initial calibration stage, without considering the pose drift and slight changes in installation angle caused by environmental changes during long-term operation of the equipment, which leads to the coordinate mapping gradually becoming inaccurate over time.

[0006] On the other hand, time synchronization only uses coarse-grained linear methods such as hardware timing or frame rate interpolation, ignoring the nonlinear delay accumulation between high-frequency radar updates and low-frequency video acquisition, which can cause significant timing misalignment in sudden traffic events or high-speed target tracking.

[0007] 2. Existing technologies lack robustness and intelligence in target association and matching: They mostly use simple geometric matching rules such as nearest neighbor to associate radar projection points with image detection boxes, lacking an effective multi-dimensional verification mechanism. In complex traffic scenarios with dense vehicles and mutual occlusion, they are prone to mismatch and target identity switching. In addition, they do not have the ability to make dynamic data fusion decisions that are adaptive to the scene. Summary of the Invention

[0008] In view of this, in order to solve the problems mentioned in the background technology, a highway monitoring method with dynamic spatiotemporal alignment of radar and video fusion is proposed.

[0009] The objective of this invention can be achieved through the following technical solution: This invention provides a highway monitoring method with dynamic spatiotemporal alignment of radar and video fusion, comprising: real-time acquisition of radar point cloud data and video image data within the highway monitoring area, and independent retrieval of radar monitoring targets and video monitoring targets therein.

[0010] The radar-monitored target and the video-monitored target are subjected to spatiotemporal alignment processing, including dynamic spatial alignment and multi-scale time synchronization.

[0011] The radar monitoring targets, which are aligned in time and space, are associated and matched with video monitoring targets. The association constraints are judged based on the consistency of spatial location, the continuity of motion trajectory and the similarity of appearance features. The association results are then output adaptively according to the actual monitoring scenario.

[0012] The target association results are structured and encapsulated to generate a fused target data packet, which is then output for monitoring feedback.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention performs spatiotemporal alignment processing operations including dynamic spatial alignment and multi-scale time synchronization on independently retrieved radar monitoring targets and video monitoring targets, optimizes spatial conversion accuracy in real time by establishing an error function guided by projection deviation, and eliminates micro-temporal errors by linear and nonlinear synchronization at the time video frame level, thereby significantly improving the alignment accuracy of spatial and temporal dimensions.

[0014] (2) In the process of matching radar monitoring targets and video monitoring targets, this invention judges whether they meet the association constraints based on spatial location consistency, motion trajectory continuity and appearance feature similarity, and outputs valid associated target pairs. Based on spatiotemporal alignment, it provides a basis for judging the perception scene context, which greatly enhances the accuracy of association judgment.

[0015] (3) The present invention introduces a trajectory prediction-assisted alignment mechanism. Through trajectory prediction, it effectively overcomes the problem of instantaneous matching failure caused by target occlusion and trajectory intersection, maintains the continuity of target tracking, improves the robustness of target association in high-density traffic scenarios, and ensures the stability and accuracy of fusion perception in complex scenarios.

[0016] (4) This invention effectively solves the problem of dynamic changes in equipment performance with distance by adaptively adjusting the partition weight. The near field area focuses on video detail features, while the far field area relies on radar ranging accuracy, thereby achieving optimal configuration of perception accuracy in the entire monitoring area and significantly improving the credibility of data fusion. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the implementation steps of the method of the present invention.

[0019] Figure 2 This is a flowchart illustrating the independent retrieval process for radar monitoring targets and video monitoring targets according to the present invention.

[0020] Figure 3 This is a flowchart illustrating the adaptive output of target association results in a real-world monitoring scenario according to the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 As shown, the present invention provides a highway monitoring method with dynamic spatiotemporal alignment of radar and video fusion, including: S11. Real-time acquisition of radar point cloud data and video image data within the highway monitoring area, and independent retrieval of radar monitoring targets and video monitoring targets therein.

[0023] See Figure 2 As shown, in a preferred embodiment of the present invention, the radar monitoring target retrieval process is implemented as follows: preprocessing the radar point cloud data, including filtering out noise point clouds and background point clouds.

[0024] It should be noted that the above noise point cloud filtering can be performed using statistical filtering, and background point cloud filtering can be performed using background difference method. Both are existing mature technologies and will not be elaborated on here.

[0025] Clustering algorithms are used to segment the preprocessed point cloud to distinguish independent radar monitoring targets, and a temporary identifier is assigned to each radar monitoring target.

[0026] It should be noted that the above clustering algorithm can be exemplarily adopted using the DBSCAN algorithm. The specific implementation process is as follows: draw a k-distance map with the distance from each point in the point cloud to its k-th nearest neighbor as the vertical axis and the index of the point as the horizontal axis. The point with the largest curvature in the map is retrieved as the inflection point where the slope changes significantly. The k-distance value corresponding to the inflection point is taken as the best estimate of the neighborhood radius. The sum of the neighborhood radius and 1 is taken as the minimum number of neighborhood points.

[0027] Traverse each point in the preprocessed point cloud and classify them into three categories: core points, boundary points, and noise points based on the neighborhood radius and the minimum number of neighborhood points. If the number of points contained in the neighborhood centered on a certain point is greater than or equal to the minimum number of neighborhood points, then the point is called a core point. Core points are the core of clustering and can expand to form clusters through neighborhood relationships.

[0028] If the number of points in the neighborhood centered on a certain point is less than the minimum number of points in the neighborhood, but the point falls within the neighborhood of other core points, then it is called a boundary point. A boundary point belongs to a certain cluster, but it cannot expand the cluster as a core point itself.

[0029] If a point is neither a core point nor a boundary point, it is called a noise point. This noise point is usually a small amount of noise that was not completely filtered out during preprocessing and will be removed during clustering.

[0030] Select an unclassified core point from the point cloud, and starting from the core point, incorporate all points in its neighborhood into a new cluster.

[0031] For each core point in the new cluster, continue searching for all points in its neighborhood, add these points to the current cluster, and mark these points as having been classified. Repeat this process until there are no more new core points in the current cluster that can be expanded to generate more points. At this point, a complete cluster is generated.

[0032] After generating a cluster, return to the unclassified points in the point cloud and repeat the cluster initialization and expansion process until all points are classified. During the cluster generation process, if two different clusters overlap or are associated through the neighborhood relationship of the core point, the two clusters need to be merged into one cluster to ensure that each independent target corresponds to a unique cluster.

[0033] In summary, the clustering algorithm was used to segment the preprocessed point cloud.

[0034] For each clustering result, its three-dimensional centroid coordinates in the radar coordinate system are calculated, and its kinematic parameters are analyzed based on the Doppler effect. The kinematic parameters include radial velocity, azimuth angle, and pitch angle.

[0035] It should be noted that the above calculation of the centroid coordinates in three-dimensional space is based on the average coordinates of all points within the cluster.

[0036] The radar sensor emits electromagnetic waves of a fixed frequency into the highway monitoring area. When the electromagnetic waves encounter a moving target, they are reflected by the target and generate a frequency shift. The Doppler frequency shift of all points in each cluster is retrieved, outliers are removed, and the mean is calculated to obtain the representative Doppler frequency shift of each radar-monitored target. The radial velocity is then calculated by substituting it into the existing Doppler frequency shift core formula. The specific calculation logic of this formula is to take the radial velocity as half of the product of the fixed configuration wavelength of the radar hardware and the representative Doppler frequency shift.

[0037] With the radar origin as the vertex and the positive x-axis as the reference, the angle formed by the centroid of each radar-monitored target in the horizontal plane is taken as the azimuth angle, and the angle between the radar-monitored target in the vertical plane and the positive x-axis is taken as the elevation angle. Both the azimuth and elevation angles can be obtained by using the arctangent function based on the centroid coordinates.

[0038] In a preferred embodiment of the present invention, the video surveillance target retrieval process is implemented as follows: a target detection algorithm is used to process the video image data, identify each video surveillance target in the image, and determine its category label and corresponding confidence level based on appearance and depth features.

[0039] It should be noted that the above-mentioned object detection algorithms specifically refer to object detection architectures based on deep learning, including but not limited to existing technologies such as the YOLO series or SSD models.

[0040] The model is pre-trained using a labeled dataset containing common target categories in highway monitoring scenarios, enabling the model to identify relevant targets, including vehicles, pedestrians, or non-motorized vehicles.

[0041] The video frames are input into the object detection architecture, and through forward propagation, the bounding box coordinates, class label confidence, and appearance depth feature vector of each potential object are output. Based on the preset confidence threshold, the detection results with low confidence are filtered out, and the candidate objects with high confidence are retained.

[0042] For each retained candidate target, its appearance depth feature vector is input into the classifier, which outputs the probability distribution of its belonging to each preset category.

[0043] The category with the highest probability is used as the final category label for the video surveillance target, and its confidence level is recorded.

[0044] Category label confidence is obtained based on video image surveillance quality assessment. The specific assessment process is as follows: for a video image sequence, extract and quantify the confidence level of each video surveillance target in a single frame. The average gradient magnitude, standard deviation of pixel values, and average gray value of the image are used to determine the blur, noise, and illumination anomaly factors of each video surveillance target image region. The quality degradation degree of each video surveillance target image region is obtained by linear weighted fusion. The normalized value of the quality degradation degree is used to deduct the preset confidence upper limit value in reverse to obtain the confidence degree of each video surveillance target in a single frame image.

[0045] Outline the bounding boxes of each video surveillance target and extract the center point within the box as the two-dimensional pixel coordinates of the video surveillance target in the image coordinate system.

[0046] Motion trajectory tracking is performed on each video surveillance target in the sequence of images, and the behavioral semantics of each video surveillance target are analyzed based on the trajectory data.

[0047] It should be added that the above implementation process of analyzing the behavioral semantics of each video surveillance target based on trajectory data is as follows: the trajectory data includes a position sequence consisting of the coordinates of the target center point in consecutive frames, as well as the velocity vector, acceleration, direction of motion and velocity variance derived therefrom.

[0048] The system predefines a feature rule base for typical traffic behaviors. Based on publicly available large-scale labeled traffic datasets, it uses trajectory mining and statistical analysis to associate each behavioral semantic with a set of pre-defined trajectory data indicator ranges and logical constraints. Each behavioral semantic has its key trajectory data indicator attributes labeled. The rule base divides behavioral semantics into three categories: basic motion states, direction change behaviors, and complex interaction behaviors. Basic motion states include going straight, standing still, accelerating, and decelerating; direction change behaviors include turning left, turning right, making a U-turn, and changing lanes; and complex interaction behaviors include overtaking, following, merging, and splitting.

[0049] During parsing, the real-time trajectory data of the video surveillance target is matched against the behavioral semantics in the rule base. The matching conditions include the following: the key data indicators of the target trajectory must fall within the core range defined by the behavioral semantics.

[0050] The changes in the target's motion state must conform to the typical temporal logic of the behavioral semantics.

[0051] When the above conditions are met, the current behavior of the target is determined to be the corresponding behavioral semantics.

[0052] It should be noted that the significance of the independent retrieval of radar monitoring targets and video monitoring targets in this invention lies in the fact that existing methods often adopt an alignment-then-retrieval process for radar and video, that is, first mapping the original radar point cloud onto the image plane, or performing early fusion at the feature layer. Such methods are highly dependent on the absolute accuracy of spatiotemporal alignment. Once the alignment is slightly off, the accuracy of subsequent target retrieval will drop significantly. The path adopted in this invention, which is independent retrieval first and then alignment matching, constructs a dual perception link that serves as a backup for each other. The radar and video can independently and in parallel complete their respective target detection tasks. Even if the performance of one sensor fluctuates due to temporary interference, the retrieval results of the other sensor can still serve as a reliable basis, effectively avoiding the risk of single point of failure and greatly improving the robustness and fault tolerance of the entire monitoring system.

[0053] S12. Perform spatiotemporal alignment processing on the radar monitoring target and the video monitoring target, including dynamic spatial alignment and multi-scale time synchronization.

[0054] In a preferred embodiment of the present invention, the dynamic spatial alignment is implemented as follows: initial spatial calibration is performed on the radar sensor and the camera device, and an initial spatial transformation matrix is ​​established. The initial spatial transformation matrix is ​​used to map the three-dimensional spatial coordinates in the radar coordinate system to the two-dimensional pixel coordinates in the image coordinate system.

[0055] By using an initial spatial transformation matrix, the radar monitoring target is mapped to the image coordinate system and its position is matched with that of the video monitoring target. A dynamic alignment error function is then constructed with the deviation of their projected positions on the image plane as the optimization direction.

[0056] Based on the dynamic spatial alignment error function, the parameters of the initial spatial transformation matrix are adjusted in real time using a gradient optimization method to achieve dynamic spatial alignment between radar monitoring targets and video monitoring targets.

[0057] It should be noted that the dynamic spatial alignment error function is defined as the sum of the squares of the Euclidean distances between the projection point of the centroid of the radar monitoring target and the center point of the video monitoring target on the image plane.

[0058] The numerical gradient of the dynamic spatial alignment error function with respect to each parameter of the spatial transformation matrix can be calculated by applying a small perturbation to the parameters of the spatial transformation matrix and then using the rate of change of the error function as an approximation of the gradient.

[0059] The spatial transformation matrix parameters are updated in the opposite direction of the gradient according to the preset learning rate.

[0060] The updated spatial transformation matrix is ​​applied to the spatial mapping of the next frame of data.

[0061] In a preferred embodiment of the present invention, the multi-scale time synchronization is implemented as follows: based on a unified clock source, timestamps are added to radar point cloud data, and frame sequence numbers and corresponding time information are added to video image data to establish a time sequence correspondence between timestamps and frame sequence numbers.

[0062] Based on the aforementioned time sequence correspondence, the radar data is resampled according to the sampling time of the video frame through linear interpolation to generate a radar data sequence synchronized with the video frame.

[0063] Nonlinear temporal matching and alignment are performed on the resampled radar data sequence and video image sequence. The nonlinear temporal matching is achieved by calculating the minimum cumulative alignment cost between the two sequences.

[0064] This invention embodiment performs spatiotemporal alignment processing operations, including dynamic spatial alignment and multi-scale temporal synchronization, on independently retrieved radar monitoring targets and video monitoring targets. It optimizes spatial transformation accuracy in real time by establishing an error function guided by projection deviation, and eliminates micro-temporal errors through linear and nonlinear synchronization at the temporal video frame level, thereby significantly improving the alignment accuracy in both spatial and temporal dimensions.

[0065] S13. The radar monitoring targets after spatiotemporal alignment are associated and matched with video monitoring targets. Based on the consistency of spatial location, the continuity of motion trajectory and the similarity of appearance features, it is determined whether the association constraints are met, and the target association results are adaptively output according to the actual monitoring scenario.

[0066] In a preferred embodiment of the present invention, the association and matching of the spatiotemporally aligned radar monitoring target and the video monitoring target is carried out as follows: based on the synchronization timestamp and spatial mapping relationship generated by the spatiotemporal alignment process, monitoring targets whose Euclidean distance between the centroid coordinates of the radar monitoring target projected onto the image plane and the center point coordinates of the video monitoring target at the same time is less than a preset threshold are retrieved, and potential target pairs are formed.

[0067] For the potential target pairs, spatial position consistency verification, motion trajectory continuity verification, and appearance feature similarity verification are performed sequentially.

[0068] If all checks pass, the pair is considered a valid associated target pair.

[0069] If the target density within the perceived highway monitoring area reaches the preset high-density standard, the trajectory prediction-assisted alignment mechanism is activated to predict the trajectory position of each video monitoring target at the next moment, and based on the predicted trajectory, the potential target pairs that failed the verification due to target occlusion or trajectory intersection are re-associated and matched.

[0070] It should be noted that the above trajectory prediction auxiliary alignment mechanism is specifically performed using the Kalman filter algorithm.

[0071] This invention introduces a trajectory prediction-assisted alignment mechanism. Through trajectory prediction, it effectively overcomes the problem of instantaneous matching failure caused by target occlusion and trajectory intersection, maintains the continuity of target tracking, improves the robustness of target association in high-density traffic scenarios, and ensures the stability and accuracy of fusion perception in complex scenarios.

[0072] In a preferred embodiment of the present invention, the spatial position consistency verification process includes: projecting the three-dimensional coordinate sequence of the potential target aligned with the radar monitoring target onto the image plane to generate a predicted bounding box.

[0073] The set of intersection-union ratios (IoU) of the predicted bounding box and the bounding boxes of potential targets in a video surveillance target pair within a video image sequence.

[0074] The cross-union ratio (CUNR) threshold is dynamically selected based on the target category and its video pixel size. The proportion of frames in the video image sequence whose CUNR exceeds the threshold is counted. If the proportion of frames reaches the preset percentage requirement, the spatial position consistency check is deemed to be qualified.

[0075] It should be noted that the basis for dynamically selecting the intersection-union ratio (IU) threshold is that targets with different attributes have different tolerances for spatial position consistency. The logic for dynamically adjusting the judgment criteria based on target size is as follows: the pixel size of the target on the image plane, i.e. the area of ​​the target bounding box. If the pixel size is less than 100, it is judged as a small target. Even if the radar and video positioning are completely consistent, the IU of a small target may be low due to the small number of bounding box pixels. If the pixel size is greater than 1000, it is judged as a large target. When the pixel size is between 100 and 1000, it is judged as a medium target. The target size should reflect a gradient relationship where the larger the size, the stricter the matching requirements.

[0076] The logic for dynamically adjusting the judgment criteria based on target category is as follows: different categories of targets differ in radar positioning accuracy and video detection bounding box stability. For example, pedestrian targets are affected by posture changes, resulting in large fluctuations in the video bounding box, so the requirement for the intersection-union ratio (IU) threshold can be relaxed. Motor vehicle targets have stable shapes, and both radar and video positioning accuracy are high, so the requirement for the IU threshold should be more stringent.

[0077] The cross-union ratio threshold is selected by combining the target category and its video pixel size.

[0078] In a preferred embodiment of the present invention, the motion trajectory continuity verification process includes: extracting the behavioral semantics of the video surveillance target in the potential target pair, and analyzing the expected performance trend of the kinematic parameters corresponding to the behavioral semantics.

[0079] It should be added that the analysis of the correspondence between behavioral semantics and the expected performance trend of kinematic parameters fundamentally relies on the combined effect of physical constraints, traffic rules, and driving intentions. This can be established through one or more combinations of three methods: physical kinematic modeling, data-driven statistical learning, and expert knowledge rule bases. Specifically, physical modeling is based on Newton's laws of motion and vehicle dynamics principles; the data-driven method is based on clustering and regression analysis of labeled traffic scene data; and the expert knowledge method is based on the rule-based expression of traffic rules and driving behavior research.

[0080] Based on the kinematic parameters measured by radar, the expected performance trends of different behavioral semantics can be exemplified as follows: if the behavioral semantic is straight, the variances of radial velocity, azimuth angle, and pitch angle are all at low levels, and the kinematic parameters maintain relative temporal consistency.

[0081] When turning, the radial velocity tends to decelerate, the azimuth angle changes monotonically, which can be specifically manifested as an increase when turning left and a decrease when turning right, and the pitch angle has a time-series deviation depending on the turning slope.

[0082] The kinematic parameter performance trend of the potential target being monitored by radar is quantified within the corresponding video image sequence time period, and the matching degree with the expected performance trend is calculated.

[0083] It should be noted that in the above matching degree calculation process, the indicators of the performance trend of each index in the kinematic parameters need to be formed into a vector and substituted into the cosine similarity calculation formula to obtain the final matching degree value.

[0084] If the matching degree reaches a high standard, the continuity of the motion trajectory is deemed to have passed the verification.

[0085] It should be added that a high standard for matching degree is defined as the trend matching degree of the current target being no less than the 95th percentile in the baseline distribution composed of historical normal trajectory data. This means that the consistency of the current target's movement trend must be better than 95% of historical normal cases.

[0086] In a preferred embodiment of the present invention, the appearance feature similarity process includes: extracting the point cloud size of the potential target to be monitored by the radar.

[0087] The point cloud size is compared with the pre-stored reasonable vehicle size range of the category to which the video surveillance target belongs in the potential target pair.

[0088] If the point cloud size is within the specified range, the appearance feature similarity check is deemed to be successful.

[0089] It should be noted that the core logic of the appearance feature similarity verification process is cross-modal physical consistency verification. It makes full use of the precise measurement capabilities of radar and the semantic recognition capabilities of video. Through size-category rationality checks, it provides an important and reliable constraint for multi-target association. It is a verification method based on the objective laws of the physical world and a key technical means to ensure the accuracy of association in multi-sensor analysis.

[0090] This invention, through the process of matching radar monitoring targets with video monitoring targets, determines whether they meet the association constraints based on spatial location consistency, motion trajectory continuity, and appearance feature similarity, and outputs valid associated target pairs. It provides a basis for judging the perception scene context on the basis of spatiotemporal alignment, greatly enhancing the accuracy of association decisions.

[0091] See Figure 3 As shown in a preferred embodiment of the present invention, the adaptive output of target association results based on the actual monitoring scenario includes: dividing the highway monitoring area into a near-field zone, a mid-field zone, and a far-field zone according to the distance between the target and the radar sensor.

[0092] For example, the near-field zone can be defined as 0 to 50 meters from the radar, the mid-field zone as 50 to 150 meters, and the far-field zone as 150 to 500 meters.

[0093] For different field areas, the weights of radar and video in the fusion decision are dynamically adjusted based on the confidence level of video category labels.

[0094] It should be added that the weight of video data in the near-field area is greater than that of radar data. For example, the weight of video data is selected as 0.8 and the weight of radar data is 0.2.

[0095] The midfield dynamically calculates the fusion weights of radar and video data based on the confidence level of the video category label.

[0096] In the far-field region, the weight of radar data is greater than that of video data. For example, the weight of video data is selected as 0.2 and the weight of radar data is selected as 0.8.

[0097] The above-mentioned dynamic calculation process for the mid-range is as follows: Establish the basic boundary range of the video data weights, for example... The radar data weights are set in reverse intervals, and high and low confidence thresholds are set as decision boundaries for weight allocation. A piecewise linear interpolation function is used to obtain the smooth allocation weights under the decision attributes corresponding to the video category label confidence. The video category label confidence is obtained by performing double mean calculation on the confidence of each video monitoring target in each frame of the video image sequence.

[0098] Based on the weights, radar point cloud data and video image data of effectively associated target pairs are fused to obtain target association results.

[0099] This invention effectively solves the problem of dynamic changes in device performance with distance by adaptively adjusting partition weights. The near-field region focuses on video detail features, while the far-field region relies on radar ranging accuracy, thereby achieving optimal configuration of perception accuracy across the entire monitoring area and significantly improving the reliability of data fusion.

[0100] S14. The target association results are structured and encapsulated to generate a fused target data packet and output it for monitoring feedback.

[0101] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A highway monitoring method based on dynamic spatiotemporal alignment of radar and video fusion, characterized in that, include: Real-time acquisition of radar point cloud data and video image data within the highway monitoring area, and independent retrieval of radar monitoring targets and video monitoring targets; The radar-monitored target and the video-monitored target are subjected to spatiotemporal alignment processing, including dynamic spatial alignment and multi-scale time synchronization; The radar monitoring targets after spatiotemporal alignment are associated and matched with video monitoring targets. The association constraints are judged based on spatial location consistency, motion trajectory continuity and appearance feature similarity. The target association results are adaptively output according to the actual monitoring scenario. The target association results are structured and encapsulated to generate a fused target data packet, which is then output for monitoring feedback. The radar monitoring target retrieval process is implemented as follows: the radar point cloud data is preprocessed, including filtering out noisy point clouds and background point clouds; a clustering algorithm is used to segment the preprocessed point cloud to distinguish independent radar monitoring targets, and a temporary identifier is assigned to each radar monitoring target; For each clustering result, its three-dimensional centroid coordinates in the radar coordinate system are calculated, and its kinematic parameters are analyzed based on the Doppler effect. The kinematic parameters include radial velocity, azimuth angle, and elevation angle. The video surveillance target retrieval process is implemented as follows: a target detection algorithm is used to process the video image data, identify each video surveillance target in the image, and determine its category label and corresponding confidence level based on appearance and depth features; the bounding box of each video surveillance target is delineated and the center point inside the box is extracted as the two-dimensional pixel coordinates of the video surveillance target in the image coordinate system; the motion trajectory of each video surveillance target in the sequence of images is tracked, and the behavioral semantics of each video surveillance target are analyzed based on the trajectory data; The dynamic spatial alignment is implemented as follows: initial spatial calibration is performed on the radar sensor and the camera device, and an initial spatial transformation matrix is ​​established. The initial spatial transformation matrix is ​​used to map the three-dimensional spatial coordinates in the radar coordinate system to the two-dimensional pixel coordinates in the image coordinate system. Through the initial spatial transformation matrix, the radar monitoring target is mapped to the image coordinate system, and position matching is performed with the video monitoring target. A dynamic alignment error function is constructed with the deviation of the projected positions of the two on the image plane as the optimization direction. Based on the dynamic spatial alignment error function, the parameters of the initial spatial transformation matrix are adjusted in real time using a gradient optimization method to achieve dynamic spatial alignment between radar monitoring targets and video monitoring targets. The process of associating and matching the spatiotemporally aligned radar monitoring targets with video monitoring targets is implemented as follows: Based on the synchronization timestamps and spatial mapping relationships generated by the spatiotemporal alignment process, monitoring targets whose Euclidean distance between the centroid coordinates of the radar monitoring target projected onto the image plane and the center point coordinates of the video monitoring target at the same time is less than a preset threshold are retrieved, forming potential target pairs; for the potential target pairs, spatial position consistency verification, motion trajectory continuity verification, and appearance feature similarity verification are performed sequentially; If all verifications pass, the target pairs are deemed valid. If the target density within the perceived highway monitoring area reaches the preset high-density standard, the trajectory prediction-assisted alignment mechanism is activated to predict the trajectory position of each video monitoring target at the next moment. Based on the predicted trajectory pairs, potential target pairs that fail verification due to target occlusion or trajectory intersection are re-associated and matched.

2. The highway monitoring method based on radar and video fusion with dynamic spatiotemporal alignment according to claim 1, characterized in that: The multi-scale time synchronization is implemented as follows: Based on a unified clock source, timestamps are added to radar point cloud data, and frame sequence numbers and corresponding time information are added to video image data to establish a time sequence correspondence between timestamps and frame sequence numbers. Based on the aforementioned time sequence correspondence, the radar data is resampled according to the sampling time of the video frame through linear interpolation to generate a radar data sequence synchronized with the video frame; Nonlinear temporal matching and alignment are performed on the resampled radar data sequence and video image sequence. The nonlinear temporal matching is achieved by calculating the minimum cumulative alignment cost between the two sequences.

3. The highway monitoring method based on radar and video fusion with dynamic spatiotemporal alignment according to claim 1, characterized in that: The spatial location consistency verification process includes: The three-dimensional coordinate sequence of the potential target being aligned with the radar target is projected onto the image plane to generate a predicted bounding box; The set of intersection-union ratios (IUR) of the predicted bounding boxes and the bounding boxes of potential targets in the video surveillance target pair within a video image sequence; The cross-union ratio (CUNR) threshold is dynamically selected based on the target category and its video pixel size. The proportion of frames in the video image sequence whose CUNR exceeds the threshold is counted. If the proportion of frames reaches the preset percentage requirement, the spatial position consistency check is deemed to be qualified.

4. The highway monitoring method based on radar and video fusion with dynamic spatiotemporal alignment according to claim 2, characterized in that: The motion trajectory continuity verification process includes: Extract the behavioral semantics of potential target pairs in video surveillance, and analyze the expected performance trend of the kinematic parameters corresponding to the behavioral semantics; The kinematic parameter performance trend of the potential target-mid-target radar monitoring target in the corresponding video image sequence time period is quantified, and the matching degree with the expected performance trend is calculated. If the matching degree reaches a high standard, the continuity of the motion trajectory is deemed to have passed the verification.

5. The highway monitoring method based on radar and video fusion with dynamic spatiotemporal alignment according to claim 2, characterized in that: The appearance feature similarity verification process includes: Extract the point cloud size of potential targets for radar monitoring; The point cloud size is compared with the reasonable vehicle size range pre-stored for the category to which the video surveillance target belongs in the potential target pair; If the point cloud size is within the specified range, the appearance feature similarity check is deemed to be successful.

6. The highway monitoring method based on radar and video fusion with dynamic spatiotemporal alignment according to claim 1, characterized in that: The adaptive output of target association results based on the actual monitoring scenario includes: Based on the distance between the target and the radar sensor, the highway monitoring area is divided into near-field, mid-field, and far-field zones; For different field areas, the weights of radar and video in the fusion decision are dynamically adjusted based on the confidence level of video category labels; Based on the weights, radar point cloud data and video image data of effectively associated target pairs are fused to obtain target association results.

Citation Information

Patent Citations

  • Vehicle holographic perception and risk behavior identification system based on thunder multi-source data deep fusion

    CN116403179A

  • Highway event recognition method and recognition device based on multi-source traffic data

    CN117649632B

  • Expressway multi-target tracking method based on Leiyu fusion

    CN116863382A

  • Method for tracking multiple objects

    EP4439478A1