Radar and video fusion dynamic space-time alignment road monitoring method

By independently retrieving and dynamically aligning radar and video surveillance targets, combined with multi-scale time synchronization and partition weight adjustment, the problems of spatiotemporal alignment accuracy and target association robustness in existing radar and video fusion monitoring technologies have been solved, achieving high-precision and stable highway monitoring results.

CN120913409AActive Publication Date: 2025-11-07SHENZHEN ARCKE INNOVATION TECH CO LTD

Patent Information

Application Number
CN202511415299.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-07
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing technologies lack sufficient spatiotemporal alignment accuracy, inaccurate time synchronization, robustness and intelligence in target association and matching, especially in complex traffic scenarios where mismatches and target identity switching are prone to occur.

Method used

By acquiring radar point cloud data and video image data in real time, the system independently retrieves radar and video surveillance targets, performs dynamic spatial alignment and multi-scale time synchronization, and performs association matching based on spatial position consistency, motion trajectory continuity and appearance feature similarity. It also introduces a trajectory prediction-assisted alignment mechanism, combined with adaptive adjustment of partition weights, to output target association results.

Benefits of technology

It significantly improves the accuracy of spatiotemporal alignment and target association, enhances robustness and stability in complex scenarios, and ensures the perception accuracy and reliability of the entire monitoring area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913409A_ABST
    Figure CN120913409A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent traffic, and relates to a radar and video fusion dynamic space-time alignment road monitoring method, which comprises the following steps of: firstly, independently retrieving a radar monitoring target and a video monitoring target in a road monitoring area, and adopting a dynamic space alignment and multi-scale time synchronization mode to obtain a radar monitoring target and a video monitoring target; performing space-time alignment processing on the radar monitoring target and the video monitoring target, performing association matching on the radar monitoring target and the video monitoring target after space-time alignment, and judging whether association constraint conditions are met or not based on spatial position consistency, movement track continuity and appearance feature similarity; according to the method and the device, a target association result is adaptively output according to an actual monitoring scene, structured packaging is carried out on the target association result, a fusion target data packet is generated and output for monitoring feedback, and the radar and video fusion perception precision and robustness in a full scene are remarkably improved, so that the accuracy and reliability of road monitoring are effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent transportation, and relates to a highway monitoring method of radar and video fusion dynamic space-time alignment. BACKGROUND

[0002] With the evolution of intelligent transportation towards all-weather, high-precision and full-factor perception, the demand for multi-source heterogeneous sensor fusion of highway monitoring is increasingly urgent. Radar and video, as the two main sensing means at present, have unique advantages in motion parameter measurement and semantic feature recognition respectively, and their collaborative application is considered as a key path to break through the performance bottleneck of single sensor.

[0003] In the prior art, there are some solutions related to radar and video fusion to perform highway monitoring, for example, a highway event identification method and identification device based on multi-source traffic data disclosed in Chinese Patent No. CN117649632B, which focuses on training and inference of an event identification model using fused features, and realizes event identification of a monitoring video stream by constructing a feature fusion module and a space-time consistency semantic alignment module.

[0004] In addition, a vehicle holographic perception and risk behavior identification system based on radar and video multi-source data deep fusion disclosed in Chinese Patent No. CN116403179A, which fuses laser radar and video information through a multi-level processing architecture, aims to realize vehicle risk identification and tracking in complex scenes.

[0005] However, the prior art has the following limitations, specifically: 1. The existing technology is insufficient in terms of space-time alignment of radar and video: on the one hand, the spatial conversion relationship depends on the static conversion matrix obtained in the initial calibration stage, and does not consider the pose drift and installation angle change caused by environmental changes during long-term operation of the device, resulting in gradual misalignment of coordinate mapping over time.

[0006] On the other hand, time synchronization only uses coarse linear methods such as hardware timing or frame rate interpolation, ignoring the nonlinear delay accumulation between radar high-frequency updates and video low-frequency acquisition, which causes significant time sequence misalignment in sudden traffic events or high-speed target tracking.

[0007] 2. The existing technology lacks robustness and intelligence in target association matching: simple geometric matching rules such as nearest neighbor are mostly used to associate radar projection points with image detection boxes, lacking effective multi-dimensional verification mechanism, and in complex traffic scenes with dense vehicles and mutual occlusion, false matching and target identity switching are prone to occur, in addition, it does not have the ability of dynamic data fusion decision-making in scene adaptation. SUMMARY

[0008] In view of this, in order to solve the problems raised in the background art, a radar and video fusion dynamic space-time alignment highway monitoring method is proposed.

[0009] The purpose of the application can be realized by the following technical solutions: The application provides a radar and video fusion dynamic space-time alignment highway monitoring method, which comprises: collecting radar point cloud data and video image data in a highway monitoring area in real time, and independently retrieving radar monitoring targets and video monitoring targets.

[0010] Performing space-time alignment processing on the radar monitoring targets and the video monitoring targets, including dynamic space alignment and multi-scale time synchronization.

[0011] Correlating and matching the radar monitoring targets and the video monitoring targets after space-time alignment, judging whether the correlation constraint condition is met based on spatial position consistency, motion trajectory continuity and appearance feature similarity, and outputting the target correlation result adaptively according to the actual monitoring scene.

[0012] Structuring and packaging the target correlation result, generating a fusion target data packet and outputting it for monitoring feedback.

[0013] Compared with the prior art, the application has the following advantages: (1) The application performs space-time alignment processing operation including dynamic space alignment and multi-scale time synchronization on the independently retrieved radar monitoring targets and video monitoring targets, establishes an error function oriented to projection deviation to optimize the spatial conversion precision in real time, eliminates micro timing errors through linear and nonlinear synchronization at the video frame level, and significantly improves the alignment accuracy in space and time dimensions.

[0014] (2) In the correlation matching process of the radar monitoring targets and the video monitoring targets, the application judges whether the correlation constraint condition is met based on spatial position consistency, motion trajectory continuity and appearance feature similarity, outputs effective correlation target pairs, provides a judgment basis for perceiving the context of the scene on the basis of space-time alignment, and greatly enhances the accuracy of the correlation decision.

[0015] (3) The application introduces a trajectory prediction assisted alignment mechanism, effectively overcomes the problem of instantaneous matching failure caused by target occlusion and trajectory intersection, maintains the continuity of target tracking, improves the robustness of target correlation in high-density traffic scenes, and ensures the stability and accuracy of fusion perception in complex scenes.

[0016] (4) The application effectively solves the problem of dynamic change of device performance with distance by adaptive adjustment of partition weight, focuses on video detail features in the near-field area, and relies on radar ranging accuracy in the far-field area, so as to realize optimal configuration of perception accuracy in the whole monitoring area, and significantly improve the data fusion reliability. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the following description of the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0018] Figure 1 The flowchart of the method embodiment of the present application is implemented.

[0019] Figure 2 The flowchart of the radar monitoring target and the video monitoring target independent retrieval process of the present application is shown.

[0020] Figure 3 The flowchart of the actual monitoring scene adaptive output target association result of the present application is shown. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort belong to the protection scope of the present application.

[0022] Please refer to Figure 1 The present application provides a radar and video fusion dynamic space-time alignment highway monitoring method, as shown in the figure, which comprises the following steps: S11. Real-time acquisition of radar point cloud data and video image data in a highway monitoring area, and independent retrieval of radar monitoring targets and video monitoring targets therein.

[0023] Please refer to Figure 2 In a preferred embodiment of the present application, the radar monitoring target retrieval process is implemented as follows: the radar point cloud data is preprocessed, including filtering out noise point cloud and background point cloud.

[0024] It should be noted that the noise point cloud filtering can be implemented by statistical filtering, and the background point cloud filtering can be implemented by background difference method, both of which are mature technologies and will not be described here.

[0025] The preprocessed point cloud is segmented by using a clustering algorithm to distinguish independent radar monitoring targets, and a temporary identifier is assigned to each radar monitoring target.

[0026] It should be noted that the above clustering algorithm can exemplarily adopt DBSCAN algorithm, and the specific implementation process is: taking the distance from each point in the point cloud to its kth nearest neighbor as the vertical axis, and taking the index of the point as the horizontal axis, a k-distance graph is drawn, the point with the maximum curvature in the graph is searched as the inflection point where the slope changes obviously, and the k-distance value corresponding to the inflection point is taken as the best estimation value of the neighborhood radius, and the sum of the neighborhood radius and 1 is taken as the minimum number of neighborhood points.

[0027] Each point in the preprocessed point cloud is traversed, and the points are divided into three categories of core points, boundary points and noise points according to the neighborhood radius and the minimum number of neighborhood points: if the number of points contained in the neighborhood with a certain point as the center is greater than or equal to the minimum number of neighborhood points, the point is called a core point, and the core point is the core of clustering and can be expanded to form a cluster through neighborhood relationship.

[0028] If the number of points contained in the neighborhood with a certain point as the center is less than the minimum number of neighborhood points, but the point falls within the neighborhood of other core points, the point is called a boundary point, and the boundary point belongs to a certain cluster, but cannot expand the cluster as a core point.

[0029] If a point is neither a core point nor a boundary point, the point is called a noise point. This noise point is usually a small amount of noise that is not completely filtered out in the preprocessing process and will be excluded in the clustering process.

[0030] A core point that has not been classified is selected from the point cloud, and all points in the neighborhood of the core point are included in a new cluster.

[0031] For each core point in the new cluster, continue to search for all points in its neighborhood, add these points to the current cluster, and mark these points as classified. Repeat this process until there are no more core points in the current cluster that can expand more points. At this time, a complete cluster is generated.

[0032] After completing the generation of a cluster, return to the points in the point cloud that have not been classified, and repeat the initialization and expansion of the cluster described above until all points are classified. In the process of generating clusters, if two different clusters overlap or are associated through the neighborhood relationship of core points, the two clusters need to be merged into one cluster to ensure that each independent target corresponds to a unique cluster.

[0033] The above completes the clustering algorithm for segmenting the preprocessed point cloud.

[0034] For each clustering result, the three-dimensional spatial centroid coordinates in the radar coordinate system are calculated, and the kinematic parameters based on the Doppler effect are analyzed, including radial velocity, azimuth angle and pitch angle.

[0035] It should be noted that the calculation of the three-dimensional spatial centroid coordinates is based on the mean value of the coordinates of all points in the cluster.

[0036] The radar sensor transmits electromagnetic waves of a fixed frequency to the highway monitoring area. When the electromagnetic waves encounter a moving target, they are reflected by the target and produce a frequency shift. The Doppler shift of all points in each cluster is retrieved, outliers are removed, and the mean value is calculated to obtain the representative Doppler shift of each radar monitoring target. The radial velocity is obtained by substituting the representative Doppler shift into the existing Doppler shift core formula. The specific calculation logic of this formula is to take the radial velocity as half the product of the fixed configuration wavelength of the radar hardware and the representative Doppler shift.

[0037] With the radar origin as the vertex and the positive direction of the x-axis as the reference, the angle formed by the centroid of each radar monitoring target in the horizontal plane is taken as the azimuth angle, and the angle between each radar monitoring target in the vertical plane of the radar and the positive direction of the x-axis is taken as the elevation angle. Both the azimuth angle and the elevation angle can be obtained by using the inverse tangent function based on the centroid coordinates.

[0038] In a preferred embodiment of the present application, the video monitoring target retrieval process is implemented as follows: a target detection algorithm is used to process the video image data, identify each video monitoring target in the image, and determine its class label and corresponding confidence based on the appearance depth feature.

[0039] It should be noted that the above-mentioned target detection algorithm specifically refers to a target detection architecture based on deep learning, including but not limited to existing technologies such as YOLO series or SSD model.

[0040] A labeled data set containing common target classes in highway monitoring scenes is used to pre-train the model, so that the model has the ability to recognize related targets. The common target classes include vehicles, pedestrians, or non-motor vehicles.

[0041] The video frame is input into the target detection architecture, and the boundary box coordinates, class label confidence, and appearance depth feature vector of each potential target are output through forward propagation calculation: according to the preset confidence threshold, filter out the detection results with low confidence, and retain the candidate targets with high confidence.

[0042] For each retained candidate target, its appearance depth feature vector is input into the classifier, and the probability distribution of its belonging to each preset class is output.

[0043] The class with the highest probability is taken as the final class label of the video monitoring target, and its confidence is recorded.

[0044] The class label confidence is obtained based on the video image monitoring quality assessment. The specific assessment process is as follows: for a video image sequence, the appearance depth feature vector of each video monitoring target in a single frame image is quantized and extracted, and the class label confidence of each video monitoring target is calculated based on the appearance depth feature vector. The gradient amplitude mean value, pixel value standard deviation, and image average gray value determine the blur, noise, and light abnormality factors of each video monitoring target image region. The quality degradation degree of each video monitoring target image region is obtained by linearly weighting and fusing. The confidence of each video monitoring target in a single frame image is obtained by inversely deducting the pre-set upper limit value of the confidence from the normalized value of the quality degradation degree.

[0045] The boundary box of each video monitoring target is outlined, and the center point in the box is extracted as the two-dimensional pixel coordinates of the video monitoring target in the image coordinate system.

[0046] The motion trajectory of each video monitoring target in the sequence image is tracked, and the behavior semantics of each video monitoring target is analyzed based on the trajectory data.

[0047] It should be noted that the above implementation process of analyzing the behavior semantics of each video monitoring target based on the trajectory data is as follows: the trajectory data includes a position sequence composed of the center point coordinates of the target in consecutive frames, as well as the derived speed vector, acceleration, motion direction, and speed variance.

[0048] A feature rule base of typical traffic behaviors is predefined, which is based on a large public labeled traffic data set. Through trajectory mining and statistical analysis, each behavior semantics is associated with a set of pre-set trajectory data index intervals and logical constraints, and each behavior semantics has been labeled with its key trajectory data index attributes. The rule base divides the behavior semantics into three categories: basic motion state, direction change behavior, and complex interaction behavior. The basic motion state includes straight, stationary, acceleration, and deceleration. The direction change behavior includes left turn, right turn, U-turn, and lane change. The complex interaction behavior includes overtaking, following, merging, and splitting.

[0049] In the analysis, the real-time trajectory data of the video monitoring target is matched with the behavior semantics in the rule base. The matching conditions include the following aspects: the key data indicators of the target trajectory must fall within the core interval defined by the behavior semantics.

[0050] The motion state change of the target must comply with the typical time sequence logic of the behavior semantics.

[0051] When the above conditions are met, it is determined that the current behavior of the target is the corresponding behavior semantics.

[0052] It should be noted that the present application aims at the independent retrieval of radar monitoring targets and video monitoring targets, and the significance thereof lies in that: in the prior art, the radar and the video are often subjected to a process of alignment first and then retrieval, that is, the original radar point cloud is first mapped to the image plane, or early fusion is performed at the feature layer, and such a method highly depends on the absolute accuracy of the time-space alignment, and once the alignment is slightly deviated, the accuracy of the subsequent target retrieval will be greatly reduced. The path of the present application, that is, independent retrieval first and then alignment and matching, constructs a dual perception link as a backup, and the radar and the video can independently and in parallel complete the target detection tasks that they are best at. Even if one sensor fluctuates in performance due to temporary interference, the retrieval result of the other sensor can still be used as a reliable basis, effectively avoiding the risk of single-point failure, and greatly improving the robustness and fault tolerance of the entire monitoring system.

[0053] S12. Perform time-space alignment processing on the radar monitoring target and the video monitoring target, including dynamic spatial alignment and multi-scale time synchronization.

[0054] In a preferred embodiment of the present application, the dynamic spatial alignment is implemented as follows: the radar sensor and the camera device are subjected to initial spatial calibration, and an initial spatial conversion matrix is established, which is used to map the three-dimensional spatial coordinates in the radar coordinate system to the two-dimensional pixel coordinates in the image coordinate system.

[0055] The radar monitoring target is mapped to the image coordinate system by the initial spatial conversion matrix, and the position matching is performed with the video monitoring target, and a dynamic alignment error function is constructed with the position deviation of the projection of the two targets on the image plane as the optimization direction.

[0056] Based on the dynamic spatial alignment error function, the parameters of the initial spatial conversion matrix are adjusted in real time by using a gradient optimization method, and the dynamic spatial alignment of the radar monitoring target and the video monitoring target is completed.

[0057] It should be noted that the dynamic spatial alignment error function is defined as the sum of the squared Euclidean distances of the projection points of the centers of the radar monitoring targets and the center points of the video monitoring targets on the image plane.

[0058] The numerical gradient of the dynamic spatial alignment error function with respect to each parameter of the spatial conversion matrix is calculated, and the calculation method can be exemplarily by applying a small perturbation to the parameters of the spatial conversion matrix, and calculating the change rate of the error function as the gradient approximation value.

[0059] The parameters of the spatial conversion matrix are updated in the opposite direction of the gradient at a preset learning rate.

[0060] The updated spatial conversion matrix is applied to the spatial mapping of the next frame of data.

[0061] In a preferred embodiment of the present application, the multi-scale time synchronization is implemented as follows: based on a unified clock source, time stamps are added to the radar point cloud data, frame numbers and corresponding time information are added to the video image data, and a time sequence correspondence between the time stamps and the frame numbers is established.

[0062] According to the time sequence correspondence, the radar data is resampled at the video frame sampling time through linear interpolation to generate a radar data sequence synchronized with the video frames.

[0063] The resampled radar data sequence and the video image sequence are subjected to nonlinear time sequence matching and alignment, and the nonlinear time sequence matching is achieved by calculating the minimum cumulative alignment cost between the two sequences.

[0064] In the embodiments of the present application, by performing spatio-temporal alignment processing operation on the independently searched radar monitoring targets and video monitoring targets, including dynamic spatial alignment and multi-scale time synchronization, by establishing an error function oriented by projection deviation to optimize the spatial conversion precision in real time, and by linear and nonlinear synchronization at the video frame level to eliminate micro temporal sequence errors, the alignment precision in the spatial and temporal dimensions is significantly improved.

[0065] S13. The radar monitoring targets and the video monitoring targets after spatio-temporal alignment are associated and matched, and it is judged whether they meet the association constraint conditions based on spatial position consistency, motion trajectory continuity and appearance feature similarity, and the target association result is adaptively output according to the actual monitoring scene.

[0066] In a preferred embodiment of the present application, the radar monitoring targets and the video monitoring targets after spatio-temporal alignment are associated and matched as follows: based on the synchronization time stamps and the spatial mapping relationship generated by the spatio-temporal alignment processing, the radar monitoring targets and the video monitoring targets whose Euclidean distances between the center of mass coordinates of the radar monitoring targets projected onto the image plane and the center point coordinates of the video monitoring targets at the same time are less than a preset threshold are searched to form potential target pairs.

[0067] The spatial position consistency check, the motion trajectory continuity check and the appearance feature similarity check are sequentially performed for the potential target pairs.

[0068] If all the checks are qualified, the potential target pairs are determined as valid associated target pairs.

[0069] If the target density in the perception highway monitoring area reaches a preset high density standard, a trajectory prediction assisted alignment mechanism is started, the trajectory positions of the video monitoring targets at the next time are predicted, and the potential target pairs that fail to pass the check due to target occlusion or trajectory intersection are re-associated and matched based on the predicted trajectories.

[0070] It should be noted that the trajectory prediction assisted alignment mechanism is specifically implemented by using Kalman filter algorithm.

[0071] This invention introduces a trajectory prediction-assisted alignment mechanism. Through trajectory prediction, it effectively overcomes the problem of instantaneous matching failure caused by target occlusion and trajectory intersection, maintains the continuity of target tracking, improves the robustness of target association in high-density traffic scenarios, and ensures the stability and accuracy of fusion perception in complex scenarios.

[0072] In a preferred embodiment of the present invention, the spatial position consistency verification process includes: projecting the three-dimensional coordinate sequence of the potential target aligned with the radar monitoring target onto the image plane to generate a predicted bounding box.

[0073] The set of intersection-union ratios (IoU) of the predicted bounding box and the bounding boxes of potential targets in a video surveillance target pair within a video image sequence.

[0074] The cross-union ratio (CUNR) threshold is dynamically selected based on the target category and its video pixel size. The proportion of frames in the video image sequence whose CUNR exceeds the threshold is counted. If the proportion of frames reaches the preset percentage requirement, the spatial position consistency check is deemed to be qualified.

[0075] It should be noted that the basis for dynamically selecting the intersection-union ratio (IU) threshold is that targets with different attributes have different tolerances for spatial position consistency. The logic for dynamically adjusting the judgment criteria based on target size is as follows: the pixel size of the target on the image plane, i.e. the area of ​​the target bounding box. If the pixel size is less than 100, it is judged as a small target. Even if the radar and video positioning are completely consistent, the IU of a small target may be low due to the small number of bounding box pixels. If the pixel size is greater than 1000, it is judged as a large target. When the pixel size is between 100 and 1000, it is judged as a medium target. The target size should reflect a gradient relationship where the larger the size, the stricter the matching requirements.

[0076] The logic for dynamically adjusting the judgment criteria based on target category is as follows: different categories of targets differ in radar positioning accuracy and video detection bounding box stability. For example, pedestrian targets are affected by posture changes, resulting in large fluctuations in the video bounding box, so the requirement for the intersection-union ratio (IU) threshold can be relaxed. Motor vehicle targets have stable shapes, and both radar and video positioning accuracy are high, so the requirement for the IU threshold should be more stringent.

[0077] The cross-union ratio threshold is selected by combining the target category and its video pixel size.

[0078] In a preferred embodiment of the present invention, the motion trajectory continuity verification process includes: extracting the behavioral semantics of the video surveillance target in the potential target pair, and analyzing the expected performance trend of the kinematic parameters corresponding to the behavioral semantics.

[0079] It needs to be added that the analysis of the correspondence between the behavior semantics and the expected performance trend of the kinematic parameters is fundamentally based on the joint action of physical constraints, traffic rules and driving intentions, which can be established by one or more combinations of three methods: physical kinematic modeling, data-driven statistical learning and expert knowledge rule base. Among them, the physical modeling is based on Newton's law of motion and vehicle dynamics principles, the data-driven method is based on clustering and regression analysis of labeled traffic scene data, and the expert knowledge method is based on the rule-based expression of traffic rules and driving behavior research.

[0080] Based on the kinematic parameters measured by the radar, the expected performance trend of different behavior semantics can be exemplified as follows: if the behavior semantics is straight, the variances of radial velocity, azimuth angle and pitch angle are all at a low level, and the kinematic parameters remain relatively consistent in time sequence.

[0081] If it is a turn, there is a deceleration trend in radial velocity, and the azimuth angle continuously changes monotonously, which can be specifically manifested as increasing when turning left and decreasing when turning right, and the pitch angle depends on the time deviation of the turning slope.

[0082] The kinematic parameter performance trend of the radar-monitored target in the potential target pair in the corresponding video image sequence period is quantified, and the matching degree is calculated with the expected performance trend.

[0083] It needs to be noted that in the above matching degree calculation process, the indications of the performance trend of each index in the kinematic parameters are constituted into a vector, which is substituted into the cosine similarity calculation formula to obtain the final matching degree value.

[0084] If the matching degree reaches a high standard, it is determined that the motion trajectory continuity verification is qualified.

[0085] It needs to be added that the high standard of the matching degree is defined as the trend matching degree of the current target being not lower than the 95th percentile in the reference distribution constituted by the historical normal trajectory data. This means that the motion trend consistency of the current target needs to be better than 95% of the historical normal situation.

[0086] In a preferred embodiment of the present application, the appearance feature similarity process includes: extracting the point cloud size of the radar-monitored target in the potential target pair.

[0087] The point cloud size is compared with the pre-stored reasonable vehicle size interval of the category to which the video-monitored target in the potential target pair belongs.

[0088] If the point cloud size is within the interval, it is determined that the appearance feature similarity verification is qualified.

[0089] It should be noted that the core logic of the appearance feature similarity verification process is based on cross-modal physical consistency verification, which fully utilizes the precise measurement capability of radar and the semantic recognition capability of video, and provides an important and reliable constraint condition for multi-target association through size-class rationality checking. It is a verification method based on objective laws of the physical world, and is also a key technical means to ensure the accuracy of association in multi-sensor analysis.

[0090] In the embodiment of the application, whether the association constraint condition is met is judged based on spatial position consistency, motion trajectory continuity and appearance feature similarity in the association matching process of radar monitored targets and video monitored targets, and an effective associated target pair is output. On the basis of space-time alignment, a judgment basis for perceiving scene context is provided, which greatly enhances the accuracy of association decision.

[0091] Referring to Figure 3 In a preferred embodiment of the application, the target association result is adaptively output according to the actual monitoring scene, which includes: dividing the highway monitoring area into a near-field zone, a middle-field zone and a far-field zone according to the distance of the target from the radar sensor.

[0092] For example, the near-field zone can be defined as 0-50 meters from the radar, the middle-field zone can be defined as 50-150 meters, and the far-field zone can be defined as 150-500 meters.

[0093] For different field zones, the weights of radar and video in fusion decision are dynamically adjusted in combination with the video category label confidence.

[0094] It should be noted that the weight of video data in the near-field zone is greater than the weight of radar data, and for example, the weight of video data is selected as 0.8 and the weight of radar data is selected as 0.2.

[0095] The weight of radar and video data fusion in the middle-field zone is dynamically calculated according to the video category label confidence.

[0096] In the far-field zone, the weight of radar data is greater than the weight of video data, and for example, the weight of video data is selected as 0.2 and the weight of radar data is selected as 0.8.

[0097] The dynamic calculation process of the above-mentioned middle-field zone is: to establish a basic boundary range of the weight of video data, for example , the weight of radar data is set in the corresponding reverse interval, high and low confidence thresholds are set as the decision boundary of weight distribution, and a piecewise linear interpolation function is used to obtain the smooth distribution weight under the decision attribute corresponding to the video category label confidence, wherein the video category label confidence is obtained by performing double mean calculation on the confidence of each video monitoring target in each frame image in the video image sequence.

[0098] Fuse the radar point cloud data and the video image data based on the weight to obtain a target association result.

[0099] The embodiment of the present application effectively solves the problem of dynamic change of device performance with distance by adaptive adjustment of partition weight, focuses on video detail features in the near field area, and relies on radar ranging accuracy in the far field area, so as to realize optimal configuration of sensing accuracy in the whole monitoring area, and significantly improve the data fusion reliability.

[0100] S14. Structure the target association result, generate a fusion target data packet, and output the fusion target data packet for monitoring feedback.

[0101] The above is only an example and description of the concept of the present application. Those skilled in the art can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, as long as they do not deviate from the concept of the present application or exceed the scope defined by the present application, which shall fall within the protection scope of the present application.

Claims

1. A method for highway monitoring by radar and video fusion dynamic spatio-temporal alignment, characterized in that, The method comprises the following steps: Real-time acquisition of radar point cloud data and video image data in a highway monitoring area, and independent retrieval of radar monitoring targets and video monitoring targets in the data; Temporal and spatial alignment processing of the radar monitoring targets and the video monitoring targets, including dynamic spatial alignment and multi-scale time synchronization; Correlation matching of the radar monitoring targets and the video monitoring targets after temporal and spatial alignment, judgment of whether the correlation constraint conditions are met based on spatial position consistency, motion trajectory continuity and appearance feature similarity, and adaptive output of target correlation results according to actual monitoring scenes; Structural encapsulation of the target correlation results, generation of fusion target data packets and output for monitoring feedback.

2. The highway monitoring method of claim 1, wherein: The radar monitoring target retrieval process is implemented as follows: Pretreatment of radar point cloud data, including filtering out noise point clouds and background point clouds; Segmentation of the pretreated point clouds by using a clustering algorithm to distinguish independent radar monitoring targets, and temporary identifiers are assigned to each radar monitoring target; For each clustering result, the three-dimensional spatial centroid coordinates in the radar coordinate system are calculated, and the kinematic parameters are analyzed based on the Doppler effect, including radial velocity, azimuth angle and pitch angle.

3. The method of claim 2, wherein the method further comprises: The video monitoring target retrieval process is implemented as follows: Processing of video image data by using a target detection algorithm to identify each video monitoring target in the image, and judging the category label and the corresponding confidence degree based on the appearance depth feature; Drawing the boundary box of each video monitoring target and extracting the center point in the box as the two-dimensional pixel coordinates of the video monitoring target in the image coordinate system; Motion trajectory tracking of each video monitoring target in the sequence image, and analysis of the behavior semantics of each video monitoring target based on the trajectory data.

4. The method of claim 1, wherein the method further comprises: The dynamic spatial alignment is implemented as follows: Initial spatial calibration of the radar sensor and the camera device, establishment of an initial spatial conversion matrix, and mapping of the three-dimensional spatial coordinates in the radar coordinate system to the two-dimensional pixel coordinates in the image coordinate system by using the initial spatial conversion matrix; Position matching of the radar monitoring target to the video monitoring target by using the initial spatial conversion matrix, and construction of a dynamic alignment error function taking the projection position deviation of the two targets on the image plane as the optimization direction; Real-time adjustment of the parameters of the initial spatial conversion matrix based on the dynamic spatial alignment error function by using gradient optimization to complete the dynamic spatial alignment of the radar monitoring target and the video monitoring target.

5. The method of claim 1, wherein the method further comprises: The multi-scale time synchronization is implemented as follows: Based on a unified clock source, time stamps are added to the radar point cloud data, frame numbers and corresponding time information are added to the video image data, a time sequence correspondence relationship between the time stamps and the frame numbers is established; According to the time sequence correspondence relationship, the radar data is resampled at the video frame sampling time by linear interpolation to generate a radar data sequence synchronized with the video frame; Nonlinear time sequence matching and alignment of the resampled radar data sequence and the video image sequence, the nonlinear time sequence matching is realized by calculating the minimum cumulative alignment cost between the two sequences.

6. The method of claim 3, wherein the method further comprises: The correlation matching of the radar monitoring targets and the video monitoring targets after temporal and spatial alignment is implemented as follows: The generated synchronization time stamp and space mapping relationship are processed based on time and space alignment, and a monitoring target whose Euclidean distance between the center point coordinates of the radar monitoring target projected onto the image plane and the video monitoring target at the same time is less than a preset threshold is searched to form a potential target pair; The spatial position consistency verification, motion trajectory continuity verification and appearance feature similarity verification are sequentially performed on the potential target pair; If all verifications are qualified, the effective associated target pair is determined; If the target density in the perceived highway monitoring area reaches a preset high-density standard, a trajectory prediction assisted alignment mechanism is started to predict the trajectory position of each video monitoring target at the next time, and the potential target pair whose verification fails due to target occlusion or trajectory intersection is re-associated and matched based on the predicted trajectory.

7. The method of claim 6, wherein the method further comprises: The spatial position consistency verification process includes: projecting the three-dimensional coordinate sequence of the radar monitoring target in the potential target pair onto the image plane to generate a predicted bounding box; measuring the intersection over union of the predicted bounding box and the video monitoring target bounding box in the potential target pair in the video image sequence set; dynamically selecting an intersection over union threshold according to the target category and its video pixel size, counting the frame number proportion of the video image sequence whose intersection over union is greater than the threshold, and if the frame number proportion meets the preset percentage requirement, the spatial position consistency verification is qualified.

8. The method of claim 6, wherein the method further comprises: The motion trajectory continuity verification process includes: extracting the behavior semantics of the video monitoring target in the potential target pair, and analyzing the expected performance trend of the kinematic parameters corresponding to the behavior semantics; quantifying the kinematic parameter performance trend of the radar monitoring target in the potential target pair in the corresponding video image sequence period, and calculating the matching degree with the expected performance trend; if the matching degree meets the high standard, the motion trajectory continuity verification is qualified.

9. The method of claim 6, wherein the method further comprises: The appearance feature similarity process includes: extracting the point cloud size of the radar monitoring target in the potential target pair; comparing the point cloud size with the pre-stored reasonable vehicle size interval of the category to which the video monitoring target in the potential target pair belongs; if the point cloud size is within the interval, the appearance feature similarity verification is qualified.

10. The method of claim 3, wherein the method further comprises: The adaptive output of the target association result according to the actual monitoring scene includes: dividing the highway monitoring area into near-field, middle-field and far-field areas according to the distance between the target and the radar sensor; for different field areas, dynamically adjusting the weight of radar and video in the fusion decision according to the video category label confidence; based on the weight, the radar point cloud data and video image data of the effective associated target pair are fused to obtain the target association result.

Citation Information

Patent Citations

  • Vehicle holographic perception and risk behavior identification system based on thunder multi-source data deep fusion

    CN116403179A

  • Highway event recognition method and recognition device based on multi-source traffic data

    CN117649632B

  • Expressway multi-target tracking method based on Leiyu fusion

    CN116863382A

  • Method for tracking multiple objects

    EP4439478A1

  • Lane detection method and system based on vision and lidar multi-level fusion

    US10929694B1

Cited By

  • Scene label generation method and device based on dynamic graph reasoning

    CN121564718A