Heterogeneous asynchronous multi-source data fusion driving track integration method and system

Through the heterogeneous asynchronous multi-source data fusion method, combined with multiple sensor data, a fusion architecture of space-time alignment is established, which solves the problem of relying on a single sensor to obtain two-dimensional information in traditional traffic monitoring management, and achieves complete trajectory data integration and cloud upload of traffic goals, providing precise support for traffic management.

CN120108197AInactive Publication Date: 2025-06-06SHENZHEN URBAN TRANSPORT PLANNING CENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510599576.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional traffic monitoring management relies on a single sensor, can only obtain two-dimensional information, and cannot analyze the motion state and itinerary data of traffic targets, resulting in poor trajectory integration effect.

Method used

Using heterogeneous asynchronous multi-source data fusion method, by receiving and preprocessing multiple sensor data (camera video and radar point cloud data), the traffic target attributes are extracted and the master-slave clock synchronization architecture is constructed, and a local BEV coordinate system based on UTM projection is established to realize the spatial and temporal alignment and fusion of the data, and finally integrate it into complete trajectory data and upload it to the cloud.

Benefits of technology

It realizes efficient fusion of cross-sensor data, generates a relatively complete itinerary trajectory of traffic goals, and provides precise support for traffic management and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108197A_ABST
    Figure CN120108197A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous asynchronous multi-source data fusion driving track integration method and system, relates to the field of image sensing, and solves the problem that the track integration effect is limited due to the fact that only two-dimensional plane information can be obtained by depending on a single sensor in the prior art. Comprising the following steps: S1, receiving input data of a sensor and preprocessing the input data to obtain standardized data; s2, traffic target attributes and unique identifiers are extracted based on the standardized data, and traffic target attribute data with labels are obtained; s3, constructing a master-slave clock synchronization architecture based on a precision time protocol of a GPS time source, establishing a local BEV coordinate system based on UTM projection, and converting the traffic target attribute data with the label into fusion data of time-space alignment; s4, integrating the fusion data subjected to space-time alignment into complete trajectory data; and S5, packaging the complete track data into a cloud message and uploading the cloud message to the cloud. The method has a good application prospect in the field of traffic monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image sensing, and in particular to a driving trajectory integration method and system based on heterogeneous asynchronous multi-point multi-source data perception fusion. Background Art

[0002] Traditional traffic monitoring management usually relies on data from a single sensor (such as a camera), and by analyzing the camera's video stream, statistics are obtained on the flow and density of traffic targets in a road scene within a specific time period. However, this method can only obtain two-dimensional information about traffic targets on the pixel plane, and cannot obtain its speed and position information, making it difficult to analyze the target's motion status and travel data.

[0003] In addition, a variety of sensors are usually installed in road facilities, such as microwave radar, depth camera, camera and GPS, etc. However, traditional traffic monitoring management methods fail to make full use of the data from these sensors, resulting in poor data fusion effects. Summary of the invention

[0004] In order to solve the problem that the existing technology relies on a single sensor to only obtain two-dimensional plane information, resulting in limited trajectory integration effect, the present invention provides a heterogeneous asynchronous multi-source data fusion driving trajectory integration method, comprising the following steps: S1. Receive input data from the sensor and preprocess it to obtain standardized data, wherein the input data includes video data collected by the camera and point cloud data collected by the radar; S2. Extracting traffic target attributes and unique identifiers based on the standardized data to obtain traffic target attribute data with labels; S3. Construct a master-slave clock synchronization architecture based on the precision time protocol of the GPS time source, establish a local BEV coordinate system based on the UTM projection, and convert the labeled traffic target attribute data into spatiotemporally aligned fused data; S4. Integrating the spatiotemporally aligned fusion data into complete trajectory data; S5. Encapsulate the complete trajectory data into a cloud message and upload it to the cloud.

[0005] Further, in S1, the preprocessing of the video data collected by the camera includes the following steps: S11. Connect the camera and the airborne edge computing platform through the internal Ethernet of the drone, and the edge computing platform obtains the real-time video stream information collected by the camera through the RTSP video stream address; S12. Decode the real-time video stream information into a single-frame image in a unified RGB format; S13. Performing color space conversion and image filtering and denoising on the single frame image; The preprocessing of the radar collected point cloud data includes the following steps: S14. Remove ground points from point cloud data using height threshold and RANSAC plane fitting; S15. Remove outliers and noise points of the point cloud data by statistical filtering to retain valid data; S16. Aligning the point cloud data of different frames using a registration algorithm to generate a continuous, seamless point cloud map; S17. Divide the point cloud map into voxel grids, and replace all points in the grid with the centroid of each grid; S18. Randomly select part of the point cloud data.

[0006] Furthermore, S2 specifically includes: S21. Perform 2D target detection on the standardized data using the Yolov5 algorithm, output the target frame coordinates and calculate the center point coordinates, and then track the target using the DeepSort algorithm and assign a unique identifier; The center point coordinates are: ; Obtain, among which, is the horizontal coordinate of the upper left corner of the target detection box, is the horizontal coordinate of the lower right corner of the target detection box, is the ordinate of the upper left corner of the target detection box, is the ordinate of the lower right corner of the target detection box, is the horizontal coordinate of the center point, is the ordinate of the center point; S22. The binocular camera obtains the depth information of the target through parallax calculation, maps it to the 3D space, and forms a 3D detection result including distance and position attributes; S23. Converting the standardized data to a world coordinate system to eliminate differences in sensor viewing angles; S24. The radar-collected point cloud data is subjected to rasterization processing and CNN network inference to generate the labeled traffic target attribute data.

[0007] Further, in S3, the establishing of a local BEV coordinate system based on UTM projection to convert the labeled traffic target attributes into spatiotemporally aligned fused data specifically includes: S31. Taking the center point of the road intersection as the origin, convert the GPS coordinates into a local coordinate system based on the UTM projection method, set the north direction as the Y axis, the east direction as the X axis, and the rotation matrix as the unit matrix; S32. Mapping the labeled traffic target attribute data to a local coordinate system through a coordinate transformation matrix; S33. Identify the overlapping area of ​​the visible light target frame and the radar grid in the local coordinate system, and calculate the spatial overlap rate between the two; S34. Prioritize the comparison of the category labels of the visible light and radar detection targets, and only targets of the same type enter the next step of matching; S35. Coincidence rate threshold determination: If the coincidence rate is ≥50%, it is determined that the two types of detection results point to the same physical target, triggering the data association process; S36. Hide the redundant radar target frame, bind the radar speed, distance and other attributes with the unique identifier of the visible light target, and form the time-space aligned fusion data.

[0008] Further, the S5 specifically includes: S51. Integrate the complete trajectory data into a structured message, and use the unique identifier as an identifier; S52. The complete trajectory data of the target in each area is associated with the unique identifier and saved, and when the unique identifier target message information is output, it is combined and sent to the cloud platform.

[0009] A heterogeneous asynchronous multi-source data fusion driving trajectory integration system is also provided, comprising: a data acquisition and processing module, a target detection module, a data fusion module, a target matching re-identification module and a trip integration transmission module; The data acquisition and processing module is used to receive input data from the sensor and perform preprocessing to obtain standardized data, wherein the input data includes video data collected by the camera and point cloud data collected by the radar; The target detection module is used to extract traffic target attributes and unique identifiers based on the standardized data to obtain traffic target attribute data with labels; The data fusion module is used to construct a master-slave clock synchronization architecture based on the precision time protocol of the GPS time source, establish a local BEV coordinate system based on the UTM projection, and convert the labeled traffic target attribute data into space-time aligned fusion data; The target matching re-identification module is used to integrate the spatiotemporally aligned fusion data into complete trajectory data; The trip integration transmission module is used to encapsulate the complete trajectory data into a cloud message and upload it to the cloud.

[0010] The beneficial effects of the present invention are: The present invention establishes a fusion coordinate system, maps the data detected by each sensor to the fusion coordinate system, and realizes cross-sensor data fusion.

[0011] The present invention integrates and associates multiple segments of fused data in a fused coordinate system by re-identifying target matching, thereby generating a relatively complete travel trajectory for a single traffic target, providing accurate support for traffic management and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flow chart of a heterogeneous asynchronous multi-source data fusion driving trajectory integration method; Figure 2 This is a structural diagram of a heterogeneous asynchronous multi-source data fusion driving trajectory integration system; Figure 3 It is a schematic diagram of multi-source data projection in the local BEV coordinate system of the intersection; Figure 4 Schematic diagram of cross-region sensor collaborative target matching and trajectory integration; Figure 5 Flowchart for cross-region sensor collaborative target matching and trajectory integration. DETAILED DESCRIPTION

[0013] In order to make the technical solutions and advantages of the embodiments of the present invention more clearly understood, the exemplary embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than an exhaustive list of all the embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0014] Embodiment 1, combination Figure 1 This embodiment describes a method for integrating driving trajectories by fusion of heterogeneous asynchronous multi-source data, including the following steps: S1. Receive input data from the sensor and preprocess it to obtain standardized data, wherein the input data includes video data collected by the camera and point cloud data collected by the radar; S2. Extracting traffic target attributes and unique identifiers based on the standardized data to obtain traffic target attribute data with labels; S3. Construct a master-slave clock synchronization architecture based on the precision time protocol of the GPS time source, establish a local BEV coordinate system based on the UTM projection, and convert the labeled traffic target attribute data into spatiotemporally aligned fused data; S4. Integrating the spatiotemporally aligned fusion data into complete trajectory data; S5. Encapsulate the complete trajectory data into a cloud message and upload it to the cloud.

[0015] Specifically, Figure 2As shown in Figure 4, in S4, the sensors traditionally mounted on the crossbar can usually only collect data in a specific direction, such as area A or area C. However, the area directly below the crossbar belongs to the blind spot of the two areas. When the traffic target travels between area A and area C, the use of traditional target re-identification algorithms may result in low recognition accuracy and high misrecognition rate.

[0016] A camera is installed directly below the crossbar to collect data in area B. Area B partially overlaps with areas A and C. After the crossbar is installed, the equipment position is usually fixed. The 77GHz or 79GHz microwave radar equipped has a coverage range of 10 to 250 meters. Therefore, the length of area A is 10 to 250 meters from the crossbar and the width is the total width of the lane. The range of area C is similar. The monitoring range of the camera in area B is usually 50 meters. It can be inferred that the overlapping range of area A and area B is 10 to 25 meters from the crossbar, and the overlapping range of area C is the same as that of area B.

[0017] When the traffic target travels from the starting point O of area A to the end point Q of area C, its distance from the sensor in area A gradually decreases, and its driving direction can be determined. When the target enters the overlapping area P of area A and area B, the system adds the target feature information collected in area A to the queue to be matched, and re-identifies and matches the target in the data collected in area B.

[0018] Prioritize matching the target's license plate features. When the license plate number cannot be completely matched, the system refers to the analysis and matching results of the convolutional neural network on the target features. If the license plate number matching degree exceeds 80%, the matching result of the convolutional neural network is accepted and deemed to be a successful match. After a successful match, the target data collected in each area is bound to the same UID. When the target enters the overlapping area of ​​area B and area C, the system re-identifies and matches the target features in the data collected in area C. After a successful match, the target data of multiple areas are integrated and bound to the UID to generate the complete trajectory data.

[0019] When the traffic target drives from the starting point Q of area C to the starting point O of area A, the recognition and matching process is the same as above.

[0020] S4 specific process is as follows Figure 3 shown.

[0021] In S1, the preprocessing of the video data collected by the camera includes the following steps: S11. Connect the camera and the airborne edge computing platform through the internal Ethernet of the drone, and the edge computing platform obtains the real-time video stream information collected by the camera through the RTSP video stream address; S12. Decode the real-time video stream information into a single-frame image in a unified RGB format; S13. Performing color space conversion and image filtering and denoising on the single frame image; The preprocessing of the radar collected point cloud data includes the following steps: S14. Remove ground points from point cloud data using height threshold and RANSAC plane fitting; S15. Remove outliers and noise points of the point cloud data by statistical filtering to retain valid data; S16. Aligning the point cloud data of different frames using a registration algorithm to generate a continuous, seamless point cloud map; S17. Divide the point cloud map into voxel grids, and replace all points in the grid with the centroid of each grid; S18. Randomly select part of the point cloud data.

[0022] S2 specifically includes: S21. Perform 2D target detection on the standardized data using the Yolov5 algorithm, output the target frame coordinates and calculate the center point coordinates, and then track the target using the DeepSort algorithm and assign a unique identifier; The center point coordinates are: ; Obtain, among which, is the horizontal coordinate of the upper left corner of the target detection box, is the horizontal coordinate of the lower right corner of the target detection box, is the ordinate of the upper left corner of the target detection box, is the ordinate of the lower right corner of the target detection box, is the horizontal coordinate of the center point, is the ordinate of the center point; S22. The binocular camera obtains the depth information of the target through parallax calculation, maps it to the 3D space, and forms a 3D detection result including distance and position attributes; S23. Converting the standardized data to a world coordinate system to eliminate differences in sensor viewing angles; S24. The radar-collected point cloud data is subjected to rasterization processing and CNN network inference to generate the labeled traffic target attribute data.

[0023] Specifically, the depth information is obtained by: ; Obtain, among which, For depth information, is the camera focal length, is the baseline distance between cameras, Pixel The parallax value of .

[0024] Map the depth information to 3D space by: ; Implementation, where ( , ) is the camera principal point coordinate, ( , , ) is the coordinate of the target point in the camera coordinate system.

[0025] The camera's extrinsic matrix Describes the position and posture of the camera in the world coordinate system. The external parameter matrix is: ; in, is the rotation matrix, is the translation vector.

[0026] The standardized data is uniformly converted to the world coordinate system by: ; Implementation, where the coordinates of the target point in the world coordinate system are ( , , ), switching to a bird's-eye view usually sets the Z-axis coordinate value to 0.

[0027] The coordinate point of the radar center point in the world coordinate system is ( , , ).

[0028] The conversion formula from the radar coordinate system to the world coordinate system is as follows: ; According to the position and posture state of the radar device, the transformation matrix can be solved. is the rotation matrix, is the translation vector. If the radar is installed horizontally, and z Axis points upward, rotation matrix Can be the identity matrix I .

[0029] Project the point cloud data collected by the radar to the BEV perspective in the world coordinate system.

[0030] Discretize the projected radar point cloud data onto a grid map with a fixed resolution. Assume that the resolution of the grid map is and , you can convert continuous coordinates to discrete raster indices: ; ; Based on the discrete grid radar point cloud data, the BEV point cloud data x , y The coordinate value and reflection intensity are used as three channels to generate the BEV feature map, and the CNN network is used to perform target detection inference on the BEV feature map.

[0031] In S3, the establishment of a local BEV coordinate system based on UTM projection to convert the labeled traffic target attributes into spatiotemporally aligned fused data specifically includes: S31. Taking the center point of the road intersection as the origin, convert the GPS coordinates into a local coordinate system based on the UTM projection method, set the north direction as the Y axis, the east direction as the X axis, and the rotation matrix as the unit matrix; S32. Mapping the labeled traffic target attribute data to a local coordinate system through a coordinate transformation matrix; S33. Identify the overlapping area of ​​the visible light target frame and the radar grid in the local coordinate system, and calculate the spatial overlap rate between the two; S34. Prioritize the comparison of the category labels of the visible light and radar detection targets, and only targets of the same type enter the next step of matching; S35. Coincidence rate threshold determination: If the coincidence rate is ≥50%, it is determined that the two types of detection results point to the same physical target, triggering the data association process; S36. Hide the redundant radar target frame, bind the radar speed, distance and other attributes with the unique identifier of the visible light target, and form the time-space aligned fusion data.

[0032] Specifically, through Figure 4 It can be seen that S3 selects a reference point at the center of the road intersection, and obtains its coordinate point in world coordinates based on the GPS information of the point and using the UTM projection method. ( , , ) The local BEV coordinate system is established with the reference point as the center point, the north direction as the positive direction of the Y axis, and the east direction as the positive direction of the X axis.

[0033] The transformation matrix from the world coordinate system to the local coordinate system of the intersection is: . T is the translation vector, which can be approximately equal to the opposite of the coordinate value of the reference point in the world coordinate system. The angle between the local coordinate system established with the north-south direction as the Y axis and the east-west direction as the X axis and the world coordinate system is 0°. Finally, the binocular camera target detection inference results from the BEV perspective in the world coordinate system and the radar point cloud rasterization inference results are projected to the local coordinate system of the intersection by the above transformation matrix.

[0034] The S5 specifically includes: S51. Integrate the complete trajectory data into a structured message, and use the unique identifier as an identifier; S52. The complete trajectory data of the target in each area is associated with the unique identifier and saved, and when the unique identifier target message information is output, it is combined and sent to the cloud platform.

[0035] Here is an example to illustrate the structured message format: the message information is as follows: { "id": "xxxx", "uuid": "0kz5jlizotnlov2amct7fd4jl3s4kqza", "route1": { "type": "1", "GPS": "XXXXXXX", "speed": "20210313155853", "time": "20210313155853", "VideoName":"01_20210313155853.avi" } } Field Description:

[0036] The video naming convention is: driving direction_time.avi For example, the driving behavior close to the installation crossbar at 15:58:53 on March 13, 2021: The picture is named: 01_20210313155853.jpg The message data and screenshots are sent to the cloud through the post function of https. Here, the edge computing gateway is used as the https client to send data, and the cloud is used as the https server to receive and store information. If the edge computing gateway fails to send, a cache mechanism will be used to save the failed data locally and resend it at another time.

[0037] A heterogeneous asynchronous multi-source data fusion driving trajectory integration system, combined with Figure 5 Description, including: data acquisition and processing module, target detection module, data fusion module, target matching re-identification module and itinerary integration transmission module; The data acquisition and processing module is used to receive input data from the sensor and perform preprocessing to obtain standardized data, wherein the input data includes video data collected by the camera and point cloud data collected by the radar; The target detection module is used to extract traffic target attributes and unique identifiers based on the standardized data to obtain traffic target attribute data with labels; The data fusion module is used to construct a master-slave clock synchronization architecture based on the precision time protocol of the GPS time source, establish a local BEV coordinate system based on the UTM projection, and convert the labeled traffic target attribute data into space-time aligned fusion data; The target matching re-identification module is used to integrate the spatiotemporally aligned fusion data into complete trajectory data; The trip integration transmission module is used to encapsulate the complete trajectory data into a cloud message and upload it to the cloud.

Claims

1. A heterogeneous asynchronous multi-source data fusion driving trajectory integration method, characterized in that: The following steps are involved: S1. Receive input data from the sensor and preprocess it to obtain standardized data, wherein the input data includes video data collected by the camera and point cloud data collected by the radar; S2. Extracting traffic target attributes and unique identifiers based on the standardized data to obtain traffic target attribute data with labels; S3. Construct a master-slave clock synchronization architecture based on the precision time protocol of the GPS time source, establish a local BEV coordinate system based on the UTM projection, and convert the labeled traffic target attribute data into spatiotemporally aligned fused data; S4. Integrating the spatiotemporally aligned fusion data into complete trajectory data; S5. Encapsulate the complete trajectory data into a cloud message and upload it to the cloud.

2. The method for integrating driving trajectories by fusion of heterogeneous asynchronous multi-source data according to claim 1 is characterized in that: In S1, the preprocessing of the video data collected by the camera includes the following steps: S11. Connect the camera and the airborne edge computing platform through the internal Ethernet of the drone, and the edge computing platform obtains the real-time video stream information collected by the camera through the RTSP video stream address; S12. Decode the real-time video stream information into a single-frame image in a unified RGB format; S13. Performing color space conversion and image filtering and denoising on the single frame image; The preprocessing of the radar collected point cloud data includes the following steps: S14. Remove ground points from point cloud data using height threshold and RANSAC plane fitting; S15. Remove outliers and noise points of the point cloud data by statistical filtering to retain valid data; S16. Aligning the point cloud data of different frames using a registration algorithm to generate a continuous, seamless point cloud map; S17. Divide the point cloud map into voxel grids, and replace all points in the grid with the centroid of each grid; S18. Randomly select part of the point cloud data.

3. The method for integrating driving trajectories by fusion of heterogeneous asynchronous multi-source data according to claim 1 is characterized in that: S2 specifically includes: S21. Perform 2D target detection on the standardized data using the Yolov5 algorithm, output the target frame coordinates and calculate the center point coordinates, and then track the target using the DeepSort algorithm and assign a unique identifier; The center point coordinates are: ; Obtain, among which, is the horizontal coordinate of the upper left corner of the target detection box, is the horizontal coordinate of the lower right corner of the target detection box, is the ordinate of the upper left corner of the target detection box, is the ordinate of the lower right corner of the target detection box, is the horizontal coordinate of the center point, is the ordinate of the center point; S22. The binocular camera obtains the depth information of the target through parallax calculation, maps it to the 3D space, and forms a 3D detection result including distance and position attributes; S23. Converting the standardized data to a world coordinate system to eliminate differences in sensor viewing angles; S24. The radar-collected point cloud data is subjected to rasterization processing and CNN network inference to generate the labeled traffic target attribute data.

4. The method for integrating driving trajectories by fusion of heterogeneous asynchronous multi-source data according to claim 1 is characterized in that: In S3, the establishment of a local BEV coordinate system based on UTM projection to convert the labeled traffic target attributes into spatiotemporally aligned fused data specifically includes: S31. Taking the center point of the road intersection as the origin, convert the GPS coordinates into a local coordinate system based on the UTM projection method, set the north direction as the Y axis, the east direction as the X axis, and the rotation matrix as the unit matrix; S32. Mapping the labeled traffic target attribute data to a local coordinate system through a coordinate transformation matrix; S33. Identify the overlapping area of ​​the visible light target frame and the radar grid in the local coordinate system, and calculate the spatial overlap rate between the two; S34. Prioritize the comparison of the category labels of the visible light and radar detection targets, and only targets of the same type enter the next step of matching; S35. Coincidence rate threshold determination: If the coincidence rate is ≥50%, it is determined that the two types of detection results point to the same physical target, triggering the data association process; S36. Hide the redundant radar target frame, bind the radar speed, distance and other attributes with the unique identifier of the visible light target, and form the time-space aligned fusion data.

5. The method for integrating driving trajectories by fusion of heterogeneous asynchronous multi-source data according to claim 1 is characterized in that: The S5 specifically includes: S51. Integrate the complete trajectory data into a structured message, and use the unique identifier as an identifier; S52. The complete trajectory data of the target in each area is associated with the unique identifier and saved, and when the unique identifier target message information is output, it is combined and sent to the cloud platform.

6. A heterogeneous asynchronous multi-source data fusion driving trajectory integration system, characterized in that: include: Data acquisition and processing module, target detection module, data fusion module, target matching and re-identification module and itinerary integration and transmission module; The data acquisition and processing module is used to receive input data from the sensor and perform preprocessing to obtain standardized data, wherein the input data includes video data collected by the camera and point cloud data collected by the radar; The target detection module is used to extract traffic target attributes and unique identifiers based on the standardized data to obtain traffic target attribute data with labels; The data fusion module is used to construct a master-slave clock synchronization architecture based on the precision time protocol of the GPS time source, establish a local BEV coordinate system based on the UTM projection, and convert the labeled traffic target attribute data into space-time aligned fusion data; The target matching re-identification module is used to integrate the spatiotemporally aligned fusion data into complete trajectory data; The trip integration transmission module is used to encapsulate the complete trajectory data into a cloud message and upload it to the cloud.

Citation Information

Patent Citations

  • Three-dimensional multi-target tracking method fusing images and laser point clouds

    CN110675431A

  • Unmanned aerial vehicle multi-target tracking method based on visual / millimeter wave radar information fusion

    CN115731268A

  • Multi-sensor data fusion method, device, equipment and medium

    CN119227011A