Traffic operation state collaborative perception method, storage medium and electronic equipment

By using cameras and LiDAR for collaborative perception and leveraging YOLOv8 and PointPillars models for multi-sensor data fusion, the limitations of single sensors in complex environments are addressed, enabling high-precision perception of traffic operation status.

CN120953932APending Publication Date: 2025-11-14JINLING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511010009.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing traffic operation status perception technologies, the differences in the sensing characteristics of single-type sensors lead to poor performance in complex environments, making it difficult to achieve effective fusion of different sensors.

Method used

Data is collected by cameras and LiDAR respectively. Vehicle detection bounding box information is extracted through YOLOv8 and PointPillars models. Time synchronization filtering and multi-coordinate system transformation are performed. Combined with decision-level fusion method and BoT-SORT algorithm, collaborative perception of multiple sensors is achieved.

Benefits of technology

It improves the accuracy and stability of traffic operation status perception, is suitable for robust perception in complex environments, and enhances the environmental perception capability of the traffic system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953932A_ABST
    Figure CN120953932A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic operation state collaborative perception method, a storage medium and electronic equipment, and the method comprises the steps: collecting image data and point cloud data through a camera and a laser radar, extracting the vertex information and image time information of a 2D vehicle detection frame from the image data through a YOLOv8 model, extracting vertex information and point cloud time information of the 3D vehicle detection frame from the point cloud data by adopting a Point Pillars model; performing time synchronization on the acquired information, acquiring corresponding 2D vehicle detection frame vertex information and 3D vehicle detection frame vertex information under the condition that the image moment information and the point cloud moment information are the same or similar, and converting the 2D vehicle detection frame vertex information and the 3D vehicle detection frame vertex information into the same coordinate form; a decision-level fusion method is adopted, a detection frame IoU is utilized, two projection frames formed by the converted information are matched, and a result is output; and a BoT-SORT algorithm is adopted to receive the result, and the traffic operation state of the target vehicle is output. According to the invention, the problem of time-space asynchronization of multiple sensors is solved, and the traffic operation state sensing effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle traffic operation status perception technology, specifically to a collaborative perception method for traffic operation status, a storage medium, and an electronic device. Background Technology

[0002] Existing traffic status perception technologies primarily rely on single types of sensors, such as infrastructure cameras or vehicle-mounted LiDAR. Due to the differences in the sensing characteristics of different sensors, existing solutions have significant shortcomings in practical applications. For example, cameras are greatly affected by lighting and weather conditions, easily resulting in missed or false detections at night, in foggy weather, or in complex traffic scenarios, and they struggle to acquire 3D information and motion characteristics of targets. While LiDAR possesses strong 3D perception capabilities, it suffers from sparse point clouds, poor detection of small targets at long distances, and an inability to acquire image texture information. Furthermore, existing technologies struggle to effectively fuse different sensors, leading to inadequate traffic status perception.

[0003] Therefore, there is an urgent need for a collaborative perception method, storage medium, and electronic device for traffic operation status to solve the problem that different sensors are difficult to integrate effectively, resulting in poor perception of traffic operation status. Summary of the Invention

[0004] This invention addresses the shortcomings of existing technologies by providing a collaborative perception method, storage medium, and electronic device for traffic operation status, thereby solving the problem of poor traffic operation status perception due to the difficulty in effectively integrating different sensors.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] A collaborative perception method for traffic operation status includes the following steps:

[0007] Image data and point cloud data are acquired using cameras and LiDAR, respectively. A YOLOv8 model is used to extract 2D vehicle detection box vertex information and image time information from the image data, while a PointPillars model is used to extract 3D vehicle detection box vertex information and point cloud time information from the point cloud data. The acquired information is time-synchronized and filtered to obtain the corresponding 2D and 3D vehicle detection box vertex information when the image time information and point cloud time information are the same or similar, and then converted to the same coordinate form. A decision-level fusion method is used, utilizing the detection box IoU, to match the two projected boxes formed by the converted 2D and 3D vehicle detection box vertex information, and the fused detection result is output based on the matching result. Finally, the BoT-SORT algorithm is used to receive the fused detection result and output the traffic operation status of the target vehicle.

[0008] To optimize the above technical solution, the specific measures also include:

[0009] Furthermore, the step of performing time-synchronized filtering on the collected information to obtain the corresponding 2D vehicle detection box vertex information and 3D vehicle detection box vertex information when the image time information and point cloud time information are the same or similar includes the following steps:

[0010] Based on the point cloud time information, the time closest to the corresponding point cloud time information timestamp is selected from the image time information for matching, and the corresponding 2D vehicle detection box vertex information and 3D vehicle detection box vertex information are obtained for each point cloud time information and its matching image time information.

[0011] Furthermore, the conversion to the same coordinate form includes the following steps:

[0012] Establish camera coordinate system, lidar coordinate system, world coordinate system, image coordinate system and pixel coordinate system. After transforming the vertex information of 2D vehicle detection box and 3D vehicle detection box into multiple coordinate systems, output the pixel coordinate form.

[0013] Furthermore, establishing the camera coordinate system, lidar coordinate system, world coordinate system, image coordinate system, and pixel coordinate system includes the following steps:

[0014] Establish camera coordinate system O C -X C Y C Z C LiDAR coordinate system O L -X L Y L Z L World coordinate system O W -X W Y W Z W Image coordinate system O I -X I Y I and pixel coordinate system O P -uv; in the camera coordinate system O C -X C Y C Z C In the middle, the optical center of the camera installed on the traffic light pole is taken as the origin O. C coordinate axis X C Y C Z C Direction X C The positive direction is the horizontal rightward direction of the image, Y C The positive direction is the bottom right of the image horizontally, Z. CThe positive direction is perpendicular to the front of the camera; in the lidar coordinate system O L -X L Y L Z L In the above, the laser emission point of the lidar sensor is taken as the origin O. L coordinate axis X L Y L Z L Direction X L The positive direction points directly ahead on the road, Y L The positive direction is perpendicular to the left of the lidar, Z L The positive direction is perpendicular to the ground and points to the sky; in the world coordinate system O W -X W Y W Z W , take the origin O W The location is the southwest corner of the world map, with the X-axis as the coordinate axis. W Y W Z W Direction X W The positive direction is due east on the map, Y W The positive direction is due north on the map, Z W The positive direction is perpendicular to the sky above the map; in the image coordinate system O I -X I Y I In the process, the center of the imaging plane is taken as the origin O. I coordinate axis X I Y I Direction X I The positive direction is horizontal to the right of the image plane, Y I The positive direction is downward from the image plane; in the pixel coordinate system O P -u a v a In the image, the top-left pixel is taken as the origin O. P coordinate axis u a v a In the direction of u a The positive direction is the pixel index to the right horizontally, v a The positive direction is the vertical downward pixel index.

[0015] Furthermore, the step of transforming the 2D vehicle detection box vertex information and the 3D vehicle detection box vertex information into pixel coordinates after multiple coordinate system transformations includes the following steps:

[0016] The coordinate information of the 3D vehicle detection box vertices in the LiDAR coordinate system is converted into coordinate information in the world coordinate system. Then, the coordinate information in the world coordinate system is converted into coordinate information in the camera coordinate system. Next, the coordinate information in the camera coordinate system, including the coordinate information of the 2D vehicle detection box vertices and the converted 3D vehicle detection box vertices, is converted into coordinate information in the image coordinate system. Finally, the coordinate information in the image coordinate system is converted into coordinate information in the pixel coordinate system and output.

[0017] Furthermore, it includes the following steps:

[0018] The coordinate information of the vertices of the 3D vehicle detection box is represented by homogeneous coordinates in the lidar coordinate system as P. L =x L ,y L ,z L ,1] T The extrinsic matrix needed to transform it to the world coordinate system is as follows:

[0019]

[0020] Wherein, the rotation matrix R from the lidar coordinate system to the world coordinate system WL For a 3×3 matrix, the translation matrix t WL It is a 3×1 vector;

[0021] Coordinates of points in the transformed world coordinate system:

[0022] P W =T WL ·P L

[0023] The coordinates obtained in the previous step are transformed from the world coordinate system to the camera coordinate system for projection onto the planar image coordinate system. The inverse matrix of the extrinsic parameter matrix required is as follows:

[0024]

[0025] Among them, the rotation matrix R for transforming the world coordinate system to the camera coordinate system CW For a 3×3 matrix, the translation matrix t CW It is a 3×1 vector;

[0026] Obtain the coordinates of the camera in the camera coordinate system:

[0027] P C =T CW ·P W

[0028] After completing the spatial transformation, the coordinates P obtained in the previous step are... C =[xC ,y C ,z C ] T By projecting the image onto the image coordinate system using the pinhole camera imaging principle, and normalizing the points in the camera coordinate system, the projection onto the image coordinate system can be obtained. The coordinates (x, y) of the points in the camera coordinate system are then determined. C ,y C ,z C The projection of ) onto the image coordinate system is (x I ,y I The formula is as follows:

[0029]

[0030] Or in homogeneous coordinate form, as follows:

[0031]

[0032] Where f is the focal length of the camera, and the image coordinate system is transformed to the pixel coordinate system, with the origin O of the image coordinate system being O. I If the position in the pixel coordinate system is (u0, v0), then the coordinate transformation relationship is as follows:

[0033]

[0034] Where dx and dy represent the x-axis of a point in the pixel coordinate system and the x-axis of the point in the image coordinate system, respectively. I axis and Y I The unit length value of the axis, written in homogeneous coordinate form, is as follows:

[0035]

[0036] Finally, the coordinates are output in pixel coordinate system.

[0037] Furthermore, the decision-level fusion method, utilizing the IoU of the detection boxes, matches the two projected boxes formed by the vertex information of the transformed 2D vehicle detection boxes and the vertex information of the 3D vehicle detection boxes, and outputs the fused detection result based on the matching result, including the following steps:

[0038] Project the vertex information of the 2D vehicle detection bounding box onto the 2D projection bounding box CBB in the pixel coordinate system. 2D Represented by the coordinates of the endpoints of the two diagonals:

[0039] CBB = (u1, v1, u2, v2)

[0040] Project the vertex information of the 3D vehicle detection bounding box onto the 2D projection bounding box LBB in the pixel coordinate system. 2D Similarly, the coordinates of the endpoints of the two diagonals are expressed as follows:

[0041] LBB = (u3, v3, u4, v4)

[0042] For CBBs with the same or similar time information 2D and LBB 2D The overlapping region S C ∩S L Represented by coordinates:

[0043] S C ∩S L =(u3,v3,u2,v2)

[0044] Among them, S C CBB 2D The area, S L Indicates LBB 2D The area is calculated using the following formula:

[0045]

[0046] According to the definition of IoU, the final formula for calculating the intersection-union ratio is as follows:

[0047]

[0048] After completing the crossover ratio (CBB) calculation, the CBB of the group is determined based on the set IoU threshold. 2D and LBB 2D Determine whether the target matches and output the fusion detection result based on the matching result.

[0049] Furthermore, the step of outputting the fusion detection result based on the matching result includes the following steps: if the calculated intersection-union ratio is greater than a preset threshold, the group of detection results is determined to be a successful match, pointing to the same target; if it is less than the preset threshold, it is determined to be a mismatch, and the next group is selected for the next round of matching.

[0050] Furthermore, a computer-readable storage medium storing a computer program is characterized in that: the computer program causes a computer to execute a traffic operation status collaborative sensing method as described above.

[0051] Furthermore, an electronic device is characterized by comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the aforementioned traffic operation status collaborative perception method.

[0052] The beneficial effects of this invention are:

[0053] This invention addresses the problem of spatiotemporal asynchrony among multiple sensors by introducing multi-coordinate system spatial synchronization procedures and time synchronization strategies based on timestamp nearest neighbor matching, thereby improving the effectiveness of traffic operation status perception. It leverages the advantages of different sensors by combining the YOLOv8 visual detection model and the PointPillars point cloud detection model to achieve high-precision detection of target vehicles. Decision-level fusion, achieved through IoU and other methods, effectively enhances the stability and accuracy of detection. Finally, the BoT-SORT algorithm enables temporal tracking and motion state estimation of the detected target, achieving dynamic and continuous perception of the traffic system's operational status.

[0054] This invention demonstrates robustness in complex environments, significantly outperforming single-sensor perception. It is suitable for infrastructure traffic perception needs in low-density vehicle-to-everything (V2X) environments and shows promising engineering application prospects. This invention overcomes the limitations of single sensors by utilizing multiple sensors, such as cameras and radar, in a collaborative manner, thereby enhancing the environmental perception capabilities of traffic systems. Attached Figure Description

[0055] Figure 1 This is a multi-coordinate system schematic diagram of a collaborative perception method for traffic operation status proposed in this invention;

[0056] Figure 2 This is a time synchronization diagram of a traffic operation status collaborative sensing method proposed in this invention;

[0057] Figure 3 This is a flowchart of the multi-coordinate transformation of a collaborative perception method for traffic operation status proposed in this invention;

[0058] Figure 4 This is a schematic diagram illustrating the principle of camera coordinate system to image coordinate system transformation in a traffic operation status collaborative perception method proposed in this invention.

[0059] Figure 5 This is a schematic diagram illustrating the principle of image coordinate system to pixel coordinate system transformation in a collaborative perception method for traffic operation status proposed in this invention.

[0060] Figure 6 This is a schematic diagram of the projection frame matching for a collaborative perception method for traffic operation status proposed in this invention. Detailed Implementation

[0061] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0062] A traffic operation status collaborative sensing method according to an embodiment of the present invention includes the following steps:

[0063] Image data and point cloud data are acquired using cameras and LiDAR, respectively. A YOLOv8 model is used to extract 2D vehicle detection box vertex information and image time information from the image data, while a PointPillars model is used to extract 3D vehicle detection box vertex information and point cloud time information from the point cloud data. The acquired information is time-synchronized and filtered to obtain the corresponding 2D and 3D vehicle detection box vertex information when the image time information and point cloud time information are the same or similar, and then converted to the same coordinate form. A decision-level fusion method is used, employing the detection box IoU and the Hungarian algorithm to match the two projected boxes formed by the converted 2D and 3D vehicle detection box vertex information, and outputting the fused detection result based on the matching result. Finally, the BoT-SORT algorithm is used to receive the fused detection result, achieve target tracking, and output the traffic operation status of the target vehicle, such as position and speed, realizing collaborative perception of traffic operation status.

[0064] In the above scheme, real traffic flow and vehicle dynamics can be simulated based on the SUMO and CARLA joint simulation platform to construct a virtual-real traffic scenario for testing. Specifically, SUMO is used to generate macro traffic flow data, which includes vehicle paths and traffic light timings, to simulate the traffic flow, congestion distribution, and signal control strategies of the real road network. CARLA is used to deploy CAV vehicles and infrastructure sensors, including cameras and LiDAR. Real-time data interaction between SUMO and CARLA is achieved through the TraCI interface to synchronously acquire traffic flow status and raw sensor data, thus constructing a virtual-real traffic scenario.

[0065] As attached Figure 2 As shown, in a further specific embodiment, the above-mentioned process of performing time-synchronized filtering on the collected information to obtain the corresponding 2D vehicle detection box vertex information and 3D vehicle detection box vertex information when the image time information and point cloud time information are the same or similar includes the following steps:

[0066] Based on the point cloud time information, the time closest to the corresponding point cloud time information timestamp is selected from the image time information for matching, and the corresponding 2D vehicle detection box vertex information and 3D vehicle detection box vertex information are obtained for each point cloud time information and its matching image time information.

[0067] In both simulations and the real world, the sampling frequency of LiDAR is typically lower than that of a camera. While the LiDAR completes one point cloud data acquisition, the camera has already acquired multiple image data. Therefore, this paper uses point cloud time information (i.e., the acquisition time of the point cloud data) as a benchmark and selects the timestamp closest to the acquisition time of the point cloud data from the image time information (i.e., the various acquisition times of the image data) for matching. This achieves time alignment between the two, i.e., extracting and utilizing timestamps that can be used for point cloud-image alignment.

[0068] As attached Figure 1 As shown, in a further specific embodiment, the above conversion to the same coordinate form includes the following steps:

[0069] Establish camera coordinate system, lidar coordinate system, world coordinate system, image coordinate system and pixel coordinate system. After transforming the vertex information of 2D vehicle detection box and 3D vehicle detection box into multiple coordinate systems, output the pixel coordinate form.

[0070] The establishment of the camera coordinate system, lidar coordinate system, world coordinate system, image coordinate system, and pixel coordinate system includes the following steps:

[0071] The camera coordinate system O was established using the SUMO+CARLA co-simulation platform. C -X C Y C Z C LiDAR coordinate system O L -X L Y L Z L World coordinate system O W -X W Y W Z W The image coordinate system O required for camera to acquire and capture images I -X I Y I and pixel coordinate system O P -uv; in the camera coordinate system O C -X C Y C Z C In the middle, the optical center of the camera installed on the traffic light pole is taken as the origin O. C coordinate axis X C Y C Z C Direction X C The positive direction is the horizontal rightward direction of the image, Y C The positive direction is the bottom right of the image horizontally, Z. C The positive direction is perpendicular to the front of the camera; in the lidar coordinate system OL -X L Y L Z L In the above, the laser emission point of the lidar sensor is taken as the origin O. L coordinate axis X L Y L Z L Direction X L The positive direction points directly ahead on the road, Y L The positive direction is perpendicular to the left of the lidar, Z L The positive direction is perpendicular to the ground and points towards the sky; the world coordinate system is an intermediate station established to connect the camera coordinate system and the lidar coordinate system. In the world coordinate system O... W -X W Y W Z W , take the origin O W The bottom left corner, or southwest corner, of the rectangular world map is located on the X-axis. W Y W Z W Direction X W The positive direction is due east on the map, Y W The positive direction is due north on the map, Z W The positive direction is perpendicular to the top of the map; the image coordinate system O I -X I Y I and pixel coordinate system O P -uv are two-dimensional coordinate systems used to describe the positions of pixels in an image. In the image coordinate system O I -X I Y I In the process, the center of the imaging plane is taken as the origin O. I coordinate axis X I Y I Direction X I The positive direction is horizontal to the right of the image plane, Y I The positive direction is downward from the image plane; in the pixel coordinate system O P -u a v a In the image, the top-left pixel is taken as the origin O. P coordinate axis u a v a In the direction of u a The positive direction is the pixel index to the right horizontally, v a The positive direction is the vertical downward pixel index.

[0072] As attached Figure 3 Appendix Figure 4 and attached Figure 5As shown, in a further specific embodiment, the vertex information of the 2D vehicle detection box and the vertex information of the 3D vehicle detection box are transformed into pixel coordinates after multiple coordinate system transformations, including the following steps:

[0073] The coordinate information of the 3D vehicle detection box vertices in the LiDAR coordinate system is converted into coordinate information in the world coordinate system. Then, the coordinate information in the world coordinate system is converted into coordinate information in the camera coordinate system. Next, the coordinate information in the camera coordinate system, including the coordinate information of the 2D vehicle detection box vertices and the converted 3D vehicle detection box vertices, is converted into coordinate information in the image coordinate system. Finally, the coordinate information in the image coordinate system is converted into coordinate information in the pixel coordinate system and output.

[0074] The steps for each coordinate information transformation include: aligning the origins of the two different coordinate systems using a translation matrix t, and then aligning the coordinate systems using a rotation matrix R. The specific transformation process is as follows:

[0075] The coordinate information of the vertices of the 3D vehicle detection box is represented by homogeneous coordinates in the lidar coordinate system as P. L =x L ,y L ,z L ,1] T The extrinsic matrix needed to transform it to the world coordinate system is as follows:

[0076]

[0077] Wherein, the rotation matrix R from the lidar coordinate system to the world coordinate system WL For a 3×3 matrix, the translation matrix t WL It is a 3×1 vector;

[0078] Coordinates of points in the transformed world coordinate system:

[0079] P W =T WL ·P L

[0080] The coordinates obtained in the previous step are transformed from the world coordinate system to the camera coordinate system for projection onto the planar image coordinate system. The inverse matrix of the extrinsic parameter matrix required is as follows:

[0081]

[0082] Among them, the rotation matrix R for transforming the world coordinate system to the camera coordinate system CW For a 3×3 matrix, the translation matrix t CW It is a 3×1 vector;

[0083] Obtain the coordinates of the camera in the camera coordinate system:

[0084] P C =T CW ·P W

[0085] After completing the spatial transformation, the coordinates P obtained in the previous step are... C =[x C ,y C ,z C ] T By projecting the image onto the image coordinate system using the pinhole camera imaging principle, and normalizing the points in the camera coordinate system, the projection onto the image coordinate system can be obtained. The coordinates (x, y) of the points in the camera coordinate system are then determined. C ,y C ,z C The projection of ) onto the image coordinate system is (x I ,y I The formula is as follows:

[0086]

[0087] Alternatively, it can be written in homogeneous coordinate form, as shown below:

[0088]

[0089] Where f is the focal length of the camera. In reality, the image captured by the camera is displayed in pixel format, therefore it is necessary to transform the image coordinate system to the pixel coordinate system. The origin of the image coordinate system is O. I If the position in the pixel coordinate system is (u0, v0), then the coordinate transformation relationship is as follows:

[0090]

[0091] Where dx and dy represent the x-axis of a point in the pixel coordinate system and the x-axis of the point in the image coordinate system, respectively. I axis and Y I The unit length value of the axis, written in homogeneous coordinate form, is as follows:

[0092]

[0093] Finally, the coordinates are output in pixel coordinate system.

[0094] The key process for aligning and converting LiDAR and camera data described above is: converting point P in the LiDAR coordinate system... L Through the extrinsic parameter matrix T WL Transform to world coordinate system P W Then, through the inverse matrix of the camera world's extrinsic parameters Convert to point P in the camera coordinate system CThen, it is projected onto the image coordinate system P using the pinhole camera imaging principle. I Finally, the pixel coordinates P are obtained through image scaling and translation. P Complete the spatial transformation from LiDAR coordinate system points, camera coordinate system points to pixel coordinate system.

[0095] In a further specific embodiment, the above-mentioned decision-level fusion method, using the detection box IoU and the Hungarian algorithm, matches the two projection boxes formed by the transformed 2D vehicle detection box vertex information and the 3D vehicle detection box vertex information, and outputs the fused detection result based on the matching result, including the following steps:

[0096] Image data captured by the camera is input into an improved YOLOv8 detection algorithm that integrates a lightweight cross-scale feature fusion module, a dynamic detection head, and a CBAM attention mechanism to enhance image target detection capabilities. The algorithm outputs 2D vehicle detection box vertex information and confidence scores. Point cloud data collected by the LiDAR is input into the PointPillars model for detection and output of 3D vehicle detection box vertex information. The specific formula is as follows:

[0097]

[0098]

[0099] Among them, C i It is the output of the i-th vehicle detection result, CBB i 2D It is the 2D vehicle detection bounding box of vehicle i, including the vertex information of the 2D vehicle detection bounding box, CL i It is confidence level, used to measure the reliability of visual inspection results. j This is the output of the j-th vehicle detection result, LBB. j 3D This is the 3D vehicle detection bounding box of vehicle j, including the vertex information of the 3D vehicle detection bounding box, (x Lj ,y Lj ,z Lj () represents the cloud coordinates of the center point of the target vehicle;

[0100] As attached Figure 6 As shown, the vertex information of the 2D vehicle detection bounding box output by the visual detection branch using the improved YOLOv8 model is projected onto a 2D projection box CBB in the pixel coordinate system. 2D Represented by the coordinates of the endpoints of the two diagonals:

[0101] CBB = (u1, v1, u2, v2)

[0102] The point cloud detection branch uses the 3D vehicle detection bounding box vertex information LBB output by the PointPillars model. 3D After spatial synchronization, the image is projected onto a 2D projection frame (LBB) in the pixel coordinate system. 2D Similarly, the coordinates of the endpoints of the two diagonals are expressed as follows:

[0103] LBB = (u3, v3, u4, v4)

[0104] For time information that is the same or similar, or at the same timestamp CBB 2D and LBB 2D The overlapping region S C ∩S L Represented by coordinates:

[0105] S C ∩S L =(u3,v3,u2,v2)

[0106] Among them, S C CBB 2D The area, S L Indicates LBB 2D The area is calculated using the following formula:

[0107]

[0108] According to the definition of IoU, CBB 2D and LBB 2D The intersection-union ratio is the ratio of the area of ​​the intersection to the area of ​​the union. From the above formula and mathematical relationships, the final formula for calculating the intersection-union ratio is as follows:

[0109]

[0110] Complete CBB 2D and LBB 2D After calculating the crossover-union ratio (CBB), the CBB of the group is determined based on the set IoU threshold. 2D and LBB 2D Determine whether the target matches and output the fusion detection result based on the matching result.

[0111] In a further specific embodiment, the above-mentioned output of the fusion detection result based on the matching result includes the following steps:

[0112] If the calculated intersection-union ratio (IU) is greater than a preset threshold, the detection results of that group are considered a successful match, pointing to the same target; if it is less than the preset threshold, they are considered a mismatch, and the next group is selected for the next round of matching. This process obtains the optimal matching result by calling the Hungarian matching function in the library to search the IU matrix.

[0113] This paper adopts the following strategies to address several situations that may arise during the matching process:

[0114] 1) Image recognition successful, point cloud recognition successful: If they are determined to be the same target, the fusion detection result is directly output.

[0115] 2) Image recognition failed, but point cloud recognition succeeded: This may be because the camera failed to detect the target due to environmental factors, such as nighttime, fog, or obstruction, and directly output the point cloud detection result.

[0116] 3) Image recognition successful, point cloud recognition unsuccessful: This may be because the target is too far away from the lidar and the target is not detected, so the visual recognition result is output directly.

[0117] 4) Image recognition failed, point cloud recognition failed: The default is that there is no target at this location.

[0118] In another embodiment, the present invention provides a computer-readable storage medium storing a computer program that causes a computer to execute a traffic operation status collaborative sensing method as described above.

[0119] In another embodiment, the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements a traffic operation status collaborative perception method as described above.

[0120] In the embodiments disclosed in this application, a computer storage medium may be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of computer storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0121] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0122] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for collaborative perception of traffic operation status, characterized in that, Includes the following steps: Image data and point cloud data are acquired using a camera and a LiDAR, respectively. The YOLOv8 model is used to extract 2D vehicle detection box vertex information and image time information from the image data, and the PointPillars model is used to extract 3D vehicle detection box vertex information and point cloud time information from the point cloud data. The acquired information is time-synchronized and filtered to obtain the corresponding 2D and 3D vehicle detection box vertex information when the image time information and point cloud time information are the same or similar, and then converted into the same coordinate form. A decision-level fusion method is adopted, which uses the IoU of the detection boxes to match the two projection boxes formed by the vertex information of the transformed 2D vehicle detection boxes and the vertex information of the 3D vehicle detection boxes, and outputs the fused detection result based on the matching result; finally, the BoT-SORT algorithm is used to receive the fused detection result and output the traffic operation status of the target vehicle.

2. The traffic operation status collaborative perception method according to claim 1, characterized in that, The step of performing time-synchronized filtering on the collected information to obtain the corresponding 2D vehicle detection box vertex information and 3D vehicle detection box vertex information when the image time information and point cloud time information are the same or similar includes the following steps: Based on the point cloud time information, the time closest to the corresponding point cloud time information timestamp is selected from the image time information for matching, and the corresponding 2D vehicle detection box vertex information and 3D vehicle detection box vertex information are obtained for each point cloud time information and its matching image time information.

3. The traffic operation status collaborative perception method according to claim 1, characterized in that, The conversion to the same coordinate form includes the following steps: Establish camera coordinate system, lidar coordinate system, world coordinate system, image coordinate system and pixel coordinate system. After transforming the vertex information of 2D vehicle detection box and 3D vehicle detection box into multiple coordinate systems, output the pixel coordinate form.

4. The traffic operation status collaborative perception method according to claim 3, characterized in that, The establishment of the camera coordinate system, lidar coordinate system, world coordinate system, image coordinate system, and pixel coordinate system includes the following steps: Establish camera coordinate system O C -X C Y C Z C LiDAR coordinate system O L -X L Y L Z L World coordinate system O W -X W Y W Z W Image coordinate system O I -X I Y I and pixel coordinate system O P -uv; in the camera coordinate system O C -X C Y C Z C In the middle, the optical center of the camera installed on the traffic light pole is taken as the origin O. C coordinate axis X C Y C Z C Direction X C The positive direction is the horizontal rightward direction of the image, Y C The positive direction is the bottom right of the image horizontally, Z. C The positive direction is perpendicular to the front of the camera; in the lidar coordinate system O L -X L Y L Z L In the above, the laser emission point of the lidar sensor is taken as the origin O. L coordinate axis X L Y L Z L Direction X L The positive direction points directly ahead on the road, Y L The positive direction is perpendicular to the left of the lidar, Z L The positive direction is perpendicular to the ground and points to the sky; in the world coordinate system O W -X W Y W Z W , take the origin O W The location is the southwest corner of the world map, with the X-axis as the coordinate axis. W Y W Z W Direction X W The positive direction is due east on the map, Y W The positive direction is due north on the map, Z W The positive direction is perpendicular to the sky above the map; in the image coordinate system O I -X I Y I In the process, the center of the imaging plane is taken as the origin O. I coordinate axis X I Y I Direction X I The positive direction is horizontal to the right of the image plane, Y I The positive direction is downward from the image plane; in the pixel coordinate system O P -u a v a In the image, the top-left pixel is taken as the origin O. P coordinate axis u a v a In the direction of u a The positive direction is the pixel index to the right horizontally, v a The positive direction is the vertical downward pixel index.

5. The traffic operation status collaborative perception method according to claim 4, characterized in that, The step of transforming the vertex information of the 2D vehicle detection box and the 3D vehicle detection box into pixel coordinates after multiple coordinate system transformations includes the following steps: The coordinate information of the 3D vehicle detection box vertices in the LiDAR coordinate system is converted into coordinate information in the world coordinate system. Then, the coordinate information in the world coordinate system is converted into coordinate information in the camera coordinate system. Next, the coordinate information in the camera coordinate system, including the coordinate information of the 2D vehicle detection box vertices and the converted 3D vehicle detection box vertices, is converted into coordinate information in the image coordinate system. Finally, the coordinate information in the image coordinate system is converted into coordinate information in the pixel coordinate system and output.

6. The traffic operation status collaborative perception method according to claim 5, characterized in that, The steps include: The coordinate information of the vertices of the 3D vehicle detection box is represented by homogeneous coordinates in the lidar coordinate system as P. L =[x L ,y L ,z L ,1] T The extrinsic matrix needed to transform it to the world coordinate system is as follows: Wherein, the rotation matrix R from the lidar coordinate system to the world coordinate system WL For a 3×3 matrix, the translation matrix t WL A 3×1 vector; coordinates of the points in the transformed world coordinate system: P W =T WL ·P L The coordinates obtained in the previous step are transformed from the world coordinate system to the camera coordinate system for projection onto the planar image coordinate system. The inverse matrix of the extrinsic parameter matrix required is as follows: Among them, the rotation matrix R for transforming the world coordinate system to the camera coordinate system CW For a 3×3 matrix, the translation matrix t CW It is a 3×1 vector; Obtain the coordinates of the camera in the camera coordinate system: P C =T CW ·P W After completing the spatial transformation, the coordinates P obtained in the previous step are... C =[x C ,y C ,z C ] T By projecting the image onto the image coordinate system using the pinhole camera imaging principle, and normalizing the points in the camera coordinate system, the projection onto the image coordinate system can be obtained. The coordinates (x, y) of the points in the camera coordinate system are then determined. C ,y C ,z C The projection of ) onto the image coordinate system is (x I ,y I The formula is as follows: Or in homogeneous coordinate form, as follows: Where f is the focal length of the camera, and the image coordinate system is transformed to the pixel coordinate system, with the origin O of the image coordinate system being O. I If the position in the pixel coordinate system is (u0, v0), then the coordinate transformation relationship is as follows: Where dx and dy represent the x-axis of a point in the pixel coordinate system and the x-axis of the point in the image coordinate system, respectively. I axis and Y I The unit length value of the axis, written in homogeneous coordinate form, is as follows: Finally, the coordinates are output in pixel coordinate system.

7. The traffic operation status collaborative perception method according to claim 6, characterized in that, The method employs a decision-level fusion approach, utilizing the IoU of the detection boxes to match two projected boxes formed by the vertex information of the transformed 2D vehicle detection boxes and the vertex information of the 3D vehicle detection boxes, and outputs the fused detection result based on the matching result. This includes the following steps: Project the vertex information of the 2D vehicle detection bounding box onto the 2D projection bounding box CBB in the pixel coordinate system. 2D Represented by the coordinates of the endpoints of the two diagonals: CBB = (u1, v1, u2, v2) Project the vertex information of the 3D vehicle detection bounding box onto the 2D projection bounding box LBB in the pixel coordinate system. 2D Similarly, the coordinates of the endpoints of the two diagonals are expressed as follows: LBB = (u3, v3, u4, v4) For CBBs with the same or similar time information 2D and LBB 2D The overlapping region S C ∩S L Represented by coordinates: S C ∩S L (u3,v3,u2,v2) Among them, S C CBB 2D The area, S L Indicates LBB 2D The area is calculated using the following formula: According to the definition of IoU, the final formula for calculating the intersection-union ratio is as follows: After completing the crossover ratio (CBB) calculation, the CBB of the group is determined based on the set IoU threshold. 2D and LBB 2D Determine whether the target matches and output the fusion detection result based on the matching result.

8. The traffic operation status collaborative perception method according to claim 6, characterized in that, The step of outputting the fusion detection result based on the matching result includes the following steps: If the calculated intersection-union ratio (IUU) is greater than the preset threshold, the detection results of this group are considered to be a successful match, pointing to the same target; if it is less than the preset threshold, it is considered a mismatch, and the next group is selected for the next round of matching.

9. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute a traffic operation status collaborative perception method as described in any one of claims 1-8.

10. An electronic device, characterized in that, include: The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements a traffic operation status collaborative perception method as described in any one of claims 1-8.