A dynamic target point cloud rapid identification and point cloud segmentation method based on a roadside perception unit

By using roadside lidar and point cloud data processing technology, dynamic and static objects can be quickly identified and segmented, solving the problems of limited perception range and insufficient decision-making ability of on-board equipment, improving the perception and decision-making capabilities of autonomous driving systems, and reducing traffic safety risks.

CN115605777BActive Publication Date: 2026-05-15TONGJI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2021-04-01
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing autonomous driving technologies, the onboard equipment has limited perception capabilities and a restricted perception range, as well as insufficient decision-making ability, making it difficult to cope with complex and ever-changing traffic environments, resulting in high traffic safety risks.

Method used

By employing roadside sensing units, point cloud data is collected using roadside lidar. Invalid data is eliminated using boundary equations, and voxelization is performed. Combined with point cloud target detection algorithms, key traffic participants are screened out, reducing the amount of data and improving recognition accuracy.

Benefits of technology

This enables rapid identification of dynamic and static objects in vehicle-road cooperative systems, reduces data transmission volume, enhances the perception range and decision-making capabilities of autonomous vehicles, and improves traffic safety and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115605777B_ABST
    Figure CN115605777B_ABST
Patent Text Reader

Abstract

A kind of quick identification and point cloud segmentation method of dynamic and static object based on roadside perception unit, comprising the following steps: 1, construct roadside holographic perception (laser radar) scene, collect a batch of point cloud data as prior information;2, establish effective target identification range, and extract static point cloud background for subsequent matching;3, by comparing the point cloud data of current frame to be identified with the voxel features of recorded static point cloud background, identify the non-static region with larger change difference, and the rest is recorded as static object;4, record non-static region as temporary static region at fixed frequency, identify short-time static object and dynamic object by comparing non-static region with recorded temporary static region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, and specifically relates to a method for rapid identification of dynamic and static objects and point cloud segmentation based on roadside sensing units, mainly for target perception on the infrastructure side in a vehicle-road cooperative environment. Background Technology

[0002] With the continuous improvement and development of my country's national economy and automotive industry technology, automobiles have become an indispensable means of transportation for our daily lives and production. However, it must be acknowledged that while automobiles have greatly improved our lifestyles, they have also brought about increasingly serious problems such as traffic accidents, traffic congestion, economic losses, and environmental pollution. To improve the overall traffic environment, governments and experts worldwide are actively exploring effective solutions to these traffic safety issues, leading to the emergence of autonomous driving technology. Using the automobile as a hardware and software platform, onboard sensors, decision-making units, and actuators provide intelligent support, enabling the vehicle to make driving decisions based on its surroundings. This avoids traffic risks caused by varying levels of driver competence, thereby improving vehicle safety. Furthermore, with the increasing maturity of dedicated short-range communication technology, sensor technology, and vehicle control technology, the pace of autonomous and driverless technologies moving from the laboratory to practical applications is accelerating.

[0003] However, judging from the current development trend of autonomous driving technology, relying solely on single-vehicle intelligent systems will be insufficient to effectively solve traffic safety risks. The main reasons are as follows: 1. Onboard equipment has limited perception capabilities, which may lead to insufficient perception data and decision-making errors in certain scenarios; 2. Onboard equipment has limited perception range. Due to installation location limitations, many sensing devices cannot fully utilize their theoretical perception range and are often obstructed by other vehicles around the autonomous vehicle, causing potential dangers to be overlooked; 3. Decision-making capabilities need improvement. The current level of intelligence in autonomous driving decision-making systems is insufficient to cope with complex and ever-changing traffic environments. Based on these problems, countries around the world have begun to propose a new solution: vehicle-road cooperative systems. The basic idea of ​​vehicle-road cooperative systems is to use a multidisciplinary approach, leveraging advanced wireless communication and sensing technologies to acquire real-time vehicle and road information. Through vehicle-to-vehicle and vehicle-to-infrastructure communication, information exchange and sharing are achieved between vehicles and between vehicles and intelligent roadside facilities, realizing intelligent collaboration between vehicles and roads. This improves road traffic safety, road efficiency, and the utilization rate of road traffic system resources. Typically, based on the sensor installation locations within the system, vehicle-to-infrastructure (V2I) systems can be divided into two subsystems: intelligent roadside systems and intelligent in-vehicle systems. The intelligent roadside system primarily handles tasks such as acquiring and disseminating traffic flow information, controlling roadside equipment, vehicle-to-infrastructure communication, and traffic management and control. The intelligent in-vehicle system mainly handles tasks such as acquiring vehicle motion status information and information about the vehicle's surrounding environment, vehicle-to-vehicle / vehicle-to-infrastructure communication, safety warnings, and vehicle auxiliary control. The intelligent roadside system and the intelligent in-vehicle system transmit and share information through vehicle-to-infrastructure communication, thereby achieving data interaction, expanding the perception range and data volume of autonomous vehicles, enhancing the decision-making basis of autonomous vehicles, and improving driving safety.

[0004] In terms of current development, the main application scenario for intelligent roadside systems is providing auxiliary perception information for autonomous vehicles. Roadside perception devices offer more variety in type and installation location compared to vehicle-mounted perception devices, and are more tolerant of stringent requirements such as energy supply. Common roadside perception devices include: 1. Loop coils; 2. Millimeter-wave radar; 3. UWB technology; 4. Vision methods; 5. LiDAR, etc. Among these perception methods, visual detection methods and LiDAR technology are the main technologies that can be used as engineering solutions. Both have the advantages of simple and easy-to-understand data formats and relatively mature target detection technologies. However, compared to vehicle-mounted systems, roadside visual data must be the detection results transmitted to serve the vehicle, because image data with large differences in viewing angles can hardly achieve data fusion at the original data level; only the detection results can be fused. LiDAR data, on the other hand, is in the form of point cloud coordinates, and data fusion can be achieved through coordinate axis transformation. This pre-fusion data fusion method results in less loss of semantic information compared to post-fusion based on detection results, which is more beneficial for improving the recognition accuracy of the detection results.

[0005] However, it should be noted that even though the front fusion method is more advantageous, it still has limitations. Its biggest limitation is the larger data transmission volume, because the transmitted data is raw data, and the size of a frame of typical point cloud data is between several megabytes and tens of megabytes. That is, the data transmitted one-to-one per second can be as high as several hundred megabytes. In the case of complex intelligent transportation environments with a large number of vehicles, the total data transmission volume per second can even reach several gigabytes. Therefore, the size of the transmitted data must be reduced.

[0006] There are many methods to reduce data size, such as point cloud sparsity sampling and skeleton structure extraction. One effective method suitable for autonomous driving scenarios is to extract key parts of the data. For example, for a target vehicle, the importance of surrounding vehicles, non-motorized vehicles, pedestrians, and road barriers far outweighs that of the road surface, green plants, and buildings on both sides. Therefore, low-value objects can be identified and filtered out from the original data to reduce the data volume. In addition, after removing these low-value objects, the usability of the remaining objects will naturally increase because the presence of important objects becomes more prominent after eliminating interference factors, avoiding the influence of low-value data.

[0007] Existing technology

[0008] CN106780458B

[0009] CN110892453A

[0010] CN110796724A

[0011] CN103530899A

[0012] Terminology Explanation

[0013] To make the description of this invention more accurate and clear, the various terms that appear in this invention are explained as follows:

[0014] Roadside sensing unit: In vehicle-road cooperative scenarios, environmental sensing devices are deployed around the road using roadside poles or gantries as mounting bases. In this invention, roadside sensing unit specifically refers to a lidar system, and the two should be considered as the same description.

[0015] Point cloud data: A collection of vectors in a three-dimensional coordinate system. These vectors typically include at least X, Y, and Z coordinate data to represent the shape of an object's outer surface. Depending on the device used, other information may also be acquired. In this invention, the symbol D is typically used to represent point cloud data and its processed components. All content expressed using the symbol D should be understood as representing point cloud data.

[0016] Static object D s This mainly refers to road surfaces and their ancillary facilities, buildings, and roadside greenery. Without considering reconstruction, expansion, or frequent road maintenance, these objects are in a state of long-term fixed position and unchanged appearance. Objects whose position and appearance do not change significantly within a month can be considered static objects. The criteria for determining whether an object has undergone a significant change are: a shift in the centroid of its horizontal projection exceeding 1 meter is considered a significant change; or a change in the outer contour length or volume exceeding 5% of the original data is also considered a significant change.

[0017] Short-time static object D st This mainly includes temporarily parked vehicles and standing pedestrians. These objects are in a short-term, unchanging position or state, but the possibility of movement in the next moment cannot be ruled out. In this invention, objects that are not considered static and do not undergo significant changes in position or appearance within 5 frames are considered short-term static objects. The criteria for judging significant changes are: a centroid offset of more than 0.3m from the horizontal projection of the object is considered a significant change, or a change in the outer contour side length or volume exceeding 5% of the original data is considered a significant change.

[0018] Dynamic Object D d This mainly includes moving vehicles, walking pedestrians, etc. These objects are in motion when observed and can be considered as not static objects. Objects that undergo significant changes in position or appearance within two consecutive frames are considered dynamic objects. The criteria for judging significant changes are: a centroid shift of the object's horizontal projection exceeding 0.3m is considered a significant change, or a change in the outer contour side length or volume exceeding 5% of the original data is considered a significant change.

[0019] Non-static object D ns The sum of short-term static and dynamic objects.

[0020] Raw data D0: The point cloud dataset used in the preprocessing part of this invention, generally containing about 1000 frames of point cloud data, and should include most common traffic scenarios on the road sections where the road test sensing units are installed.

[0021] Point cloud data to be identified: This is different from the original data D0 mentioned above. It is the point cloud data frame that is actually used for identification during the use of this invention. Its identification results can be used to support subsequent point cloud target detection and other tasks. In the description, it is usually referred to as the i-th frame data D. i .

[0022] Key road surface areas: Areas containing road surfaces that are identified using the method described in this invention, typically road areas that are clearly distinguishable within the range of lidar scanning and are major traffic routes.

[0023] Data value: The extent to which point cloud data influences driving decisions when used by autonomous vehicles. In this invention, the criterion is the potential danger that an object may pose to an autonomous vehicle. Without considering the autonomous driving system's ability to understand data, it is generally believed that dynamic objects have the greatest data value, followed by short-term static objects, and static objects have the least.

[0024] Invalid data refers to point cloud data that has almost no data value. It typically includes building edges, slopes, and open spaces along roads, and is generally point cloud data located more than 5 meters away from the road edge. In practical applications, the scope of invalid data can be adjusted according to the needs of identification. For example, in road scenarios with roadside guardrails, such as urban expressways, point cloud data outside the guardrails can be considered invalid data.

[0025] Valid data: Point cloud data remaining after invalid data has been removed. In this invention, the operation of extracting valid data is indicated by adding a single quote (').

[0026] Boundary equations E b The function boundary is used to separate invalid and valid data. After projecting the point cloud data into a bird's-eye view, the boundary points are manually selected and then established by fitting using the least squares method.

[0027] Static point cloud background B: A pure static background space that does not contain any short-term static or dynamic objects. For the vehicle-road cooperative scenario that this invention addresses, it means that the traffic environment does not contain any traffic participants such as non-permanently parked vehicles, non-motorized vehicles, and pedestrians.

[0028] Valid point cloud data to be identified: The point cloud data to be identified is processed using the boundary equation system E b After cropping, the resulting point cloud dataset is the valid point cloud data to be identified, denoted as the valid i-th frame data D′ in the description. i .

[0029] Statistical Space: Point cloud data has the characteristic of being dense in the near and sparse in the far. In order to avoid this characteristic affecting the subsequent recognition work, the point cloud data is divided into multiple statistical spaces and the point cloud density is made approximately equal in all places by using overlay and downsampling operations.

[0030] Point cloud density: An indicator used to describe the density of point clouds, characterized by the number of points per unit volume. Specific calculation methods and parameters can be determined based on the equipment available.

[0031] Scanning distance L: refers to the distance between the point cloud and the center of the lidar, which can be characterized by the planar distance after projecting the point cloud data into a bird's-eye view.

[0032] Point cloud target detection algorithm: refers to the target detection algorithm used to detect and determine the point cloud data into a specific category (such as large vehicles, small vehicles, pedestrians, etc.). Its function is different from the target recognition method described in this article. The method described in this article only recognizes specific point cloud datasets and does not detect their specific categories.

[0033] Detection confidence P: The confidence level of the output result obtained by inputting point cloud data that can characterize an object into a point cloud target detection algorithm. In particular, if the point cloud target detection algorithm does not detect any result, its detection confidence is considered to be 0.

[0034] The trigger distance threshold DT is defined as the range of perception distances within which most point cloud object detection algorithms can perform well. In this invention, a good performance of a point cloud object detection algorithm is defined as the ability to detect all non-static objects within the region with a detection confidence level of not less than 75%.

[0035] Point cloud data for identification: The result of cropping point cloud data using an identification trigger distance threshold. In this invention, the operation of extracting point cloud data for identification is indicated by double quotes (”).

[0036] Point cloud data to be identified: The point cloud dataset obtained after cropping the valid point cloud data to be identified using the identification trigger distance threshold DT is the point cloud data to be identified, denoted as the i-th frame data D″ in the description. i .

[0037] Voxel: short for volume pixel, similar to the definition of a pixel in two-dimensional space. It is the smallest unit of three-dimensional space segmentation, represented as a spatial cube. The side length of the cube can be determined by the user. Different sizes of voxels describe models with different levels of detail.

[0038] Voxelization: The operation of converting point cloud data into voxels, which is referred to as voxels in the description of this invention. v express.

[0039] Voxelized point cloud data to be identified: The point cloud dataset obtained after voxelizing the point cloud data to be identified is the voxelized point cloud data to be identified, denoted as the voxelized i-th frame data D″ in the description. v .

[0040] Origin of point cloud coordinate system: Point cloud data is usually represented in three-dimensional coordinate form. In this invention, the origin of the coordinate system of point cloud data is denoted as the origin of point cloud coordinate system.

[0041] Annular spare recognition area: To avoid some static or non-static objects being incomplete due to the aforementioned cropping operation, an annular spare recognition area is added outside the recognition trigger distance threshold and divided into multiple sub-regions at a fixed angle; when a non-static region is recognized near the edge of the recognition trigger distance threshold, the horizontal angle between the non-static region and the X-axis of the point cloud coordinate system is recorded, and an annular spare recognition sub-region corresponding to this angle is added to the non-static region.

[0042] Sliding window method: Using a fixed-size data filtering box, a portion of the data is continuously filtered out from the original data along a certain direction, so that the operation is only applied to these sub-data, reducing the amount of data processing and enhancing the recognition of local features.

[0043] Background subtraction method: This usually refers to a method that uses the current frame in an image sequence to compare with a background reference model to detect non-static objects. In this invention, it is applied to 3D point cloud data.

[0044] Static area A s A static region is defined as a spatial region containing multiple voxels. If the number, location, and distribution of voxels within this region all change less than a threshold compared to the static point cloud background, then the objects or scene within this region are considered unchanged. Essentially, it is a subset of static objects.

[0045] Non-static region A ns A spatial region containing multiple voxels is considered to be a non-static region when the number, position, distribution, and other characteristics of the voxels change more than the threshold compared to the static point cloud background.

[0046] Temporary static area A st A spatial region containing multiple voxels is formed by preserving non-static regions at a fixed frequency. Short-term static regions are used to distinguish between dynamic objects and short-term static objects; the portion after separating the dynamic object is the short-term static object.

[0047] Dynamic Region A dA spatial region containing multiple voxels is considered a non-static region. When the number, position, and distribution of voxels within this region change more than a threshold compared to a short-term static region, the objects or scene within that region can be considered to have changed, thus constituting a dynamic region. Essentially, it is a subset of dynamic objects. Summary of the Invention

[0048] This invention provides a method for rapid identification of static and dynamic objects and point cloud segmentation based on roadside sensing units. Designed for real traffic environments, it assumes the roadside lidar is relatively fixed in position and uses previously collected point cloud background data as prior information. By comparing the frame to be identified with the background data, it quickly filters out areas with significant changes, greatly reducing invalid point cloud data and data transmission volume. Furthermore, the applicability of this invention is not limited to single-radar environments; it is also applicable to multi-radar networks. The extracted areas to be identified can also be fused with other data formats to improve detection accuracy. Considering the hardware and software requirements for data processing, it is recommended that the processing method described in this invention be deployed on roadside equipment or cloud-based devices.

[0049] The flowchart of this invention is as follows Figure 1 and Figure 2 As shown, its features include the following steps:

[0050] (I) Data Acquisition

[0051] For vehicle-road cooperative environments, a roadside LiDAR perception scenario is established, and a batch of point cloud data is collected for the preprocessing stage. The collected point cloud data is required to be fully covered and unobstructed, and should include static objects, mainly the road surface and its ancillary facilities, buildings, and roadside green plants, and should also include enough non-static objects, such as pedestrians, non-motorized vehicles, and vehicles, which can be manually distinguished. It is recommended that the total number should not be less than 300.

[0052] The constructed roadside lidar scene should ensure that there is no large-area obstruction within the scanning range, that is, all key road surface areas within the scanning range should be clearly visible. Figure 3 The left figure shows poorly positioned data points that lost half of the road width data due to the influence of the central divider. Figure 3 The image on the right is a better example.

[0053] The point cloud data format acquired by this invention is shown in the table below.

[0054]

[0055] Point cloud data includes three-dimensional coordinate data (X, Y, Z), reflection intensity value (Intensity), three-channel color values ​​(RGB), and return number. This invention uses only the three-dimensional coordinate data (X, Y, Z) as the basis for point cloud data extraction. When transmitting the filtering results, the three-dimensional coordinate data, reflection intensity value, and three-channel color values ​​are all transmitted together to avoid missing information that would prevent the vehicle from using the data. It should be understood that the point cloud data format applicable to this invention is not limited to the examples above; any point cloud data containing three-dimensional coordinate data (X, Y, Z) can be considered within the scope of this invention.

[0056] The data acquisition described in this invention is divided into two stages: the first is a data acquisition stage serving preprocessing, which involves collecting data to provide a data source for preprocessing. This stage requires the collected data to meet the requirements described below, aiming to comprehensively reflect the normal road conditions of the roadside sensing unit installation scenario. The second is a data acquisition stage serving daily use, which has no detailed requirements for data acquisition; it only needs to ensure the normal operation of the roadside sensing unit. In this invention, the point cloud dataset obtained in the preprocessing stage is collectively referred to as raw data D0, and the point cloud data obtained in the usage stage is named sequentially (frame number) as D1, D2, etc. The following focuses on describing the data acquisition work in the preprocessing stage.

[0057] Considering the need for post-processing, it is generally required to collect no less than 1000 frames of data. The specific number of frames can be adjusted appropriately according to the deployment scenario of the roadside sensing unit. During the data collection process, the influence of environmental factors must be considered, such as the reduction in the number of lidar echoes due to wet ground in rainy weather, resulting in a double decrease in the density and quality of road surface point clouds, or the interference of high-brightness headlights from vehicles at night, causing abnormal point cloud quality in some areas. Furthermore, although it is generally believed that the intensity of visible light has a small impact on lidar, for the sake of rigor, the collected data should be appropriately distributed across multiple time periods. Based on the above requirements, the original data collection scheme recommended by this invention is as follows:

[0058] ① When conditions permit, in heavy rain or by simulating rainy weather by spraying a large amount of water on the road surface, run the roadside lidar to collect more than 200 frames of data;

[0059] ② Illuminate the road surface at night using a high-intensity light-generating device, especially covering road signs and markings made of highly reflective materials, and collect more than 200 frames of data;

[0060] ③ In a normal environment, data is collected in three time periods: morning, noon, and evening, with more than 200 frames of data collected in each time period.

[0061] It should be understood that, apart from the above-mentioned data collection scheme, other data collection schemes adjusted according to the actual scenario should all be considered as one of the variations of the data collection scheme applicable to this invention, such as the variations exemplified below.

[0062] For regions with little rainfall throughout the year but frequent sandstorms, the following variant data collection scheme can be adopted:

[0063] ① In weather conditions with poor visibility, such as sandstorms or strong winds, the roadside lidar is operated to collect more than 200 frames of data;

[0064] ② Illuminate the road surface at night using a high-intensity light-generating device, especially covering road signs and markings made of highly reflective materials, and collect more than 200 frames of data;

[0065] ③ In a clear environment, data were collected in three time periods: morning, noon, and evening, with more than 200 frames collected in each time period.

[0066] For northern regions with long winters and severe road icing and snow cover, the following alternative data collection scheme can be adopted:

[0067] ①When there is heavy snow or the road surface is covered with ice and snow, the roadside lidar will collect more than 200 frames of data.

[0068] ② Illuminate the road surface at night using a high-intensity light-generating device, especially covering road signs and markings made of highly reflective materials, and collect more than 200 frames of data;

[0069] ③ In a clear environment, data were collected in three time periods: morning, noon, and evening, with more than 200 frames collected in each time period.

[0070] The data acquisition schemes applicable to this invention include, but are not limited to, the examples described above.

[0071] Under the aforementioned conditions, data collection should be arranged based on traffic conditions within the scanning range of the roadside lidar. At least 50% of the collected data should contain single frames with fewer than two vehicles or pedestrians, and the total number of samples of non-static objects such as vehicles or pedestrians should be no less than 300. Furthermore, there should be no prolonged obstruction of any visible area; that is, at least 90% of the frames should show clearly visible key road surfaces. If the collected data fails to meet these conditions, a new time period should be selected for re-collection.

[0072] (II) Pretreatment

[0073] First, establish the boundary equation system E. bInvalid data is removed from the original data. Then, a linear relationship between point cloud density and scanning distance is obtained by sampling the point cloud data at equal intervals. Next, static and non-static objects in each frame of point cloud data are separated. A point cloud object detection algorithm is used to detect non-static objects to establish a distribution curve of detection confidence versus scanning distance, and the scanning distance range where the confidence level is always above a threshold is selected as the recognition trigger distance threshold. Then, all static objects are cropped using the recognition trigger distance threshold. The cropped multi-frame recognition static objects are then superimposed and appropriately sampled to establish a static point cloud background B. v Finally, voxelization is performed on the static point cloud background. The flowchart for the preprocessing stage is as follows: Figure 2 As shown.

[0074] First, roadside lidar typically has a wide scanning range, with high-level equipment capable of scanning distances exceeding 100 meters, and its effective scanning range generally reaching over 50 meters. Therefore, the scanned data inevitably includes numerous objects of various types, such as surrounding buildings and greenery. Compared to vehicles, pedestrians, and road surfaces, these objects have very low data value for traffic environment detection and can be directly discarded. Therefore, in this invention, point cloud data located more than 5 meters from the road edge and possessing almost no data value are defined as invalid data, typically including buildings, slopes, and open areas on both sides of the road. Conversely, the road and non-static objects are defined as valid data. Invalid data can be removed before subsequent calculations, reducing the amount of data processing. The proposed method for removing invalid data in this invention is as follows:

[0075] ① First, project the point cloud data onto a horizontal plane, that is, only consider the X and Y values ​​in the point cloud data to form a bird's-eye view.

[0076] ② The elimination boundary is established manually. This can be done using some common point cloud data visualization processing tools, such as 3D Reshaper. The elimination boundary is to clearly separate invalid data from valid data. In practice, it is recommended to adopt a conservative strategy. If there is no obvious road boundary, it is difficult for a person to distinguish whether a certain area is a road or not. In this case, it can be considered a road area and thus considered as valid data.

[0077] ③ Construct boundary equations based on points on the boundaries. According to the previously selected boundaries, choose an appropriate number of points on each boundary and fit the plane equations of the boundaries using methods such as least squares. Generally, the number of points selected should not be less than 30, and the distance between any two points should not be less than 50 cm. If the length of the boundary cannot meet the above requirements, integration between adjacent boundaries can be considered. For computational reasons, it is recommended that the number of boundary equations not exceed 6.

[0078] ④ Finally, all boundary equations are integrated into a system of boundary equations E. bThe direction of data removal is recorded and used as the filtering condition for subsequent data filtering in the actual identification process.

[0079] It should be understood that due to differences in the installation scenarios of roadside sensing units, the definition of invalid data may reasonably vary. For example, for urban expressways or elevated roads, there should be no objects interfering with traffic flow on both sides of the road; therefore, point cloud data outside the guardrails can be removed. For the above scenarios, the following variant of invalid data removal methods can be adopted:

[0080] If, after the point cloud data is projected onto a horizontal plane, objects such as green plants and roadside facilities do not intrude into the road space, i.e., there is no situation where the point cloud data of these objects covers the road data, then the processing method is the same as the method proposed in this invention; if they have intruded into the road space, such as Figure 5 As shown, since the tree canopy has covered part of the road surface area, the bottom data should be filtered out according to the height threshold before being projected onto the horizontal plane. The height threshold is any value between the bottom of the tree canopy and the maximum vehicle height of common vehicles passing through the key road surface area, which satisfies the requirement of removing the tree canopy area while retaining vehicle data.

[0081] For gentle road sections, i.e., road sections with a longitudinal slope of no more than 3%, a fixed height threshold can be selected. For steep road sections, i.e., road sections with a slope greater than 3%, the height threshold can adopt a stepped distribution or construct a spatial plane equation. A stepped distribution means that for road sections whose planar coordinates lie within a certain area, the height threshold is selected as the same fixed value, and its form is as follows:

[0082]

[0083] Where X and Y correspond to the X and Y values ​​of the point cloud data, x1,x2,x3,x4 and y1,y2,y3,y4 represent the upper and lower region thresholds in the X and Y directions, respectively, H represents the height threshold, and h1 and h2 represent the height thresholds selected for different planar coordinate regions.

[0084] The spatial plane equation is obtained by fitting a plane equation using the X, Y, and Z coordinates of road surface points, and then translating it upwards to ensure the plane meets the height threshold segmentation condition. Its representation is as follows:

[0085] Ax + By + Cz + D = 0

[0086] Where A, B, C, and D are coefficients of the plane equation, and x, y, and z correspond to the X, Y, and Z coordinates of the point cloud data, respectively. During fitting, random sampling of the road segment data is required, with at least 100 sampling points. The plane equation is then fitted using the least squares method. In practical use, the Z value can be calculated based on the X and Y coordinates of the point cloud; the resulting Z value is the height threshold for the current area.

[0087] After the above processing is completed, the subsequent processing method is the same as the invalid data removal method suggested in this invention.

[0088] The invalid data removal method applicable to this invention is not limited to the method described above; other methods that achieve the same effect can be considered as variations. Furthermore, the invalid data removal method can be placed after other steps. This invention uses it as a pre-step to obtain the linear relationship between point cloud density and scanning distance only to reduce the amount of data computation. Any other technical solutions that change the order of the steps in this invention should be considered variations. After the above steps, a set of valid data D0′ is obtained.

[0089] Secondly, due to limitations in the physical acquisition capabilities of the hardware, point cloud data exhibits a characteristic of being denser closer to the interior and sparser closer to the exterior; that is, the point cloud density is closely related to the scanning distance. Point cloud target recognition algorithms are also closely related to the point cloud density of the target; generally, objects with higher point cloud density are easier to identify. Therefore, to improve the accuracy of subsequent point cloud target recognition, it is necessary to first establish the relationship between point cloud density and scanning distance. Based on the scanning principle of LiDAR, the distance between two points on the same ring line is linearly related to the distance of the ring line from the center; therefore, it can be inferred that the point cloud density and scanning distance also have a linear relationship.

[0090] like Figure 7 As shown, with a sampling interval of 0.5 meters, sample points were collected at equal intervals along each scanning ring from the inside out. The number of points within a 10cm radius centered on each sampling point was recorded as the point cloud density. An average of 30 sampling points were selected on each ring, and the statistical results were filled into the table below.

[0091] Point cloud density Sampling point 1 Sampling point 2 …… Sampling point 30 Mean point cloud density First Ring Road Second Ring Road 3rd Ring Road ……

[0092] The mean point cloud density of each loop is calculated, and the distance between the mean point cloud density and its corresponding loop is used as the x and y values, respectively. A linear relationship between point cloud density and scanning distance is established by fitting the data using the least squares method, expressed as:

[0093] ρ=k·L

[0094] Where ρ is the point cloud density, L is the scanning distance, and k is the linear function parameter.

[0095] After completing the above steps, each frame of point cloud data needs to be manually extracted and divided into two parts: the static object D. s and non-static object D nsStatic objects refer to road surfaces and their ancillary facilities, buildings, and roadside greenery, which remain unchanged in position and state over a long period without considering reconstruction, expansion, or frequent road maintenance. Objects whose position and appearance show no significant change within a month are generally considered static. Non-static objects, on the other hand, are the sum of all objects in the point cloud data excluding static objects, and are further divided into dynamic objects (D). d and short-term static object D st Dynamic objects refer to moving vehicles, walking pedestrians, etc. These objects are in motion when observed and can be considered non-static objects. Objects that undergo significant changes in position or appearance within two consecutive frames are considered dynamic objects. Short-term static objects refer to temporarily parked or standing pedestrians, etc. These objects are in a short-term state without change in position or state, but the possibility of movement in the next moment cannot be ruled out. In this invention, objects that do not belong to static objects and do not undergo significant changes in position or appearance within five frames are considered short-term static objects.

[0096] Static object D s Used to extract static point cloud background B. A static point cloud background refers to a purely static background space that does not contain any short-term static or dynamic objects. For the vehicle-road cooperative scenario addressed in this invention, this means a traffic environment that does not include any non-permanently parked vehicles, pedestrians, or other traffic participants. The effect is as follows: Figure 4 As shown.

[0097] Non-static object D ns This is used to obtain the recognition trigger distance threshold DT. The recognition trigger distance threshold refers to the range of perception distances that allows most point cloud object detection methods to perform well. Because point cloud data exhibits a phenomenon of sparse outer edges and dense inner edges, distant non-static objects may only be described by one or two scan lines. Such a sparse point cloud is difficult for most point cloud object detection algorithms to detect. Therefore, to meet the detection requirements of most point cloud detection algorithms, it is necessary to establish a suitable recognition trigger distance threshold, which represents the trigger distance for subsequent methods. For point cloud data located outside the recognition trigger distance threshold, even if non-static objects exist, they may not be detected, or the confidence of the detection result may be too low, potentially leading to incorrect decisions by the vehicle. Therefore, point cloud data outside the recognition trigger distance threshold will be considered low-value data and discarded.

[0098] To obtain the recognition trigger distance threshold DT, the relationship between the scanning distance L and the detection confidence P must first be established. The final results of this invention will be provided to the vehicle. However, the point cloud object detection algorithms built into the systems of various autonomous vehicle manufacturers differ. Therefore, for practical considerations, this invention selects several common point cloud object detection algorithms as test algorithms for the preprocessing stage, including VoxelNet, PIXOR, and Point R-CNN. Among them:

[0099] VoxelNet is a typical voxelized point cloud processing method. It divides a 3D point cloud into a certain number of voxels. After random sampling and normalization of the points, it uses several voxel feature encoding layers to extract local features for each non-empty voxel, obtaining Voxel-wise features. Then, it further abstracts the features through 3D convolution operations, increasing the receptive field and learning geometric spatial representations in the process. Finally, it uses a Region Proposal Network to classify, detect, and regress the objects.

[0100] PIXOR is a typical image-based point cloud processing method. It projects the point cloud to obtain a two-dimensional bird's-eye view with height and reflectance channels, and then uses a structurally fine-tuned RetinaNet for object detection and localization. The overall processing is more similar to traditional image object detection methods.

[0101] PointR-CNN is a typical point cloud processing method that utilizes raw point cloud data structures. The entire framework consists of two stages: the first stage generates 3D proposals from the bottom up, and the second stage modifies the proposals in canonical coordinates to obtain the final detection results. The first stage sub-network does not generate a small number of high-quality 3D examples directly from the point cloud in a bottom-up manner by segmenting the entire scene's point cloud into foreground and background points. The second stage sub-network converts the pooled points of each example into canonical coordinates to better learn local spatial features. This process is combined with the global semantic features learned for each point in the first stage for Box optimization and confidence prediction.

[0102] The three methods described above are typical representatives of the three most mainstream point cloud object detection algorithms, and can effectively simulate the intelligent perception of autonomous vehicles. It should be understood that the three algorithms selected in this invention cannot fully represent all point cloud object detection algorithms. Therefore, it is reasonable to use other point cloud object detection algorithms as test algorithms in the preprocessing stage, and these should be considered as variations.

[0103] Since the length of common vehicles is generally between 3.5 and 4.5 meters, they typically cross multiple scan lines, which may result in dense point cloud data in the front of the vehicle and sparse point cloud data in the rear, or even complete occlusion with almost no point cloud data. Therefore, it is necessary to determine the average point cloud density of non-static objects, such as vehicles.

[0104] This invention also employs random sampling to obtain the average point cloud density for non-static objects, but the sampling method here differs slightly from the point cloud sampling described above, as detailed below. The point cloud sampling ratio needs to be determined with reference to hardware device parameters and the actual total number of point clouds in the target. The random sampling method recommended by this invention is as follows:

[0105] ① For non-static objects with a total number of point clouds greater than 3000, random sampling is performed 3 times, with 300 points sampled each time, and the average point cloud density of the 3 samplings is calculated.

[0106] ② For non-static objects with a total number of point clouds greater than 1000, random sampling is performed 3 times, with 100 points sampled each time, and the average point cloud density of the 3 samplings is calculated.

[0107] ③ For non-static objects with a total number of point clouds greater than 500, random sampling is performed 3 times, with 75 points sampled each time, and the average point cloud density of the 3 samplings is calculated.

[0108] ④ For non-static objects with a total number of point clouds less than 500, randomly sample 100 points and calculate the average point cloud density.

[0109] In the sampling method described above, the calculation method for point cloud density is the same as described above, that is, the number of points within a radius of 10cm centered on each sampling point is taken as the point cloud density.

[0110] Besides the sampling method described above, other point cloud sampling methods used to obtain the point cloud density of non-static objects can be considered as variations of the sampling method applicable to this invention. For example, the following are variations that can employ different sampling ratios for different types of objects:

[0111] ① For pedestrians in non-static objects, random sampling is performed twice, with 100 points sampled each time. If the number of points is insufficient, all points are sampled. Finally, the average point cloud density of the two samplings is calculated.

[0112] ② For non-motorized vehicles among non-static objects, random sampling is performed 3 times, with 100 points sampled each time. If the number of points is insufficient, all points are sampled. Finally, the average point cloud density of the 3 samplings is calculated.

[0113] ③ For small cars among non-static objects, random sampling is performed 3 times, with 200 points sampled each time. If the number of points is insufficient, all points are sampled. Finally, the average point cloud density of the 3 samplings is calculated.

[0114] ④ For large vehicles among non-static objects, random sampling is performed 3 times, with 300 points sampled each time. If the number of samples is insufficient, all points are sampled. Finally, the average point cloud density of the 3 samplings is calculated.

[0115] After establishing the average point cloud density for each non-static object, each non-static object is input into different backend algorithms for detection, yielding each detection result and its detection confidence P. For example... Figure 8 As shown, plot the distribution curve of scanning distance L versus detection confidence P using the following formula:

[0116]

[0117] Where j and i represent the upper and lower limits of the recognition trigger distance threshold DT, i is the nearest distance threshold, j is the farthest distance threshold, and n i ,n j These represent the total number of dynamic targets at distances i and j from the origin, respectively. j -n i ) p>75% This represents the total number of non-static targets within the range i and j with a detection confidence level greater than 75%. It should be understood that 75% is only a recommended threshold for this invention, and its actual value can be adjusted according to the scenario in which the roadside sensing unit is installed.

[0118] It's important to note that the reason for the lower limit of the trigger distance threshold is that LiDAR devices generally have a vertical scanning angle parameter. This means that when the physical distance to the LiDAR is too close and the height is lower than the LiDAR, the vehicle cannot be scanned. In this case, some vehicles may be very close to the interior but only half of the body may be scanned. Therefore, a lower limit must be established to ensure that the entire vehicle body can be captured.

[0119] Choose appropriate values ​​for i and j to make P ij A value greater than 75% is used as the target extraction range for actual recognition. Initial values ​​for i and j are generally established by directly observing the image distribution. Based on these initial values, the upper and lower limits of the range are repeatedly adjusted with an offset of 0.5m until the region with the largest range is established, which is the final i and j value. Finally, the i and j values ​​need to be transformed into boundary equation form, generally using the equation of a circle. Therefore, the recognition trigger distance threshold DT should be represented as a ring-shaped interval, composed of the boundary equations of two circles, as shown in the following formula.

[0120] DT:i 2 ≤x 2 +y 2 ≤j 2

[0121] After obtaining the recognition trigger distance threshold DT, it is used to crop the aforementioned static object D. sSimilar to removing invalid data, point cloud data outside the trigger distance threshold DT is removed using an analogy to linear programming, resulting in the static object D″ for identification. s .

[0122] Next, the static objects to be identified need to be transformed into a static point cloud background B. A single frame of data can only reflect the scene at a specific moment; therefore, it is undoubtedly necessary to use multiple frames of data to overlay and integrate them into a point cloud background that satisfies the vast majority of scene conditions. However, due to the characteristic of point cloud data being sparse on the outside and dense on the inside, simple overlay can easily lead to denser areas becoming even denser, while sparser areas remain relatively sparse. Therefore, it is possible that the external point cloud is too sparse and thus treated as noise data, making it difficult for the system to distinguish changes in point cloud distribution, or that the internal point cloud is too dense and therefore overly sensitive to changes in point cloud distribution, even identifying point cloud changes caused by its own shaking as features of object movement. Therefore, this invention employs an interval-sampling overlay method to avoid the above problems.

[0123] Considering that lidar operates on a rotating scanning principle, the scanned data can be assumed to be distributed in a ring shape. For example... Figure 6 As shown, for each frame of point cloud data, it is divided into n statistical spaces with gradually increasing spacing from the innermost to the outermost edge. The specific spacing needs to be determined based on the scanning range and point cloud density parameters of the hardware device. The spacing recommended in this invention is as follows:

[0124]

[0125] r represents the width of the inner ring, l represents the length of the square, and R represents the distance of the inner ring from the origin.

[0126] Starting from the initial frame, the static objects for identification are sequentially superimposed from the next frame. During each superposition, the point cloud density of each statistical space needs to be calculated. The formula for calculating the point cloud density at this time is:

[0127]

[0128] Where ρ is the point cloud density, n is the total number of points contained in the statistical space, S is the horizontal projected area of ​​the statistical space, and r, l, and R have the same meaning as above.

[0129] If the point cloud density in a certain statistical space is greater than a preset threshold α, then the point cloud in that space is randomly downsampled to maintain its point cloud density. This invention suggests using a threshold of 2000 points / m². 2 .

[0130] It should be understood that the above parameter values ​​are for reference only, and the actual values ​​should be determined based on the actual performance of the roadside sensing unit. The main basis for determining these values ​​is as follows:

[0131] After subsequent processing, the point cloud density in each statistical space should be approximately equal, based on the statistical space defined by each parameter. In particular, it is necessary to check the point cloud density between the outermost and innermost statistical spaces.

[0132] The number of statistical spaces should not exceed 500, otherwise the computational load may be too large, and should not be less than 100, otherwise a single statistical space is too large, which is not conducive to the establishment of the point cloud density threshold α, and may also lead to large differences in point cloud distribution in statistical spaces with the same point cloud density, which is not conducive to subsequent calculations.

[0133] Any parameter that meets the above requirements can be used as a practical parameter for calculation.

[0134] After overlay and downsampling, a static point cloud background B is obtained. Finally, to ensure comparability in subsequent matching, it needs to be voxelized. Because of the limitations of its mechanical structure, LiDAR cannot guarantee that the laser points emitted in the previous round and the laser points emitted in the next round will hit the same location during scanning. In other words, comparing the positions of points is not only tedious and complex but also meaningless. Therefore, this invention introduces voxel features as a comparison basis.

[0135] A voxel, or volume pixel, is a unit of measurement similar to a pixel in two-dimensional space. In three-dimensional space, a voxel is the smallest unit of measurement. Voxelization represents a 3D model using voxels. It can be understood as a generalization of two-dimensional pixelation to three-dimensional space. The simplest form of voxelization is binary voxelization, where the value is either 0 or 1. Voxelization describes 3D scene data, enabling the representation of complex 3D scenes and accurately depicting vertically overlapping objects. Figure 9 As shown in the example of voxelization of point clouds, it can be seen that voxel representation involves less data and less loss of semantic information.

[0136] The size of the voxel, i.e., the side length v of the cube, needs to be determined based on the density of the collected point cloud data. If it's too large, semantic information may be lost; if it's too small, voxelization won't effectively reduce the amount of data. Testing has shown that a voxel size of 1 / 20 to 1 / 40 of the statistical space size in the first step generally yields good results. This invention recommends a voxel size of 30cm*30cm*30cm, but the actual value should be adjusted based on actual usage results.

[0137] This invention proposes to calculate the voxel position of any point in point cloud data based on the following formula:

[0138]

[0139] Where x, y, and z are the coordinates of any point in the point cloud data; x0, y0, and z0 represent the origin coordinates of the point cloud data, which are not directly 0 because there may be multiple radars in a network, causing the radar center to be not (0, 0, 0); v is the side length of the voxel. r, c, and h are the coordinate indices of the voxel.

[0140] Then, based on the coordinate index of the voxels, the number of point clouds contained in each voxel is counted. If the number of point clouds is less than 5, it is considered an empty voxel and deleted from the voxel list. Finally, all non-empty voxels are retained.

[0141] In addition to the methods described above, other voxelization methods are theoretically applicable to this invention, such as using functions from the PCL library, and should be considered as variations of the methods used in this invention.

[0142] (III) Static Object Recognition

[0143] When actually identifying a frame of point cloud data, invalid data is first removed from the point cloud data and the identification region is separated using the identification trigger threshold. Then, it is voxelized into a three-dimensional matrix. The result is compared with the static point cloud background after being voxelized using the sliding window method. If the rate of change between a certain continuous region and the background value in the sliding window is greater than the judgment threshold, the region is marked as a non-static region; otherwise, it is marked as a static region.

[0144] This invention draws on the algorithmic approach of foreground and background separation in image processing. Based on the static point cloud background obtained in the first step, it uses the sliding window method to compare it with the point cloud data distribution of the current frame. If the distribution of some areas changes too much, it can be considered that there are newly added objects in the area that do not belong to static objects, thereby extracting non-static objects.

[0145] First, the roadside sensing unit is installed in the scene to be detected, and the above preprocessing procedure is used to complete the preliminary data processing, resulting in the following: Voxelized static point cloud background B v Identify trigger distance threshold DT and boundary equation system E b .

[0146] The roadside sensing unit is then used in practical applications. The data collected during system startup is generally unstable; typically, it takes 3-5 minutes before the method described in this invention can be used for identifying moving and static targets. The first frame of point cloud data at the start is denoted as D1, and subsequent frames are denoted as D2, D3, ..., D... i .

[0147] Then, the boundary equation system E obtained through preprocessing is used. b Each frame of data is cropped once to obtain valid data D′1, D′2, ..., D′. iThen, the recognition trigger distance threshold DT is used to perform a second cropping on each frame of data, resulting in recognition data D″1, D″2, ... D″. i Finally, the identified data is voxelized to obtain voxelized data D″. v1 D″ v2 ...D″ vi .

[0148] Comparing the presence or absence of individual voxels is meaningless and fails to reflect the semantic information of objects. Therefore, this invention borrows from image processing methods such as the sliding window method and background subtraction method to match the voxelized data of the current frame with the voxelized static background from the inside out. This invention suggests using a sliding window size that is 5 times the size of the voxel frame, meaning that one sliding window can contain a maximum of 125 voxels. It should be understood that the above parameter values ​​are only for reference, and the actual values ​​should be determined based on the actual performance of the method during operation.

[0149] Because environmental vibrations can cause fluctuations in scan results, resulting in differences in point cloud distribution between frames containing static objects, a threshold needs to be set to avoid this. If the difference in voxel distribution within a window exceeds the threshold β, all voxels contained in the sliding window are marked as non-static region A. ns Otherwise, mark it as static region A. s The comparison process recommended in this invention is as follows:

[0150] ① If the highest value of the voxel Z-axis within the sliding window does not change by more than 20% compared to the same area in the static point cloud background, then the area where the window is located is marked as a static area; otherwise, proceed to the next comparison step.

[0151] ②If the total number of voxels in the sliding window does not change by more than 20% compared with the same area in the static point cloud background, the area where the window is located is marked as a static area; otherwise, proceed to the next comparison step.

[0152] ③ Calculate the centroid position of the voxels within the sliding window. Compared to the same region in the static point cloud background, if the position offset does not exceed 2 voxel side lengths, the area where the window is located is marked as a static region; otherwise, it is marked as a non-static region. The centroid calculation method is as follows:

[0153]

[0154] Where x, y, and z represent the coordinates of the centroid, x i y i z i This represents the coordinate index of each voxel, where n is the total number of voxels contained in the sliding window.

[0155] ④ When a non-static region is identified again, if it has already merged with the known non-static region A nsIf -1 is adjacent, it is also marked as A. ns -1, otherwise mark as non-static region A ns -2, and so on.

[0156] After the above comparison process is completed, the set of all static regions is the static object D in the final result. s The remaining non-static objects are further classified using the dynamic object recognition method described below.

[0157] It should be understood that the above comparison process is only one of the solutions proposed by this invention. Other methods for comparing point clouds within a sliding window with static point cloud backgrounds are applicable to this invention and should be considered as variations of this solution.

[0158] For edge regions where the identification trigger distance threshold is located, there may be cases where only half of a vehicle is included. For example, if a vehicle enters the LiDAR scanning range from outside the scanning range, it will still be identified due to the change in the point cloud distribution within the region. However, it may not be detected by the point cloud target detection algorithm or may be detected incorrectly due to incomplete data. Since the identification trigger distance threshold is usually smaller than the actual scanning range of the LiDAR, the incomplete vehicle in this case corresponds to a complete vehicle in the original data.

[0159] Therefore, this invention adds a redundancy value to the trigger distance threshold, that is, it extends a 1.5m wide annular backup recognition area outward from the boundary of the trigger distance threshold, and then divides it into sub-regions based on the scanning angle (10° in this invention), such as... Figure 10 As shown. If a dynamic region is identified in the edge region, the nearest external backup region is added to the dynamic region. Calculating the nearest external backup region can be done by directly calculating the distance between each backup region and the centroid of the dynamic region.

[0160] (iv) Non-static object recognition

[0161] In the above recognition process, all regions to be identified in a certain frame of data are recorded as temporary static regions A according to a certain frequency. st When subsequent frames are used for recognition, the area to be recognized is matched with temporary static areas. If two areas are found to have no changes in size or position, it is considered a short-term static object; otherwise, it is considered a dynamic object. Finally, after traversing the entire recognition area, the recognition results are distributed to each vehicle at different frequencies.

[0162] Based on the preceding analysis, the danger posed by short-term static objects is greater than that of static objects but less than that of dynamic objects, making them secondary identification targets for this invention. Therefore, their transmission frequency can fall between the two. Static objects are transmitted at no frequency or at a low frequency on the order of minutes, while dynamic data is transmitted at a high frequency in real time. Short-term static objects are transmitted at a medium frequency on the order of seconds.

[0163] The identification method for short-term static objects differs slightly from that for non-static objects. Since the position of a short-term static object does not change, the difference in point cloud distribution between two consecutive frames is almost negligible. In other words, if two objects that are non-static but whose features show almost no change between two consecutive frames are identified, they can be considered the same object. Based on this idea, the following method is used for identification.

[0164] During the actual recognition process, at fixed frequency intervals, all non-static regions A in the recognition frame are recorded. ns As a temporary static area A st Used for secondary matching. Specifically, the first frame of data serves as the starting frame, and its non-static regions have no temporary static regions available for comparison. Therefore, all its non-static regions are recorded as temporary static regions, but the output results are all treated as dynamic objects. The fixed frequency interval used can generally be selected as 1 / 2 to 1 / 5 of the lidar acquisition frequency. In this invention, the frequency selected is once every 5 frames.

[0165] After all non-static regions are extracted from the frame to be identified in step (3), since the non-static regions are obviously discontinuous point cloud data, each individual point cloud space is used as the matching object. A corresponding relationship table can be established in the previous matching step, or the data can be uniformly recorded and then clustered into sub-regions using Eulerian distance. This invention recommends the former. The sub-regions of both methods are denoted as A. ns -i and A st -j.

[0166] Compare non-static regions A in sequence ns With temporary static area A st For each sub-region in the above, the present invention suggests the following method for comparison:

[0167] ①Non-static region A ns and temporary static area A st The sub-regions in the two are sorted according to the scanning distance of their centroids. The centroid positions of the sub-regions in the two are compared in turn. If a sub-region A exists... ns -i、A st -j If the centroid distance between the two is no greater than 0.3 meters and there are no other matching objects within 1 meter, then proceed to the next step; otherwise, mark all non-static sub-regions that do not meet the conditions as dynamic regions.

[0168] ②A ns -i、A st -j Comparing the horizontal projection sizes of the two regions, if the two regions are located in the edge region and the rate of change is within 15%, or the two regions are located in the inner region and the rate of change is less than 5%, then proceed to the next step; otherwise, mark all non-static sub-regions that do not meet the conditions as dynamic regions.

[0169] ③A ns -i、A st Comparing the highest voxel Z-axis values ​​of the two regions, if the two regions are located at the edge and the rate of change is within 15%, or if the two regions are located in the inner region and the rate of change is less than 5%, then A is considered... ns -i can be considered as A st -j, about to A ns -i、A st -j represents the same object as a short-term static object; otherwise, all non-static sub-regions that do not meet the conditions are still marked as dynamic regions.

[0170] ④ Since dynamic objects are also recorded in the temporary static region each time it is recorded, by comparing the previous and next frames, if a sub-region A exists in the temporary static region... st -j means that if no sub-region in the non-static region of the subsequent frame data can match it, then sub-region A is considered to be... st -j is not a short-term static object, so it can be removed from the temporary static region to reduce the amount of comparison in the temporary static region later.

[0171] Finally, after the above comparison is completed, the set of all dynamic regions is the dynamic object D. d After removing dynamic objects, the remaining part of the temporary static region is the short-term static object D. st This is also the temporary static area A for the next comparison. st .

[0172] It should be understood that the above comparison process is only one of the solutions proposed by this invention. Other methods for comparing sub-regions within a non-static region with sub-regions within a temporary static region are applicable to this invention and should be considered as variations of this solution.

[0173] When recording temporary static areas, a counter can be added to the temporary static area. If a short-term static object exists after the comparison, the counter value is incremented by 1. The system can be manually set for the transmission frequency of short-term static objects. For example, if the counter is set to send every 3 increments, the transmission frequency of short-term static objects is 1 / 3 of the transmission frequency of dynamic objects.

[0174] Brief description of the attached figures

[0175] Figure 1 A flowchart of the preprocessing stage of a method for fast identification of dynamic and static objects and point cloud segmentation based on roadside sensing units.

[0176] Figure 2 A flowchart illustrating the application stages of a method for rapid identification of dynamic and static objects and point cloud segmentation based on roadside sensing units.

[0177] Figure 3 Examples of bad data samples and examples of usable data samples

[0178] Figure 4 Static point cloud background data visualization example

[0179] Figure 5 Example of data visualization of roadside green belt encroachment on road space

[0180] Figure 6 Statistical spatial division diagram

[0181] Figure 7 Point cloud sampling method illustration

[0182] Figure 8 Distribution curve of distance L versus confidence level P

[0183] Figure 9 Point cloud voxelization example

[0184] Figure 10 Additional explanation of the edge area

[0185] Figure 11 Visualization of single-frame data for test cases Detailed Implementation

[0186] According to the patent description, roadside sensing units are deployed. The example uses a pole-mounted installation method, with the lidar installed at a height of approximately 5 meters. The scanning range covers a section of two-way, two-lane road, as well as surrounding buildings, trees, and other objects. The furthest scanning distance for actual data recognition is approximately 120 meters, the data acquisition frequency is 10Hz, and the number of point clouds per frame exceeds 100,000. Visualization examples are provided. Figure 11 As shown.

[0187] First, following step one, data was collected and processed. Since the entire test was conducted under sunny conditions, lidar scanning data from rainy days could not be collected. Instead, extensive water spraying on the road surface was used to obtain approximately 720 frames of point cloud data. Approximately 600 frames were collected each around 8:00 AM, 1:00 PM, and 9:00 PM, for three consecutive days to avoid randomness. At night, the experimental vehicle was driven to the scanning area with high beams on to acquire approximately 250 frames of point cloud data under strong light conditions. The total number of frames was close to 6000. After manual filtering, approximately 2000 frames of point cloud data were obtained for extracting the static point cloud background.

[0188] The boundary equations are established manually and then visualized using common point cloud data processing tools. This invention selects a domestically produced point cloud processing software as the visualization tool. Since the selected example road segment has a clear curb as the road boundary, the road area and non-road area can be clearly distinguished. Point cloud sampling is performed in the curb area, using manual sampling at 50cm intervals along the curb's extension direction to obtain point cloud samples for fitting the road boundary. Based on these sampling points, the plane equations of the road boundary are fitted using methods such as least squares, ultimately obtaining the plane equations of the left and right road boundaries. The direction of data removal is recorded and used as the invalid data removal boundary as the first data screening condition in the subsequent actual recognition process.

[0189] Next, sampling points were collected at equal intervals along each scanning ring from the inside out, with a sampling interval of 0.5 meters. The number of points within a 10cm radius centered on each sampling point was recorded as the point cloud density. An average of 30 sampling points were selected on each ring. Some examples of the case sampling results are shown in the table below.

[0190] Point cloud density Sampling point 1 Sampling point 2 …… Sampling point 30 Mean point cloud density First Ring Road 28 23 26 27 Second Ring Road 25 24 27 26 3rd Ring Road 24 25 23 24 …… 29th Ring Road 7 8 6 9 30th Ring Road 9 6 7 9 31st Ring Road 9 7 8 8 ……

[0191] 58th Ring Road 3 2 2 2 59th Ring Road 2 2 1 1 60th Ring Road 2 0 1 1

[0192] Calculate the mean point cloud density for each loop, and use the mean point cloud density and its corresponding loop distance as x and y values, respectively. By fitting the data using the least squares method, a linear relationship between point cloud density and scanning distance can be established. The result in this case is expressed as:

[0193] ρ = 0.97·L

[0194] Where ρ is the point cloud density, L is the scanning distance, and 0.97 is the linear function parameter.

[0195] Then, using manual methods, static and non-static objects in each frame of data are separated, with static objects used to obtain a static point cloud background. However, unlike the invention described above, the point cloud target detection algorithm used in this case only requires the input of the original point cloud data; there is no need to extract non-static objects separately for detection. All point cloud target detection algorithms should be trained to achieve good recognition performance before use. This invention considers an algorithm with a precision greater than 85% to have good recognition performance. When sampling the point cloud for detected targets, the bounding boxes obtained by the algorithm can generally be used as extraction boundaries, and all point clouds within the bounding boxes are considered to belong to the detected targets.

[0196] The average point cloud density is obtained using a random sampling method. The sampling ratio is determined by referring to the parameters of the selected LiDAR equipment and the actual total number of point clouds of the target. In this case, the method used is as follows:

[0197] For non-static objects with a total number of point clouds greater than 3000, random sampling is performed 3 times, with 300 points sampled each time, and the average point cloud density of the 3 samplings is calculated.

[0198] For non-static objects with a total number of point clouds greater than 1000, random sampling is performed 3 times, with 100 points sampled each time. Finally, the average point cloud density of the 3 samplings is calculated.

[0199] For non-static objects with a total number of point clouds greater than 500, random sampling is performed 3 times, with 75 points sampled each time. Finally, the average point cloud density of the 3 samplings is calculated.

[0200] For non-static objects with a total point cloud count of less than 500, randomly sample 100 points and calculate the average point cloud density.

[0201] In the sampling method described above, the calculation method for point cloud density is the same as described above, that is, the number of points within a radius of 10cm centered on each sampling point is taken as the point cloud density.

[0202] After establishing the average point cloud density for each non-static object, the non-static objects are input into the backend algorithm for detection, yielding each detection result and the detection confidence score P. The distribution curve of the scanning distance L versus the detection confidence score P is plotted using the following formula:

[0203]

[0204] Where j and i represent the upper and lower limits of the recognition trigger distance threshold, i is the nearest distance threshold, j is the farthest distance threshold, and n i ,n j Let i and j represent the total number of non-static targets at distances i and j from the origin, respectively. j -n i ) p>75% This represents the total number of non-static targets with a confidence level greater than 75% within the range of i and j.

[0205] Choose appropriate values ​​for i and j to make P ij A value greater than 75% is used as the extraction range for non-static objects during actual recognition. In this invention, the values ​​of i and j are 3 and 45, respectively, corresponding to a non-static object extraction range from a horizontal distance of 4m from the center of the lidar to a horizontal distance of 25m from the center of the lidar.

[0206] For each frame of point cloud data, referring to the scanning range and point cloud density parameters of the LiDAR equipment used, it is divided into 93 statistical spaces with gradually increasing spacing from the innermost to the outermost edge. The spacing used in this case is as follows:

[0207]

[0208] Where r represents the width of the inner ring, l represents the length of the square, and R represents the distance of the inner ring from the origin.

[0209] Starting from the initial frame, point cloud data from the next frame is superimposed sequentially. During each superposition, the point cloud density in each statistical space is detected. The formula for calculating the point cloud density is:

[0210]

[0211] Where ρ is the point cloud density, n is the total number of points contained in the statistical space, S is the area of ​​the statistical space, r represents the width of the inner ring, l represents the length of the grid, and R represents the distance of the inner ring from the origin.

[0212] If the point cloud density of a certain statistical space is greater than the preset threshold of 2000 points / m 2 Then, the point cloud in the space is randomly downsampled to maintain its point cloud density, and finally a more ideal static point cloud background B is obtained.

[0213] To facilitate subsequent calculations, the static point cloud background is first voxelized. The voxel size used in this invention is 30cm*30cm*30cm. The voxel position of any point in the point cloud data is calculated based on the following formula:

[0214]

[0215] Where x, y, and z are the coordinates of any point in the point cloud data; x0, y0, and z0 represent the origin coordinates of the point cloud data, which are not directly 0 because there may be multiple radars in a network, causing the radar center to be not (0, 0, 0); v is the side length of the voxel. r, c, and h are the coordinate indices of the voxel.

[0216] Sort the voxels by their coordinate indices, count the number of points contained in each voxel, and if the number of points is less than 5, it is considered an empty voxel and deleted from the voxel list. Finally, all non-empty voxels are retained, which is the voxelized static point cloud background B. v .

[0217] The third step involves voxelizing the newly acquired data and using a sliding window method to filter out non-static regions. First, invalid data removal, region separation, and voxelization are performed sequentially on each frame of point cloud data, using the same method as for static point cloud background processing. After two data filtering steps, the average number of voxels in each frame of point cloud data is approximately 15,000.

[0218] Drawing inspiration from image processing methods such as the sliding window method and background subtraction, this approach matches the voxelized data of the current frame with the voxelized static background from the inside out. In this case, the sliding window size is five times the size of the voxel bounding box, meaning that a single sliding window can contain a maximum of 125 voxels.

[0219] Because environmental vibrations can cause fluctuations in scan results, resulting in differences in point cloud distribution between frames showing static objects, a trigger threshold needs to be set to avoid this. If the difference in voxel distribution within a window exceeds a fixed threshold, all voxels contained in the sliding window are marked as non-static region A. ns Otherwise, mark it as static region A. s The specific comparison process is as follows:

[0220] ⑤ If the highest value of the voxel Z-axis within the sliding window does not change by more than 20% compared to the same area in the static point cloud background, then the area where the window is located is marked as a static area; otherwise, proceed to the next comparison step.

[0221] ⑥ If the total number of voxels in the sliding window does not change by more than 20% compared with the same area in the static point cloud background, the area where the window is located is marked as a static area; otherwise, proceed to the next comparison step.

[0222] ⑦ Calculate the centroid position of the voxels within the sliding window. Compared with the same region in the static point cloud background, if the position offset does not exceed 2 voxel side lengths, the area where the window is located is marked as a static region; otherwise, it is recorded as a non-static region. The centroid calculation method is as follows:

[0223]

[0224] Where x, y, and z represent the coordinates of the centroid, x i y i z i This represents the coordinate index of each voxel, where n is the total number of voxels contained in the sliding window.

[0225] When a non-static region is identified again, if it has already merged with the known non-static region A... ns If -1 is adjacent, it is also marked as A. ns -1, otherwise mark as non-static region A ns -2, and so on.

[0226] Finally, point cloud data from all static areas are extracted as static objects. Since static objects consist of continuous point cloud data, it's difficult to compare their recognition rates; therefore, the comparison is shifted to non-static objects. In this case, the recognition rate for non-static objects in the internal region (areas less than 23m horizontally from the LiDAR center) can reach over 97%. In the edge region (areas greater than 23m but less than 25m horizontally from the LiDAR center), due to the varying sizes of the vehicle segments crossing the boundary, the recognition rate is relatively weaker, reaching over 85%. The average non-static object recognition rate is approximately 92%.

[0227] A test vehicle was used to simulate roadside parking behavior to test the method in step four. Every 5 frames of data, all non-static regions A in the detection frame were recorded.ns As a temporary static area A st Used for secondary matching. The recorded features include the horizontal projection size of each sub-region of non-static region A, the location of the centroid, and the highest value of the Z-axis.

[0228] After all non-static regions of the frame to be identified are extracted in step three, its sub-regions are compared sequentially with the sub-regions recorded in the temporary static region. The features and order of comparison are as follows:

[0229] ①Non-static region A ns and temporary static area A st The sub-regions in the two are sorted according to the scanning distance of their centroids. The centroid positions of the sub-regions in the two are compared in turn. If a sub-region A exists... ns -i、A st -j If the centroid distance between the two is no greater than 0.3 meters and there are no other matching objects within 1 meter, then proceed to the next step; otherwise, mark all non-static sub-regions that do not meet the conditions as dynamic regions.

[0230] ②A ns -i、A st -j Comparing the horizontal projection sizes of the two regions, if the two regions are located in the edge region and the rate of change is within 15%, or the two regions are located in the inner region and the rate of change is less than 5%, then proceed to the next step; otherwise, mark all non-static sub-regions that do not meet the conditions as dynamic regions.

[0231] ③A ns -i、A st Comparing the highest voxel Z-axis values ​​of the two regions, if the two regions are located at the edge and the rate of change is within 15%, or if the two regions are located in the inner region and the rate of change is less than 5%, then A is considered... ns -i can be considered as A st -j, about to A ns -i、A st -j represents the same object as a short-term static object; otherwise, all non-static sub-regions that do not meet the conditions are still marked as dynamic regions.

[0232] ④ Since dynamic objects are also recorded in the temporary static region each time it is recorded, by comparing the previous and next frames, if a sub-region A exists in the temporary static region... st -j means that if no sub-region in the non-static region of the subsequent frame data can match it, then sub-region A is considered to be... st -j is not a short-term static object, so it can be removed from the temporary static region to reduce the amount of comparison in the temporary static region later.

[0233] Finally, after the above comparison is completed, the set of all dynamic regions is the dynamic object D. dAfter removing dynamic objects, the remaining part of the temporary static region is the short-term static object D. st This is also the temporary static area A for the next comparison. st .

[0234] In this case, the dynamic object recognition rate in the internal area (the area less than 23m away from the center of the lidar) can reach over 93%. The recognition rate in the edge area (the area more than 23m but less than 25m away from the center of the lidar) is relatively weaker due to the varying sizes of the cross-boundary vehicle segments, but can still reach over 80%. The average dynamic object recognition rate is around 88%.

Claims

1. A method for fast identification of moving and static objects and point cloud segmentation based on roadside sensing units, comprising the following steps: (I) Data Acquisition For the vehicle-road cooperative environment, a roadside lidar perception scenario is built, and raw point cloud data D0 is collected for preprocessing; (II) Pretreatment 2.1) First, establish the boundary equation system E b Invalid data is removed from the original data D0 to obtain valid point cloud data; 2.2) Sample the effective point cloud data at equal intervals to establish a linear relationship between point cloud density and scanning distance; 2.3) Identify and separate static and non-static objects in each frame of valid point cloud data to establish a static point cloud background B; 2.4) Perform a voxelization operation on the static point cloud background to obtain the voxelized static point cloud background B. V ; (III) Static Object Recognition 3.1) Separate the recognition area from the effective point cloud data using the recognition trigger distance threshold; 3.2) The recognition region is voxelized into a 3D matrix, and the 3D matrix is ​​compared with the voxelized static point cloud background B using the sliding window method. V Compare the data; if a continuous area within the sliding window matches B... V If the rate of change between regions at the same location is greater than the static region determination threshold, the continuous region is marked as a non-static region; otherwise, it is marked as a static region. The union of all static regions is the static object in the valid point cloud data. (iv) Non-static object recognition 4.1) In the above static object recognition process, according to a certain frequency, all static regions in a certain frame of voxelized point cloud data are recorded as temporary static regions A. st ; 4.2) When identifying subsequent frames, the non-static region of the voxelized point cloud data to be identified is compared with the temporary static region A. st The matching is performed. If the rate of change of the size and position features of the two regions is less than the threshold for determining short-term static objects, the non-static region of the voxelized point cloud data to be identified can be considered as a short-term static object; otherwise, it is considered as a dynamic object. 4.3) Repeat steps 4.1) and 4.2) until the entire valid point cloud data has been traversed.

2. The method as described in claim 1, characterized in that, The collected point cloud data has comprehensive and unobstructed coverage and includes static objects, including road surfaces and their ancillary facilities, buildings and roadside greenery; it also includes a sufficient number of non-static objects, including pedestrians, non-motorized vehicles and / or vehicles; the non-static objects are manually identified and the total number of non-static objects is no less than 300.

3. The method as described in claim 1, characterized in that, The method for establishing the linear relationship between point cloud density and scanning distance: With a sampling interval of 0.5 meters, sample points are collected at equal intervals along each scanning ring from the inside out. The number of points within a 10cm radius centered on each sampling point is recorded as the point cloud density. The average point cloud density of each ring is then calculated to establish the linear relationship between point cloud density and scanning distance.

4. The method as described in claim 1, characterized in that, The static point cloud background B is established according to the following method: 2.3.1) Use point cloud target detection algorithm to detect non-static objects to establish the distribution curve of detection confidence and scanning distance, and select the scanning distance range where the confidence is higher than the threshold as the recognition trigger distance threshold; 2.3.2) All static objects are cropped using a recognition trigger distance threshold. The cropped multi-frame recognition static objects are then overlaid and appropriately sampled to construct a static point cloud background B. V .

5. The method as described in claim 1, characterized in that, The method for establishing the identification trigger distance threshold is as follows: Extract various non-static objects, sample the point cloud of each non-static object at a certain ratio and calculate its average point cloud density to determine its corresponding scanning distance L; Each target is input into a point cloud target detection algorithm for detection, and the detection results and detection confidence P are obtained. Plot the distribution curve of distance L versus confidence level P using the following formula: in, These represent the total number of dynamic targets at distances i and j from the origin, respectively. Let P represent the total number of non-static targets within the range i and j with a detection confidence level greater than 75%; with a sampling interval of 0.5 meters as the upper and lower limits, select appropriate values ​​of i and j to make P... ij If the value of i is greater than 75% and maximizes the difference between i and j, then the boundary equation for the transformation of the values ​​of i and j is the identification trigger distance threshold DT.

6. The method as described in claim 5, characterized in that, The voxelized static point cloud background B V The construction method is as follows: Based on manual extraction methods, static objects and non-static objects in each frame of valid point cloud data are separated. For static objects, the recognition region is separated using the recognition trigger distance threshold DT. Then, starting from the origin of the point cloud coordinate system, it is divided into n statistical spaces with gradually increasing spacing. Starting from the initial frame, the valid point cloud data of the next frame are superimposed sequentially. During each superposition, the point cloud density of each statistical space is detected. If it is greater than the threshold α, the valid point cloud data in that space is randomly sampled to maintain its density. Finally, a static point cloud background B with a suitable point cloud density is obtained; finally, voxelization is performed on the static point cloud background to obtain the voxelized static point cloud background B. V .

7. The method as described in claim 1, characterized in that, The method for marking the static and non-static regions is as follows: The point cloud data frame to be identified is first processed using the boundary equation system E. d The data is first identified by clipping the trigger distance threshold DT, and then voxelized at a density of 1 / 20 to 1 / 40 of the statistical space size to obtain the voxelized point cloud data D to be identified. v ''; From the outside in, a sliding window and background difference method are used to match a continuous region in the voxelized point cloud data to be identified with a region at the same location in the voxelized static point cloud background; If the difference in the rate of change between the two exceeds the threshold for determining a static region, it is marked as a non-static region A. ns Otherwise, mark it as static region A. s ; When a non-static region is detected again, if it is similar to the known non-static region A... ns If -1 is adjacent, it is also marked as A. ns -1, otherwise mark as non-static region A ns -2, and so on.

8. The method as described in claim 7, characterized in that, For the edge region, an annular backup identification area with a width of 1.5m is extended outward, and then divided into several sub-regions according to the scanning angle; if a non-static region is detected in the edge region, the annular backup identification area within the angle range of the non-static region is included in the non-static region according to the scanning angle of the non-static region.

9. The method as described in claim 1, characterized in that, During the detection of short-term static objects, non-static region A is recorded at fixed frequency intervals. ns As a temporary static area A st Used for secondary matching; after extracting all non-static regions from the voxelized point cloud data to be identified using a sliding window and background subtraction method, each sub-region belonging to the non-static region to be identified is sequentially compared with each sub-region in the temporary static region. If there are two sub-regions A between them... ns -i、A st -j makes the difference in position and shape features between the two objects less than the short-term static object determination threshold, then the two objects are considered to be the same object and marked as a short-term static object; otherwise, they are marked as dynamic objects. For short-term static objects, a lower data transmission frequency is set.

10. The method as described in claim 1, characterized in that, The short-term static objects include temporary parking spaces.