FPGA-based laser radar point cloud online sample self-generation method
By realizing the Euclidean distance segmentation and geometric feature merging of lidar point clouds on FPGA, and combining Bayesian cross-validation and weak semantic features for labeling and error sample filtering, the problems of low sample generation efficiency and poor adaptability of traditional point cloud training are solved, and efficient and automatic sample generation is achieved.
Patent Information
- Application Number
- CN202510508351.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional lidar point cloud training sample generation relies on manual annotation, is inefficient and cannot adapt to complex large scenarios, and lacks a robust unsupervised self-generating method.
Using FPGA-based pipeline design, point cloud data is segmented and merged through Euro-style distance and geometric features, combined with Bayesian cross-validation and weak semantic features for marking and error sample filtering, realizing full process automation.
The efficiency and quality of sample generation are improved, the cost of manual labeling is reduced, and the adaptation to complex large scenarios is achieved. The automatic sample extraction accuracy reaches more than 94.7%.
Smart Images

Figure CN120032209A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automatic generation of laser radar point cloud training samples, and in particular relates to an FPGA-based laser radar point cloud online sample self-generation method. Background Art
[0002] Autonomous driving technology is driving profound changes in various industries. LiDAR point cloud technology plays a vital supporting role in the development of autonomous driving. By capturing high-precision three-dimensional data of the vehicle's surroundings, LiDAR can build an accurate environmental model in real time, providing key perception capabilities for autonomous driving systems. However, the amount of point cloud data acquired in real time is very large. In the traditional point cloud supervised classification process, training samples of each category need to be selected in advance. This process often requires manual work. In a small scene study area, this method may be completed quickly. However, in large scenes with complex city levels, it is extremely laborious to label training data sets that rely on extensive and dense annotations, which restricts a series of subsequent applications such as recognition and perception.
[0003] The current status of research on automatic generation of LiDAR point cloud training samples mainly focuses on how to efficiently convert the collected raw point cloud data into high-quality training sets for the training of supervised learning models. Research in this field mainly covers data annotation automation, data enhancement, multi-view data generation, and training sample generation based on simulated data.
[0004] Traditionally, the annotation of point cloud data relies on manual work, which is time-consuming and susceptible to human factors. To solve this problem, researchers have developed automated annotation tools and semi-automated annotation methods. These tools usually combine deep learning algorithms and rule engines to improve annotation efficiency by initially identifying objects and annotating them, and then using a small amount of manual correction. The semantic segmentation network is used to preprocess the point cloud data to automatically generate some annotations. Data enhancement technology generates more training samples by transforming existing point cloud samples (such as rotation, translation, noise addition, etc.), thereby improving the robustness of the model. Researchers have also proposed multi-resolution data enhancement methods based on point cloud data. These methods can generate training data of different perspectives and distances while maintaining the structural information of the point cloud, thereby enhancing the model's adaptability to different scenes. In order to solve the problem of sparsity of point cloud data, researchers expand the training set by generating point cloud data from different perspectives. Specific methods include generating point cloud samples observed from multiple angles through simulation of virtual cameras or sensors.
[0005] In addition, the latest research shows the use of generative adversarial networks (GANs) to generate realistic point cloud training data. Through this method, the model can learn the underlying structure of point cloud data and generate diverse training samples, but this method first requires training samples to generate, and lacks a robust mathematical and physical process.
[0006] In summary, the research on the automatic generation technology of LiDAR point cloud training samples is gradually improving the efficiency and quality of data set generation, providing richer sample support for the training of supervised learning models. The above research more or less requires sample annotation, does not solve the problem of the origin of the "No. 0" sample, and lacks a robust unsupervised self-generation method for samples. Summary of the invention
[0007] In view of this, the present invention aims to provide an FPGA-based lidar point cloud online sample self-generation method. Based on the FPGA pipeline design, the lidar point cloud data is segmented by using Euclidean distance and geometric features, and low-quality samples are filtered out by using point cloud geometric weak semantic features and Bayesian cross-validation. Through hardware acceleration and algorithm optimization, the full process automation of point cloud segmentation, merging, labeling and error sample filtering is realized, which effectively improves the processing efficiency and data quality.
[0008] To achieve the above object, the technical solution created by the present invention is implemented as follows: The present invention provides an FPGA-based laser radar point cloud online sample self-generation method, comprising: S1: performing Euclidean distance segmentation on input point cloud data based on FPGA pipeline design, and generating an initial segmentation point set by iteratively merging the minimum distance point set; S2: Calculate the geometric features of the initial segmentation point set, where the geometric features include at least one of an adjacency matrix, normal dispersion, planarity, and horizontal angle, and merge the initial segmentation point set based on an online equivalence strategy; S3: Initialize and label the merged point set based on weak semantic knowledge to generate labeled training samples, where the weak semantic knowledge includes at least one of height feature, planarity feature, spatial projection feature, and cylindrical model feature; S4: Construct a representative feature set of the target, combine Bayesian cross-validation and linear SVM classifier to calculate the confidence of the training sample labeling, and filter out the training samples with confidence less than the set threshold as error samples.
[0009] Preferably, in S1, the point cloud data consists of a floating-point array of the spatial position and color of each point. The point cloud data is converted into a binary file through a floating-point binary algorithm based on C language, read using a FIFO core, and stored in DDR4.
[0010] Preferably, in S2, the online equivalent strategy dynamically merges the initial segmentation point set including: Each initial segmentation point set is initialized as an independent equivalence class, and the geometric features are used as equivalence relations to merge the initial segmentation point sets.
[0011] Preferably, during the merging process of the initial segmentation point sets, the point sets are merged by managing equivalence classes through a tree structure.
[0012] Preferably, the adjacency matrix is used to represent the positional relationship between different initial segmentation point sets, and the distance between any two initial segmentation point sets is: ; in, Indicates i An initial set of segmentation points, Indicates j An initial set of segmentation points, express and The distance between express Middle k Points, express Middle l points; Normal dispersion is used to represent the average dispersion degree of the normal angles between adjacent points in each initial segmentation point set. The average normal dispersion of each initial segmentation point set is: ; in, Indicates i The average normal dispersion of the initial segmentation point set, r is the number of points in the initial segmentation point set, and Represents the normal vector of any two adjacent points; The calculation formula for planarity is: ; ; in, Indicates i The planarity of the initial partition point set, Indicates i Any point in the initial segmentation point set Data items used to judge Is it the point a, b, c in the random sampling consistent fitting plane? It is the plane parameters of the random sampling consistent fitting plane. , , Represent any point The x, y, and z coordinates of ε is the tolerance parameter; The calculation formula of the horizontal angle is: ; in, Indicates i The initial segmentation point set and the horizontal direction The angle of Represents any point The normal vector of .
[0013] Preferably, in S3, the method for initializing the mark is: The merged point set is initialized and labeled through a decision tree model.
[0014] Preferably, construct a representative feature set of the target for: ; in, Represents the normal vector discreteness of the merged point set, represents the planarity of the merged point set, represents the normal azimuth of the merged point set, Represents the spatial projection features of the merged point set, represents the local maximum height of the merged point set, represents the local minimum height of the merged point set, Represents the cylindrical feature tolerance parameter of the merged point set, Represents the angle between the central axis of the cylinder of the merged point set and the zenith direction.
[0015] Preferably, the merged point set is initialized and labeled based on height features, planarity features, spatial projection features and cylindrical model features, and is divided into four categories of labeled training samples.
[0016] Preferably, the process of calculating the confidence of the training sample is: a. Perform cross-validation for each type of labeled training samples and perform multiple iterations. In a single iteration, each type of labeled training samples is divided into a training set and a test set; b. Select a set number of training samples from any one of the labeled training samples as positive samples, and select a set number of training samples from the remaining three types of labeled training samples as negative samples; use the target representative feature set corresponding to the positive and negative samples Train a linear SVM classifier; c. Use the SVM classifier to classify the samples in the test set into positive samples and negative samples; d. Repeat the process from a to c and record the number of times each sample is classified as a positive sample; e. For a training sample that is marked as a positive sample n times after N iterations, the probability that the initial label of the training sample is correct is obtained by transforming the Bayesian formula: ; Where N represents the total number of iterations, n represents the number of times the sample is marked as a positive sample in N iterations, It indicates the probability that a training sample is labeled as a positive sample n times after N iterations, and the initial label of the sample is correct. represents the prior probability that the initial label of a training sample is correct, represents the prior probability that the initial label of a training sample is wrong, It indicates the probability that a training sample is labeled as a positive sample n times in N iterations when the initial label of the training sample is correct; It represents the probability that a training sample will be labeled as a positive sample n times in N iterations when the initial label of the training sample is wrong. represents the number of combinations marked as positive samples from N iterations, represents the probability of the linear SVM classifier being correct each time, ; Determine whether the probability of the initial label of the training sample being correct is less than the set threshold, and filter out the training samples whose correct probability is less than the set threshold as error samples.
[0017] Compared with the prior art, the invention can achieve the following beneficial effects: Through the FPGA pipeline design, the present invention realizes the full process automation of point cloud segmentation, merging, marking and error sample filtering in the point cloud data processing process, and through the online sample self-generation model, sample generation and processing can be performed while data is collected, reducing the delay of data transmission and storage, further improving the efficiency of the overall system, and solving the problem that traditional sample generation relies on labeling, manual labeling of "No. 0" sample is required, and error samples cannot be filtered out.
[0018] The present invention utilizes the Bayesian cross-validation method to filter out low-quality samples through self-learning, thereby ensuring that the generated samples have high quality, effectively reducing noise and incorrectly labeled samples, and improving the accuracy of samples, so that the automatic sample extraction accuracy reaches more than 94.7%, and the processing efficiency can reach 1.7GB / min. In addition, in the sample classification process, a variety of geometric features (such as height, planarity, spatial projection features, etc.) are combined to more accurately identify and classify different objects, such as cars, trees, pedestrians, buildings, etc., to further improve the accuracy of sample generation. In addition, through the online equivalent class point set merging strategy, various situations in point cloud data, such as uneven point density and data missing, can be flexibly handled, so that the method of the present invention can adapt to different data distribution and scene requirements, and overcome the problems that traditional methods are only applicable to small scene study areas and are difficult to classify in large and complex scenes at the urban level. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings constituting part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation on the present invention. In the drawings: Figure 1 It is a flowchart of online sample self-generation of laser radar point cloud based on FPGA according to an embodiment of the present invention; Figure 2 It is a schematic diagram of automatically generating training samples and filtering out incorrectly labeled samples by taking a ground-based point cloud scene as an example according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical scheme and advantages of the invention clearer, the invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and do not constitute a limitation to the invention. Similar components in different embodiments use associated similar component numbers. In the following embodiments, many detailed descriptions are to enable the invention to be better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other components, materials, and methods. In some cases, some operations related to the invention are not shown or described in the specification, in order to avoid the core part of the invention being overwhelmed by too much description, and for those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations according to the description in the specification and the general technical knowledge in the art.
[0021] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to form various implementation methods. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a necessary sequence, unless otherwise specified that a certain sequence must be followed.
[0022] In the description of the invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the invention, unless otherwise specified, the meaning of "multiple" is two or more.
[0023] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the invention can be understood according to specific circumstances.
[0024] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0025] Automatically generate online point cloud samples for urban scene point cloud data. Due to the uneven density of urban scene point clouds, occlusion and noise, traditional segmentation algorithms are inefficient and have poor accuracy. Existing deep learning models require a large amount of labeled data, such as PointNet, which has high labor costs and is prone to introducing incorrect labels. Sample generation relies on a large amount of labeled data, so it consumes a lot of hardware resources, and rule-based geometric segmentation methods are difficult to adapt to complex scenes. Therefore, please refer to Figure 1In one embodiment of the present invention, a method for self-generating online samples of lidar point cloud based on FPGA is provided, which mainly solves the problem that traditional urban scene point cloud data relies on manual annotation, the amount of generated samples is limited, the efficiency of generating samples is low, and it cannot meet the application of complex scenes.
[0026] Specifically, the process of implementing the full process automation of point cloud segmentation, merging, labeling and error sample filtering in the embodiment of the present invention is as follows: S1: First, the hardware is optimized and updated, and FPGA (field programmable gate array) is used to generate training samples for the laser radar point cloud. In the embodiment of the present invention, the radar that collects the point cloud file is the ground-based laser radar RIEGL vz-1000, and the original point cloud file format collected by the laser radar is .las / .pcd, which mainly consists of the spatial position and color information of each point, usually represented by a floating-point array of spatial position and color information, specifically represented as , xyz represents the spatial position coordinates, and rgb represents the color information. When generating training samples for point cloud data, the floating point binary algorithm of C language is first used to convert the point cloud data into a binary file to reduce the data storage space and increase the reading speed. The data is read through the FIFO core and stored in the DDR4 memory. The FIFO core can ensure the orderly reading of data and avoid data loss or confusion.
[0027] For the training sample generation design process, in the embodiment of the present invention, Verilog HDL language is used to perform online sample self-generation algorithm design in Xilinx ISE software. For point clouds of large-scale urban scenes, data is usually missing due to occlusion and point density is unevenly distributed, and the generated point set is an incomplete object.
[0028] The point cloud data is segmented through FPGA. The segmentation principle is to divide the points according to spatial, geometric and texture features so that the point sets in the same segmentation have similar structural features. The effective segmentation of point clouds is related to its specific application scenarios. For example, in the field of intelligent driving, effective segmentation is required according to actual needs.
[0029] For a scene P where the laser radar obtains point cloud data, assuming that it contains m points, it needs to be segmented into n point sets according to the attribute characteristics, and points with similar characteristics are divided into the same set. The embodiment of the present invention performs the segmentation process based on the Euclidean distance on the input point cloud data based on the FPGA pipeline design: Initially, each point is regarded as an independent point set, and the Euclidean distance between any two point sets is calculated. The minimum Euclidean distance is usually used as the segmentation metric, that is, the two point sets with the smallest Euclidean distance are and Merge into a new point set, and this method will generate a series of new point sets. Iterate the above process, recalculate the Euclidean distance between the new point sets, and then merge the point sets until the distance between any two new point sets obtained by merging is greater than the preset segmentation threshold. At this time, the iteration is completed and the initial segmentation point set is generated, which is reflected in the image as the pixel block segmentation. It should be noted that the segmentation threshold here is set according to the application scenario and there is no fixed value.
[0030] In addition, outliers in the scene can be further identified. Outliers are usually located away from most points in the scene and may be external noise or isolated objects in the actual scene. Specifically, before merging the point set, the outliers can be determined and counted by calculating the average distance between each point and multiple neighboring points.
[0031] S2: After obtaining the initial segmentation point set, the segmented bodies formed after segmentation (i.e., the pixel presentation of the initial segmentation point set) will appear partially fragmented due to interference from factors such as uneven point cloud density or data loss due to occlusion. In order to generate a more complete and accurate independent point cloud object, these segmented bodies need to be merged. In an embodiment of the present invention, the geometric features of the initial segmentation point set obtained by S1 are still calculated by FPGA, and the efficient merging of the initial segmentation point set is achieved by calculating the geometric features of the point set and combining the online equivalence strategy. Specifically, in an embodiment of the present invention, the geometric features used include adjacency matrix, normal discreteness, planarity, and horizontal angle. Based on the above four geometric features as geometric semantic equivalence relations, the initial segmentation point set is merged twice. It should be noted that in different application scenarios, one or more of the above four geometric features may be selected.
[0032] The four methods for calculating the geometric characteristics of a point set are as follows: For the adjacency matrix of a point set , the adjacency matrix is used to represent the positional relationship between different initial segmentation point sets, then the distance between any two initial segmentation point sets can be expressed as: ; in, Indicates i An initial set of segmentation points, Indicates j An initial set of segmentation points, express and The distance between express Middle k Points, express Middle lpoints. When calculating the shortest distance between different initial segmentation point sets, it is assumed that there are n initial segmentation point sets. In order to reduce the time complexity and avoid calculating all possible point pairs, the kd-tree neighbor search can be used to find the neighboring set of any initial segmentation point set, thereby reducing the number of point pairs that need to be calculated and effectively reducing the time complexity. Then the problem of finding the shortest distance between two initial segmentation point sets is converted into finding the shortest distance between two convex hulls. The convex hull is the smallest convex polygon or convex polyhedron containing the point set. The convex hull of a point set refers to the smallest convex polygon or convex polyhedron boundary that contains all points. When calculating the shortest distance between point sets, the points in the point set are all located inside or on the boundary of the convex hull, so the shortest distance between point sets must appear between the vertices of the convex hull. By calculating the shortest distance between convex hulls, the shortest distance between the initial segmentation point sets can be approximated. This method greatly reduces the amount of calculation.
[0033] For the normal dispersion , which is used to represent the average discrete degree of the normal angle between adjacent points in each initial segmentation point set. In the planar point set merge, the normal of the point set with good planarity is usually more consistent, and the average discrete degree is lower. However, the trees or people detected during the driving process of the car are often curved surfaces with higher average discrete degree, and the structure of the crown part is very complex and the normal angle diverges.
[0034] The average normal dispersion of each initial segmentation point set is: ; in, Indicates i The average normal dispersion of the initial segmentation point set, r is the number of points in the initial segmentation point set, and Represents the normal vector of any two adjacent points; For merging and the next step of building model search, planarity is a very important feature. The planarity metric is used to determine whether the point set is in the same plane. Therefore, it is necessary to measure the planarity of the initial segmentation point set. When measuring planarity, the random sampling consensus algorithm can be used to obtain the plane parameters in the initial segmentation point set and calculate the probability of points that meet the plane equation. By calculating the probability that the points in the initial segmentation point set fall within the plane, the planarity of the initial segmentation point set can be measured, and then it can be determined whether the object detected by the radar has a high degree of planarity.
[0035] The calculation formula for planarity is: ; ; in, Indicates iThe planarity of the initial partition point set, Indicates i Any point in the initial segmentation point set Data items used to judge Is it the point a, b, c in the random sampling consistent fitting plane? It is the plane parameters of the random sampling consistent fitting plane. , , Represent any point The x, y, and z coordinates of ε is the tolerance parameter; The horizontal angle is used to determine whether two initial segmentation point sets are similar. Therefore, it is necessary to calculate the angle between each initial segmentation point set and the horizontal direction. .
[0036] The angle between the initial segmentation point set and the horizontal direction is calculated as: ; in, Indicates i The initial segmentation point set and the horizontal direction The angle of Represents any point The normal vector of .
[0037] The embodiment of the present invention dynamically merges the initial segmentation point sets through an online equivalence strategy. Specifically, each initial segmentation point set is regarded as an independent equivalence class, and the above four geometric features are used as the geometric semantic equivalence relationship for judging the point set merge. The geometric features of each initial segmentation point set are judged. If the geometric features of two initial segmentation point sets meet the merging conditions, the two point sets are merged. The labeled input merge operation In essence, the merging process is to determine whether the geometric features of the two initial segmentation point sets are similar, and if the degree of similarity meets the set merging requirements, the two initial segmentation point sets are merged. This process designs the query of equivalence classes and the point set merging operations. The embodiment of the present invention manages equivalence classes through a tree structure to merge point sets. Specifically, three domains are allocated to each point set (or element), namely the root root, the parent node parent and the child node child, where root[i] is of bool type, true means that the current element is the root, otherwise it is not; parent[i] represents the parent node of the current element. If the current element is the root node, parent[i] represents the number of all its children and is stored in child[i]. All root[i] are initialized to true, parent[i] is initialized to 1, and child[i] is initialized to an empty set. In the point set merging process, according to the quantity principle, the nodes with fewer points are merged into the nodes with more points, and their child nodes, that is, the member labels of each class after the merger, are saved.
[0038] S3: After S1 performs point set division and S2 performs point set merging, a complete point set is obtained. At this time, a decision tree model is used to complete the initialization marking of segmented bodies of different categories (cars, trees, pedestrians, buildings). The principle and basis of initialization marking is weak semantic knowledge. The merged point set is initialized and marked based on weak semantic knowledge to generate marked training samples. The weak semantic knowledge used in the embodiment of the present invention includes height features, planarity features, spatial projection features, and cylindrical model features. According to application requirements, only one or more of the above-mentioned weak semantic knowledge can be used. The specific method of using weak semantic knowledge for identification and judgment is as follows: The height features of the segmented body reflected by the merged point set include the maximum height of all points in the point set and the minimum height of all points. The segmented body reflected by the point set can be judged according to the determined height range. For example, buildings usually have a higher maximum height and a larger height range. Trees also have a higher maximum height and a larger height range, but their maximum height is usually smaller than that of buildings. Cars and pedestrians usually have smaller maximum heights and height ranges, and have clear intervals.
[0039] The planarity features of the segmented bodies reflected by the merged point set are used to measure whether the point sets are in the same plane. They can well distinguish some segmented bodies with good planarity, such as building facades, roofs, and tree crowns.
[0040] The spatial projection characteristics of the segmented body reflected by the merged point set are mainly used to judge trees. The canopy projection area ratio is the area of the tree projected to the ground divided by the trunk diameter at breast height. For trees, this ratio is very large (10~40). For other ground objects, the ratio of the vertical projection area to the bottom projection area is close to 1.
[0041] For the cylindrical model features of the segmented body reflected by the merged point set, for a car, a cylindrical model can be used to judge the terrain feature according to its geometric shape. First, the parameters of the fitted cylindrical model are calculated by the RANSAC model. Generally speaking, the cross-sectional radius of the car can be considered to be between 0.3 and 1m. The degree to which the point set satisfies the cylinder can be calculated, that is, the cylindrical feature tolerance parameter. , and the cylindrical axis direction parameters are obtained according to Direction to the zenith , find the angle between the central axis of the cylinder and the zenith direction By judging still It can distinguish people and cars very well. and Represent the cylindrical axis direction parameters of cars and people respectively. The above distinctive features are regarded as weak prior knowledge and marked using the decision tree model. The merged point set can be marked as four types of labeled training samples through the above weak semantic knowledge.
[0042] S4: After the point set is marked by FPGA, the training samples need to be further verified to remove the wrong samples. Specifically: First, seven types of features are extracted from each merged point set to form a feature vector, and a target representative feature set is constructed. The embodiment of the present invention is applied to intelligent driving, so it can also be called a representative feature set of ground objects. for: ; in, Represents the normal vector discreteness of the merged point set, represents the planarity of the merged point set, represents the normal azimuth of the merged point set, Represents the spatial projection features of the merged point set, represents the local maximum height of the merged point set, represents the local minimum height of the merged point set, Represents the cylindrical feature tolerance parameter of the merged point set, Represents the angle between the central axis of the cylinder of the merged point set and the zenith direction.
[0043] Since each dimension of the feature set represents a different feature meaning, and the dimension and value range of each feature are quite different, normalization is required. Train a linear SVM classifier.
[0044] In order to eliminate erroneous samples in the training samples, multiple rounds of cross-validation and probability modeling are used to quantify the confidence of the initial labels of the training samples, filter out low-confidence samples, and improve the quality of training data. Specifically, the process of calculating and evaluating the confidence of the training samples is as follows: a. Perform cross validation for each type of labeled training samples and perform multiple iterations, setting the total number of iterations to N. During a single iteration, each type of labeled training samples is divided into a training set and a test set.
[0045] b. To achieve subsequent training, positive samples and negative samples need to be constructed. In the training set, a set number of training samples are selected from any one category of labeled training samples as positive samples, and a set number of training samples are selected from the remaining three categories of labeled training samples as negative samples. The number of selected positive and negative samples is the same to maintain category balance. Use the target representative feature set corresponding to the positive and negative samples Train a linear SVM classifier; c. Use the SVM classifier to classify the samples in the test set into positive and negative samples. The classification performance of the linear SVM classifier, that is, the classification reliability, is significantly better than random guessing. If a sample from the test set is classified as a positive sample, then the probability of it being correctly labeled during the initial labeling will be greater than 50%. Then in each round, samples selected from a certain category are randomly divided into the test set and the training set for cross-validation. If a sample is classified as a positive sample many times, then the probability that it is correctly labeled initially is higher.
[0046] d. Based on the above principle, the embodiment of the present invention can continue to perform iterative judgment, repeat the process of iterating from a to c, and record the number of times each sample is classified as a positive sample.
[0047] e. Assuming that the classification accuracy of the linear SVM classifier is greater than 0.5, the number of times a sample is classified as a positive sample should obey the binomial distribution, that is: When the initialization mark is correct, ; When the initialization mark is wrong, .
[0048] For a training sample, if n times it is labeled as a positive sample after N iterations, the probability that the initial label of the training sample is correct can be obtained by transforming the Bayesian formula: ; Where N represents the total number of iterations, n represents the number of times the sample is marked as a positive sample in N iterations, It indicates the probability that a training sample is labeled as a positive sample n times after N iterations, and the initial label of the sample is correct. represents the prior probability that the initial label of a training sample is correct, represents the prior probability that the initial label of a training sample is wrong, It indicates the probability that a training sample is labeled as a positive sample n times in N iterations when the initial label of the training sample is correct; It represents the probability that a training sample will be labeled as a positive sample n times in N iterations when the initial label of the training sample is wrong. represents the number of combinations marked as positive samples from N iterations, represents the probability of the linear SVM classifier being correct each time, ; Determine whether the probability of the initial label of the training sample being correct is less than a set threshold, and filter out the training samples whose probability of being correct is less than the set threshold as error samples. In this embodiment of the present invention, the threshold is set to 85%. <85%, it is determined to be a low-confidence sample, that is, the training sample is an erroneous sample and is filtered out. The set threshold can be adjusted dynamically. When the noise is high, the set threshold can be appropriately lowered. The present invention fully automates the point cloud segmentation, merging, labeling and erroneous sample filtering processes through FPGA, solving the problem of traditional reliance on manual calibration and the drawback of having to manually label "No. 0" samples. The fully automated design greatly reduces the sample generation cost and improves the sample generation efficiency, which can meet the needs of large-scale complex scenes at the city level for a large number of training samples.
[0049] The method of the present invention has been tested in practice. Figure 2 As shown, for the point cloud data collected by the ground-based laser radar RIEGL vz-1000, the method of the present invention realizes the process of automatically extracting training samples in the ground-based scene and filtering out the wrongly marked training samples. Figure 2 The scene point cloud shown in (a) is divided into 260 point sets in step S1, and the point sets are merged in step S2 to obtain Figure 2 In (b), the discrete point sets that originally belonged to the same independent point set were merged to obtain 182 independent point set objects. Figure 2 (b) in the figure is used for initialization marking, starting from Figure 2In (c), we can see that the point set in the rectangle is incorrectly labeled using weak prior knowledge. Among them, the part of the building facade close to the tree is classified as a tree, and the car in the lower left corner is incorrectly classified as low vegetation. Then, in step S4, low confidence value samples are filtered out, and cross-validation learning is performed based on linear support vector machine. In the actual experiment, training samples with a probability (i.e., confidence) of 85% or less of the initial label being correct will be filtered out. After filtering, Figure 2 (d) and (f) are magnified images of (c) and (e), respectively. It can be seen that the misclassified samples are filtered out. After manual annotation verification, the automatic sample extraction accuracy of the proposed method reaches 94.7%, and the processing efficiency can reach 1.7 GB / min. The generated samples have high quality, and the overall generation efficiency is also greatly improved.
[0050] In short, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.
[0051] The systems, devices, modules or units described in one or more of the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0052] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0053] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0054] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A method for self-generating online samples of laser radar point cloud based on FPGA, characterized in that: include: S1: Based on FPGA pipeline design, the input point cloud data is segmented by Euclidean distance, and the initial segmentation point set is generated by iteratively merging the minimum distance point set; S2: Calculate the geometric features of the initial segmentation point set, where the geometric features include at least one of an adjacency matrix, normal dispersion, planarity, and horizontal angle, and merge the initial segmentation point set based on an online equivalence strategy; S3: Initializing and marking the merged point set based on weak semantic knowledge to generate marked training samples, wherein the weak semantic knowledge includes at least one of a height feature, a planarity feature, a spatial projection feature, and a cylindrical model feature; S4: Construct a representative feature set of the target, combine Bayesian cross-validation and linear SVM classifier to calculate the confidence of the training sample labeling, and filter out the training samples with confidence less than the set threshold as error samples.
2. The FPGA-based laser radar point cloud online sample self-generation method according to claim 1 is characterized in that: In S1, the point cloud data consists of a floating-point array of the spatial position and color of each point. The point cloud data is converted into a binary file through a floating-point binary algorithm based on C language, read using a FIFO core, and stored in DDR4.
3. The FPGA-based laser radar point cloud online sample self-generation method according to claim 1, characterized in that: In S2, the online equivalent strategy dynamically merges the initial segmentation point set, including: Each of the initial segmentation point sets is initialized as an independent equivalence class, and the initial segmentation point sets are merged using the geometric features as equivalence relations.
4. The FPGA-based laser radar point cloud online sample self-generation method according to claim 3 is characterized in that: During the merging process of the initial segmentation point set, the point set is merged by managing the equivalence classes through a tree structure.
5. The FPGA-based online laser radar point cloud sample self-generation method according to claim 1, characterized in that: The adjacency matrix is used to represent the positional relationship between different initial segmentation point sets. The distance between any two initial segmentation point sets is: ; in, Indicates i The initial segmentation point set, Indicates j The initial segmentation point set, express and The distance between express Middle k Points, express Middle l points; The normal dispersion is used to represent the average dispersion degree of the normal angles between adjacent points in each of the initial segmentation point sets. The average normal dispersion of each of the initial segmentation point sets is: ; in, Indicates i The average normal dispersion of the initial segmentation point set, r is the number of points in the initial segmentation point set, and Represents the normal vector of any two adjacent points; The calculation formula for the planarity is: ; ; in, Indicates i The planarity of the initial segmentation point set, Indicates i Any point in the initial segmentation point set Data items used to judge Is it the point a, b, c in the random sampling consistent fitting plane? It is the plane parameters of the random sampling consistent fitting plane. , , Represent any point The x, y, and z coordinates of ε is the tolerance parameter; The calculation formula of the horizontal angle is: ; in, Indicates i The initial segmentation point set and the horizontal direction The angle of Represents any point The normal vector of .
6. The FPGA-based laser radar point cloud online sample self-generation method according to claim 1, characterized in that: In S3, the method for initializing the mark is: The merged point set is initialized and labeled through a decision tree model.
7. The FPGA-based laser radar point cloud online sample self-generation method according to claim 1, characterized in that: Constructing a representative feature set of the target for: ; in, Represents the normal vector discreteness of the merged point set, represents the planarity of the merged point set, represents the normal azimuth of the merged point set, Represents the spatial projection features of the merged point set, represents the local maximum height of the merged point set, represents the local minimum height of the merged point set, Represents the cylindrical feature tolerance parameter of the merged point set, Represents the angle between the central axis of the cylinder of the merged point set and the zenith direction.
8. The FPGA-based laser radar point cloud online sample self-generation method according to claim 7, characterized in that: The merged point set is initialized and labeled based on height features, planarity features, spatial projection features and cylindrical model features, and is divided into four categories of labeled training samples.
9. The FPGA-based laser radar point cloud online sample self-generation method according to claim 8, characterized in that: The process of calculating the confidence of the training sample is: a. Perform cross-validation for each type of labeled training samples and perform multiple iterations. In a single iteration, each type of labeled training samples is divided into a training set and a test set; b. Select a set number of training samples from any one of the labeled training samples as positive samples, and select a set number of training samples from the remaining three types of labeled training samples as negative samples; use the target representative feature set corresponding to the positive and negative samples Train a linear SVM classifier; c. Use the SVM classifier to classify the samples in the test set into positive samples and negative samples; d. Repeat the process from a to c and record the number of times each sample is classified as a positive sample; e. For a training sample that is marked as a positive sample n times after N iterations, the probability that the initial label of the training sample is correct is obtained by transforming the Bayesian formula: ; Where N represents the total number of iterations, n represents the number of times the sample is marked as a positive sample in N iterations, It indicates the probability that a training sample is labeled as a positive sample n times after N iterations, and the initial label of the sample is correct. represents the prior probability that the initial label of a training sample is correct, represents the prior probability that the initial label of a training sample is wrong, It indicates the probability that a training sample is labeled as a positive sample n times in N iterations when the initial label of the training sample is correct; It represents the probability that a training sample will be labeled as a positive sample n times in N iterations when the initial label of the training sample is wrong. represents the number of combinations marked as positive samples from N iterations, represents the probability of the linear SVM classifier being correct each time, ; Determine whether the probability of the initial label of the training sample being correct is less than the set threshold, and filter out the training samples whose correct probability is less than the set threshold as error samples.
Citation Information
Patent Citations
Laser radar data acquisition method and system based on FPGA and ARM
CN114167385A
Pavement scene target identification method based on laser radar point cloud
CN114612795A
FPGA-based point cloud compression acceleration device and method
CN117314723A