Closed scene data enhanced point cloud generation method

By collecting and processing data in open scenarios, more true data is generated and filled in closed scenarios, the problem of insufficient data in closed scenarios is solved, and the performance and adaptability of deep learning models are improved.

CN120014173APending Publication Date: 2025-05-16安徽中科星驰自动驾驶技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510163892.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-23
Filing Date
2025-02-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In closed scenarios, the deep learning model lacks sufficient truth data, resulting in limited detection performance and convergence speed, affecting the accuracy and stability of three-dimensional object detection.

Method used

By collecting enough samples in an open scenario, completing data and sampling, generating more truth-value data, and filling it into a specific closed scenario, increasing the diversity and richness of the data, thereby improving the performance of deep learning models.

Benefits of technology

It effectively improves the detection effect and adaptability of the point cloud detection model in specific scenarios, enhances the quantity and quality of data, and ensures the stability and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014173A_ABST
    Figure CN120014173A_ABST
Patent Text Reader

Abstract

The invention relates to a point cloud generation method for closed scene data enhancement, and aims to generate a large number of true value targets when a closed scene does not have a large number of true values so as to improve the detection effect and robustness of a model for performing 3D target detection by using deep learning. According to the method, all truth values in open road data are collected, point cloud completion is carried out on the truth values through an interpolation method, coordinates of all points of truth value point clouds are changed, random downsampling and farthest point sampling methods are selected, and a rich truth value sample library is obtained according to different sampling proportions. Then, selecting enough samples from the truth value sample library, filling the samples into closed scene point cloud frames, and adding corresponding truth value labels; and finally, the generated new point cloud data set is used for iterative training, so that the detection effect of different types of objects in the corresponding closed scene can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing for deep learning training, and in particular to a closed scene data enhancement point cloud generation method. The present invention can be widely used in different fields such as autonomous driving and industrial automation. Background Art

[0002] With the rapid development of deep learning, it has been widely used in various fields such as autonomous driving and industrial automation, greatly improving the efficiency of production and life. Autonomous driving requires accurate perception and understanding of the surrounding environment to achieve safe and efficient driving. To achieve this goal, autonomous driving vehicles are usually equipped with a variety of sensors, such as radar, camera, lidar, etc. Among them, lidar is a sensor that can collect three-dimensional point cloud data. It can accurately record the outer surface of objects and scenes, with high precision, high resolution, and no influence of light. In recent years, many point cloud data processing and learning methods based on deep learning have been proposed. One of the important tasks is three-dimensional target detection, that is, identifying objects of different categories in point cloud data and representing their position and posture with three-dimensional bounding boxes, which is of great significance for environmental perception and decision-making of autonomous driving.

[0003] However, unlike 2D object detection tasks, the true values ​​in 3D object detection tasks are smaller and more sparse. In the entire 3D space, they usually only occupy a small fraction of all points, which means that most of the space does not contain useful information. At the same time, for some specific closed scenes in practical applications, such as airports and factories, there is usually no large amount of true value data to support model training in these scenes, resulting in deep learning models only being able to learn features from very limited information, which severely limits the detection performance and convergence speed of the model, and affects the accuracy and stability of 3D object detection in specific scenes. Summary of the invention

[0004] In response to the above-mentioned problems, the present invention proposes a point cloud generation method for closed scene data enhancement. By collecting enough samples in open scenes, the original data is supplemented and sampled according to the morphology and distance to generate more true value data and fill in specific closed scenes, so as to increase the diversity and richness of the data, thereby improving the performance of the deep learning model in specific scenarios.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A closed scene data enhanced point cloud generation method, which adds true value data for a specific closed scene by completing and sampling the true value of point clouds in different scenes, and the method comprises the following steps:

[0007] S1: Collect the original true value cloud points and annotate them;

[0008] S2: interpolation to complete the true value cloud points;

[0009] S3: changing the coordinates of all points of the true point cloud;

[0010] S4: downsample the newly generated true value point cloud to obtain a true value sample library;

[0011] S5: Select samples from the true value sample library to fill in a specific closed scene, and add the corresponding true value labels to generate a new point cloud data set;

[0012] S6: Perform iterative training based on the generated new point cloud dataset.

[0013] As a further technical solution of the present invention, the step of collecting original true value cloud points and marking them includes:

[0014] Get a batch of original point cloud frames, annotate the point cloud frames, get the set of all annotated truth values ​​{GT}, and calculate the center point coordinates of all truth values ​​in {GT} Set to , that is, all true values ​​are translated to the origin in the XY plane.

[0015] As a further technical solution of the present invention, the step of interpolating and completing the true value cloud points includes: based on the true value cloud point samples with the densest and richest points in the sample library, interpolating, filling and completing other true value cloud point samples.

[0016] As a further technical solution of the present invention, the step of changing the coordinates of all points of the true value point cloud includes:

[0017] Randomly select N point cloud samples and their corresponding annotation boxes from the interpolated and completed true value cloud point samples, and perform random translation and rotation perturbations on these N point cloud samples, that is:

[0018] ;

[0019] ;

[0020] ;

[0021] in, The newly generated true value coordinates of all points, are the coordinates of all points of the original true value, x_offset, y_offset are the translation distances in the x and y directions, yaw is the angle of rotation around the z axis, the center of the rear axle of the acquisition vehicle is taken as the origin, the direction of the front of the vehicle is the positive direction of the x axis, the left direction of the vehicle is the positive direction of the y axis, and the vertical upward direction is the positive direction of the z axis.

[0022] As a further technical solution of the present invention, in step S4, the downsampling method is as follows:

[0023] The samples are randomly downsampled according to the distance ratio K between the newly generated target and the origin, that is, K true value points are randomly retained. The formula is as follows;

[0024] ;

[0025] in, To generate the center point coordinates of the new target, It is half of the length and width of the point cloud detection range.

[0026] As a further technical solution of the present invention, in step S4, the downsampling method is as follows:

[0027] The farthest point sampling is performed on the true point cloud samples: that is, first find the centroid of the entire point cloud, select the point farthest from the centroid, recorded as P0, and continue to select the point farthest from P0 among all the remaining points, recorded as P1; for each remaining point, calculate the distance to P0 and P1 respectively, and select the shortest one as the distance from this point to P0 and P1 as a whole; select the point with the largest distance, recorded as P2; repeat this step until the true value point that meets the proportion is selected.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] The present invention mainly provides a large number of true values ​​for training for real closed scenes, while retaining the authenticity of the data to a great extent, effectively improving the quantity and quality of point cloud data, and can effectively improve the detection effect and adaptability of the point cloud detection model in specific scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 The original point cloud ground truth map collected in the point cloud generation method for closed scene data enhancement.

[0031] Figure 2 This is a completion diagram for the point cloud truth value in the point cloud generation method for closed scene data enhancement.

[0032] Figure 3 Sampling diagram for completing point cloud in point cloud generation method for closed scene data enhancement.

[0033] Figure 4 A point cloud frame image of a specific closed scene in the point cloud generation method for closed scene data enhancement.

[0034] Figure 5 Point cloud frame after adding a large amount of ground truth to the point cloud generation method for closed scene data enhancement.

[0035] Figure 6 Flowchart of the point cloud generation method for closed scene data augmentation. DETAILED DESCRIPTION

[0036] The technical solution of this patent is further described in detail below in conjunction with specific implementation methods.

[0037] Example 1

[0038] See also Figure 6 The embodiment of the present invention provides a closed scene data enhanced point cloud generation method, which adds true value data for a specific closed scene by completing and sampling the point cloud true values ​​in different scenes. The method includes the following steps:

[0039] S1: Collect the original true value cloud points and annotate them, such as Figure 1 As shown, it is the collected real target point cloud;

[0040] Get a batch of original point cloud frames, which can be taken on an open road or in a closed scene, annotate the point cloud frames, get the set of all annotated truth values ​​{GT}, and calculate the center point coordinates of all truth values ​​in {GT}. Set to , that is, all true values ​​are translated to the origin in the XY plane.

[0041] S2: interpolation to complete the true value cloud points, such as Figure 2 The completed target point cloud is shown;

[0042] Based on the most densely populated and abundant true value cloud point samples in the sample library, other true value cloud point samples are interpolated and filled to complete, so that the samples in the sample library can be homogenized while retaining the complete shape information of each sample.

[0043] S3: changing the coordinates of all points of the true point cloud;

[0044] Randomly select N point cloud samples and their corresponding annotation boxes from the interpolated and completed true value cloud point samples, and perform random translation and rotation perturbations on these N point cloud samples, that is:

[0045] ;

[0046] ;

[0047] ;

[0048] in, The newly generated true value coordinates of all points, are the coordinates of all points of the original true value, x_offset, y_offset are the translation distances in the x and y directions, yaw is the angle of rotation around the z axis, the center of the rear axle of the acquisition vehicle is taken as the origin, the direction of the front of the vehicle is the positive direction of the x axis, the left direction of the vehicle is the positive direction of the y axis, and the vertical upward direction is the positive direction of the z axis.

[0049] S4: Downsample the newly generated true value point cloud to obtain a true value sample library, such as Figure 3 As shown, it is the target point cloud after downsampling;

[0050] Method 1: In the original point cloud data obtained by the LiDAR, the sample density varies at different locations; the farther the object is from the sensor, the sparser the points it contains; the closer the object is to the sensor, the denser the points it contains; therefore, we randomly downsample the samples according to the distance ratio K between the newly generated target and the origin, that is, randomly retain K true value points, the formula is as follows;

[0051] ;

[0052] in, To generate the center point coordinates of the new target, It is half of the length and width of the point cloud detection range.

[0053] Method 2: Sampling the farthest point of the true point cloud samples: first find the centroid of the entire point cloud, select the point farthest from the centroid, record it as P0, and continue to select the point farthest from P0 among all the remaining points, record it as P1; for each remaining point, calculate the distance to P0 and P1 respectively, and select the shortest one as the distance from this point to P0 and P1 as a whole; after calculating these distances, select the point with the largest distance, record it as P2; repeat this step until the true value point that meets the proportion is selected.

[0054] S5: Select samples from the true value sample library to fill in a specific closed scene, and add its corresponding true value label to generate a new point cloud dataset, such as Figure 4 Shown is the point cloud frame of the original specific closed scene, such as Figure 5 Shown is the generated new scene point cloud frame;

[0055] S6: Perform iterative training based on the generated new point cloud dataset.

[0056] In this embodiment, when the true value point cloud selected from the true value sample library is added to the original point cloud file of a specific closed scene, it is necessary to determine whether a point cloud already exists at the added position; the corresponding center point coordinates of the camera coordinate system are generated according to the intrinsic parameter matrix of the added frame and the center point in the vehicle coordinate system, and added to the label file in the format of the true value label.

[0057] When judging whether there are already point clouds at the added location: if the generated location was originally the foreground, that is, there is already a true value, the generated true value is discarded; if the generated location was originally the background, the points within the generated true value range are removed. The circumscribed rectangle of the circumscribed circle is The length l and width w of the circumscribed circle are used to calculate the radius r of the circumscribed circle. The center point and r to calculate the extent of the circumscribed square.

[0058] The functions that can be achieved by the above-mentioned point cloud generation method for closed scene data enhancement are all completed by a computer device, and the computer device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories. The program code is loaded and executed by the one or more processors to achieve the functions of the point cloud generation method for closed scene data enhancement.

[0059] The processor takes out instructions from the memory one by one, analyzes the instructions, and then completes the corresponding operations according to the instruction requirements, generating a series of control commands, so that the various parts of the computer can automatically, continuously and coordinately move to become an organic whole, realize the input of programs, the input of data, and the calculation and output of results. The arithmetic operations or logical operations generated in this process are all completed by the operator; the memory includes a read-only memory (ROM), which is used to store computer programs, and a protection device is provided outside the memory.

[0060] Exemplarily, the computer program may be divided into one or more modules, one or more modules are stored in a memory and executed by a processor to implement the present invention. One or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in a terminal device.

[0061] Those skilled in the art will understand that the description of the above service equipment is merely an example and does not constitute a limitation on the terminal equipment. It may include more or fewer components than described above, or a combination of certain components, or different components, for example, it may include input and output devices, network access devices, buses, etc.

[0062] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire user terminal.

[0063] The above-mentioned memory can be used to store computer programs and / or modules. The above-mentioned processor realizes various functions of the above-mentioned terminal device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as an information collection template display function, a product information release function, etc.), etc.; the data storage area can store data created according to the use of the berth status display system (such as product information collection templates corresponding to different product types, product information that different product providers need to release, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0064] If the module / unit integrated in the terminal device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the modules / units in the above-mentioned embodiment system, and can also be completed by instructing the relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium, and the computer program can realize the functions of the above-mentioned various system embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Computer-readable media can include: any entity or device that can carry computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0065] It should be noted that, in this article, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0066] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A closed scene data enhanced point cloud generation method, characterized in that: The point cloud truth values ​​in different scenes are added to the truth value data for a specific closed scene through completion and sampling. The method includes the following steps: S1: Collect the original true value cloud points and annotate them; S2: interpolation to complete the true value cloud points; S3: changing the coordinates of all points of the true point cloud; S4: downsample the newly generated true value point cloud to obtain a true value sample library; S5: Select samples from the true value sample library to fill in a specific closed scene, and add the corresponding true value labels to generate a new point cloud data set; S6: Perform iterative training based on the generated new point cloud dataset.

2. The closed scene data enhanced point cloud generation method according to claim 1, characterized in that: The steps of collecting original true value cloud points and marking them include: Get a batch of original point cloud frames, annotate the point cloud frames, get the set of all annotated truth values ​​{GT}, and calculate the center point coordinates of all truth values ​​in {GT} Set to , that is, all true values ​​are translated to the origin in the XY plane.

3. The closed scene data enhanced point cloud generation method according to claim 1, characterized in that: The step of interpolating and completing the true value cloud points includes: based on the true value cloud point samples with the densest and richest points in the sample library, interpolating, filling and completing other true value cloud point samples.

4. The closed scene data enhanced point cloud generation method according to claim 3, characterized in that: The step of changing the coordinates of all points of the true value point cloud comprises: Randomly select N point cloud samples and their corresponding annotation boxes from the interpolated and completed true value cloud point samples, and perform random translation and rotation perturbations on these N point cloud samples, that is: ; ; ; in, The newly generated true value coordinates of all points, are the coordinates of all points of the original true value, x_offset, y_offset are the translation distances in the x and y directions, yaw is the angle of rotation around the z axis, the center of the rear axle of the acquisition vehicle is taken as the origin, the direction of the front of the vehicle is the positive direction of the x axis, the left direction of the vehicle is the positive direction of the y axis, and the vertical upward direction is the positive direction of the z axis.

5. The closed scene data enhanced point cloud generation method according to claim 1, characterized in that: In step S4, the downsampling method is as follows: The samples are randomly downsampled according to the distance ratio K between the newly generated target and the origin, that is, K true value points are randomly retained. The formula is as follows; ; in, To generate the center point coordinates of the new target, It is half of the length and width of the point cloud detection range.

6. The closed scene data enhanced point cloud generation method according to claim 1, characterized in that: In step S4, the downsampling method is as follows: The farthest point sampling is performed on the true point cloud samples: that is, first find the centroid of the entire point cloud, select the point farthest from the centroid, recorded as P0, and continue to select the point farthest from P0 among all the remaining points, recorded as P1; for each remaining point, calculate the distance to P0 and P1 respectively, and select the shortest one as the distance from this point to P0 and P1 as a whole; select the point with the largest distance, recorded as P2; repeat this step until the true value point that meets the proportion is selected.