Point cloud simulation model training method and related equipment
By constructing a point cloud simulation model and optimizing reflection intensity and loss rate using camera intrinsic parameters and hybrid spatiotemporal coding, the problems of long rendering time and insufficient accuracy in point cloud simulation methods are solved, and real-time high-precision point cloud rendering is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-03-10
AI Technical Summary
Existing point cloud simulation methods cannot achieve real-time rendering and have insufficient rendering accuracy. The neural radiation field method is time-consuming, while the three-dimensional Gaussian splashing method can only be used for image rendering and cannot be used for point cloud rendering.
Multiple point cloud data are acquired through training equipment to construct an initial point cloud simulation model. The model is then converted into a depth map using camera intrinsic parameters. Hybrid spatiotemporal coding is used to optimize reflection intensity and loss rate. The point cloud components of the background, road, and target object are decoupled and optimized to achieve real-time rendering of the point cloud and improve accuracy.
It achieves real-time rendering and accuracy improvement of point clouds, and is suitable for simulation of different types of LiDAR, solving the problems of long rendering time and insufficient accuracy in existing technologies.
Smart Images

Figure CN121639879A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence (AI), and in particular to a point cloud simulation model training method and related equipment. BACKGROUND
[0002] In the past two years, with the rapid rise of neural radiance fields (NeRF) in the field of computer vision, many image generation methods based on neural radiance fields have emerged. These methods have shown the unique advantages of neural radiance fields over other traditional methods in application scenarios such as autonomous driving, digital humans, and three-dimensional scene modeling.
[0003] Neural radiance fields (NeRF) is a new method for scene representation and image rendering. It records the scene representation in a deep neural network through implicit representation, uses a deep neural network to implicitly learn a static three-dimensional scene, and indirectly completes tasks such as three-dimensional reconstruction of the scene and generation of new view images. This method uses a fully connected (non-convolutional) deep network, inputs a single continuous 5D coordinate (spatial position (x, y, z) and viewing direction (θ, φ)), and outputs the volume density of the spatial position and the color information related to the view. The view is synthesized by querying the 5D coordinate along the camera light ray, and the output color and density are projected into the image using classical volume rendering techniques. In addition, when used for point cloud simulation, the point cloud is converted into a distance image, and the scene geometry distribution is supervised using the point cloud depth information, which can achieve scene reconstruction and end-to-end point cloud simulation. However, the simulation process takes a long time, and real-time simulation of the point cloud cannot be achieved.
[0004] Three-dimensional Gaussian splatting (3DGS) is another emerging method for scene reconstruction and image rendering. Compared with the neural radiance field rendering method, 3DGS discards the process of emitting a ray for each pixel and sampling a large number of points on the ray, and instead represents the scene as a Gaussian sphere to obtain the rendering result by projecting the Gaussian sphere, achieving real-time rendering. However, this method can only be used for image rendering and cannot be used for point cloud rendering. SUMMARY
[0005] The embodiments of the present application provide a point cloud simulation model training method and related equipment. The use of the embodiments of the present application is beneficial to achieve real-time rendering of point clouds and improve the rendering accuracy of point clouds.
[0006] In a first aspect, the embodiments of the present application provide a point cloud simulation model training method. The method comprises:
[0007] The training device obtains a plurality of first point cloud data at a plurality of different positions in a target scene; the training device constructs an initial point cloud simulation model and determines a plurality of pose information of a target object based on the plurality of first point cloud data, parameters of the initial point cloud simulation model including reflection intensity and loss rate of a Gaussian sphere corresponding to a point in the point cloud; the target object is an object in the plurality of first point cloud data; the training device converts the plurality of first point cloud data into a plurality of first depth maps based on camera intrinsic parameters; and the training device trains the initial point cloud simulation model based on the plurality of pose information and the plurality of first depth maps to obtain a target point cloud simulation model.
[0008] It can be seen that by virtually converting the plurality of first point cloud data sequences into a plurality of first depth maps, and then training the initial point cloud simulation model using the first depth maps, the joint optimization of the depth information, reflection intensity and loss rate of the point cloud is realized. When rendering the point cloud using the trained point cloud simulation model, real-time point cloud simulation is facilitated, and the simulation accuracy of the point cloud is improved. Furthermore, the conversion of the first point cloud data into the first depth maps used to train the initial point cloud simulation model is based on camera intrinsic parameters, so that the trained point cloud simulation model can be applied to the simulation of point clouds of different types of laser radars.
[0009] In combination with the first aspect, in a possible implementation, the parameters of the initial point cloud simulation model include initial parameters of a background point cloud, initial parameters of a target object point cloud, and initial parameters of a road point cloud. The training device constructs the initial point cloud simulation model based on the plurality of first point cloud data, including:
[0010] The training device determines point cloud data of the target object in the plurality of first point cloud data, and removes the point cloud data of the target object from each first point cloud data to obtain a plurality of second point cloud data. The training device constructs a point cloud map of the target scene based on the plurality of second point cloud data. The initial parameters of the background point cloud and the initial parameters of the road point cloud are determined based on the point cloud map of the target scene. The training device obtains mesh data of the target object based on the point cloud data of the target object in the plurality of first point cloud data. The training device down-samples the mesh data of the target object to obtain the initial parameters of the target object point cloud.
[0011] It can be seen that by decoupling the point cloud simulation model into a background point cloud part, a road point cloud part, and a target object point cloud part, subsequent optimization of the road point cloud part and the target object point cloud part during training is facilitated, which in turn facilitates improving the accuracy of the point cloud simulation model. Furthermore, the mesh data of the target object obtained based on the point cloud data of the target object in the plurality of first point cloud data is used to initialize the point cloud parameters of the target object, which solves the problem of inaccurate geometry reconstruction and penetration of the target object point cloud due to the sparsity of the target object point cloud in the first point cloud data.
[0012] With reference to the first aspect, in a possible implementation manner, the initial parameters of the background point cloud include parameters of Gaussian spheres corresponding to all points in the background point cloud, the initial parameters of the target object point cloud include parameters of Gaussian spheres corresponding to all points in the target object point cloud, the initial parameters of the road point cloud include parameters of Gaussian spheres corresponding to all points in the road point cloud, and a principal axis normal of the Gaussian sphere corresponding to all points in the road point cloud is consistent with a direction of the road in the target scene; the parameters of the Gaussian sphere include a position parameter, a rotation parameter, an opacity, a scale parameter, a reflection intensity, a loss rate and a material characteristic parameter.
[0013] It can be seen that, for the parameters of the Gaussian sphere, the reflection intensity, the loss rate and the material characteristic parameter are introduced on the basis of the original position parameter, the rotation parameter, the opacity and the scale parameter, and the parameters of the Gaussian sphere are trained, so that subsequent use of the point cloud simulation model is conducive to improving the accuracy of the obtained point cloud.
[0014] With reference to the first aspect, in a possible implementation manner, the training device trains the initial point cloud simulation model based on the plurality of pose information and the plurality of first depth maps to obtain the target point cloud simulation model, including:
[0015] The training device performs hybrid space-time coding based on the plurality of pose information, the plurality of time points, and the initial parameters of the target object point cloud to obtain reflection intensity offsets of the plurality of time points and loss rate offsets of the plurality of time points; the plurality of time points are respectively collection time points of the plurality of first point cloud data; the training device obtains first parameters of the target object point cloud at the plurality of time points based on the reflection intensity offsets of the plurality of time points, the loss rate offsets of the plurality of time points, and the initial parameters of the target object point cloud, wherein the first parameters include a first reflection intensity and a first loss rate, the first reflection intensity is determined based on a reflection intensity in the initial parameters of the target object point cloud and the reflection intensity offset, the time point corresponding to the first reflection intensity is the same as the time point corresponding to the reflection intensity offset, the first loss rate is determined based on a loss rate in the initial parameters of the target object point cloud and the loss rate offset, and the time point corresponding to the first loss rate is the same as the time point corresponding to the loss rate offset; the training device performs Gaussian rasterization processing on the plurality of third point cloud data to obtain a plurality of second depth maps; the plurality of third point cloud data correspond to the plurality of time points, and each third point cloud data includes initial parameters of a background point cloud, initial parameters of a road point cloud, second parameters, and the first parameters of the same time point as the third point cloud data; the second parameters include parameters other than the reflection intensity and the loss rate in the initial parameters of the target object point cloud; the training device calculates a depth loss value, a reflection intensity loss value, and a loss rate loss value based on the plurality of first depth maps and the plurality of second depth maps; and the training device adjusts the initial parameters of the background point cloud, the initial parameters of the target object point cloud, and the initial parameters of the road point cloud based on the depth loss value, the reflection intensity loss value, and the loss rate loss value to obtain target parameters of the background point cloud, target parameters of the target object point cloud, and target parameters of the road point cloud, and the parameters of the target point cloud simulation model include the target parameters of the background point cloud, the target parameters of the target object point cloud, and the target parameters of the road point cloud.
[0016] It can be seen that by introducing hybrid space-time coding to predict the offsets of the reflection intensity and the loss rate of the target object point cloud at different time points, and adjusting the reflection intensity and the loss rate in the initial parameters of the target object point cloud based on the offsets of the reflection intensity and the loss rate, the problem of inconsistency of the reflection intensity and the loss rate of the target object point cloud due to different positions / locations of the target object at different time points in the continuous frames (i.e., the plurality of first point cloud data) is overcome, and the simulation accuracy of the target object point cloud in a dynamic scene is ensured. When training the point cloud simulation model, the training device adjusts the parameters of the initial point cloud simulation model based on the depth loss value, the reflection intensity loss value, and the loss rate loss value, which is conducive to improving the accuracy of the point cloud simulation model.
[0017] With reference to the first aspect, in a possible implementation manner, the training device performs mixed space-time coding based on the plurality of pose information, the plurality of time instants and the initial parameters of the target object point cloud, to obtain the reflection intensity offset of the plurality of time instants and the loss rate offset of the plurality of time instants, including:
[0018] The training device encodes the plurality of pose information to obtain pose features, and encodes the plurality of time instants to obtain time features. The training device extracts three three-dimensional plane features based on the plurality of time instants and the plurality of pose information and the initial parameters of the target object point cloud. The training device processes the pose features, the time features and the three three-dimensional plane features to obtain the reflection intensity offset of the plurality of time instants and the loss rate offset of the plurality of time instants.
[0019] It can be seen that the mixed space-time coding is introduced, the offset of the reflection intensity and the offset of the loss rate corresponding to the target object point cloud at different time instants are predicted, which is convenient for subsequent adjustment of the reflection intensity and the loss rate in the initial parameters of the target object point cloud based on the offset of the reflection intensity and the offset of the loss rate, and the problem of inconsistency of the reflection rate and the loss rate corresponding to the target object point cloud due to different positions / locations of the target object point cloud at different time instants between continuous frames (i.e. a plurality of first point cloud data) is overcome, and the accuracy of simulation of the target object point cloud in a dynamic scene is ensured.
[0020] With reference to the first aspect, in a possible implementation manner, the training device converts the plurality of first point cloud data into a plurality of first depth maps based on the camera intrinsic parameters, including:
[0021] The training device projects the first point cloud data into a first depth map based on the camera intrinsic parameters, wherein the number of points in the first point cloud data projected onto a unique pixel is maximum.
[0022] The number of points in the first point cloud data projected onto a unique pixel is maximized during projection, and thus as much point cloud data information as possible is retained in the first depth map, which reduces the loss of point cloud data information compared to directly projecting the point cloud data into a distance image, and the use of the first point cloud data to train the point cloud simulation model is conducive to improving the accuracy of the point cloud simulation model.
[0023] In a second aspect, an embodiment of the present application provides a training device. The training device includes an acquisition unit, a determination unit and a training unit.
[0024] The acquisition unit is configured to acquire a plurality of first point cloud data at a plurality of different positions in a target scene.
[0025] The determining unit is configured to construct an initial point cloud simulation model based on the plurality of first point cloud data and determine a plurality of pose information of the target object, parameters of the point cloud simulation model including reflection intensity and loss rate of a Gaussian sphere corresponding to a point in the point cloud; the target object is an object in the plurality of first point cloud data; and the plurality of first point cloud data is converted into a plurality of first depth maps based on camera intrinsic parameters.
[0026] The training unit is configured to train the initial point cloud simulation model based on the plurality of pose information and the plurality of first depth maps to obtain a target point cloud simulation model.
[0027] In a feasible implementation, the parameters of the initial point cloud simulation model include initial parameters of a background point cloud, initial parameters of a target object point cloud, and initial parameters of a road point cloud; and in the aspect of constructing the initial point cloud simulation model based on the plurality of first point cloud data, the determining unit is configured to:
[0028] determine point cloud data of the target object in the plurality of first point cloud data, remove the point cloud data of the target object from each first point cloud data to obtain a plurality of second point cloud data, construct a point cloud map of a target scene based on the plurality of second point cloud data, determine initial parameters of the background point cloud and initial parameters of the road point cloud based on the point cloud map of the target scene, obtain mesh data of the target object based on the point cloud data of the target object in the plurality of first point cloud data, and down-sample the mesh data of the target object to obtain initial parameters of the target object point cloud.
[0029] In a feasible implementation, the initial parameters of the background point cloud include parameters of a Gaussian sphere corresponding to all points in the background point cloud, the initial parameters of the target object point cloud include parameters of a Gaussian sphere corresponding to all points in the target object point cloud, and the initial parameters of the road point cloud include parameters of a Gaussian sphere corresponding to all points in the road point cloud, a principal axis normal of the Gaussian sphere corresponding to all points in the road point cloud is consistent with a direction of a road in the target scene; and the parameters of the Gaussian sphere include a position parameter, a rotation parameter, an opacity, a scale parameter, a reflection intensity, a loss rate, and a material characteristic parameter.
[0030] In a feasible implementation, the training unit is configured to:
[0031] The initial parameters of the target object point cloud are mixed based on the plurality of pose information, the plurality of time points and the initial parameters of the target object point cloud to obtain a plurality of time point reflection intensity offsets and a plurality of time point loss rate offsets; the plurality of time points are respectively collection time points of the plurality of first point cloud data; the first parameters of the target object point cloud at the plurality of time points are obtained based on the plurality of time point reflection intensity offsets, the plurality of time point loss rate offsets and the initial parameters of the target object point cloud, wherein the first parameters include a first reflection intensity and a first loss rate, the first reflection intensity is determined based on the reflection intensity in the initial parameters of the target object point cloud and the reflection intensity offset, the time point corresponding to the first reflection intensity is the same as the time point corresponding to the reflection intensity offset, the first loss rate is determined based on the loss rate in the initial parameters of the target object point cloud and the loss rate offset, and the time point corresponding to the first loss rate is the same as the time point corresponding to the loss rate offset; the plurality of third point cloud data are subjected to Gaussian rasterization processing to obtain a plurality of second depth maps; the plurality of third point cloud data correspond to the plurality of time points, and each of the third point cloud data includes the initial parameters of the background point cloud, the initial parameters of the road point cloud, a second parameter, and the first parameters of the same time point as the time point corresponding to the third point cloud data; the second parameter includes parameters other than the reflection intensity and the loss rate in the initial parameters of the target object point cloud; depth loss values, reflection intensity loss values and loss rate loss values are calculated based on the plurality of first depth maps and the plurality of second depth maps; the initial parameters of the background point cloud, the initial parameters of the target object point cloud and the initial parameters of the road point cloud are adjusted based on the depth loss values, the reflection intensity loss values and the loss rate loss values to obtain target parameters of the background point cloud, target parameters of the target object point cloud and target parameters of the road point cloud, and the parameters of the target point cloud simulation model include the target parameters of the background point cloud, the target parameters of the target object point cloud and the target parameters of the road point cloud.
[0032] In a feasible implementation, in the aspect of mixing space-time coding based on the plurality of pose information, the plurality of time points and the initial parameters of the target object point cloud to obtain the plurality of time point reflection intensity offsets and the plurality of time point loss rate offsets, the training unit is specifically configured to:
[0033] The plurality of pose information is coded to obtain pose features; the plurality of time points is coded to obtain time features; three three-dimensional plane features are extracted based on the initial parameters of the target object point cloud, the plurality of time points and the plurality of pose information; the pose features, the time features and the three three-dimensional plane features are processed to obtain the plurality of time point reflection intensity offsets and the plurality of time point loss rate offsets.
[0034] In a feasible implementation, in the aspect of converting the plurality of first point cloud data into the plurality of first depth maps based on the camera intrinsic parameters, the determination unit is specifically configured to:
[0035] Project the first point cloud data into a first depth map based on the camera intrinsic parameters, wherein the number of points in the first point cloud data that project onto a unique pixel is the largest.
[0036] In a third aspect, an embodiment of the present application provides a training device, including a processor and a memory. The memory is configured to store program code. The processor is configured to invoke the program code stored in the memory to execute the method provided in the first aspect or any possible implementation manner of the first aspect.
[0037] In a fourth aspect, an embodiment of the present application provides a computer storage medium, including computer instructions. When the computer instructions are run on an electronic device, the electronic device is caused to execute the method provided in any possible implementation manner of the first aspect.
[0038] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the method provided in any possible implementation manner of the first aspect.
[0039] It can be understood that the training device provided in the second aspect or the third aspect is used to execute the method provided in any of the first aspect, and the computer storage medium provided in the fourth aspect and the computer program product provided in the fifth aspect are both used to execute the method provided in any of the first aspect. Therefore, the beneficial effects that can be achieved are referred to the beneficial effects in the corresponding method, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 A system architecture schematic diagram is provided for the embodiment of the present application;
[0041] Figure 2 A point cloud simulation model training process schematic diagram is provided for the embodiment of the present application;
[0042] Figure 3 A point cloud simulation model training method flow schematic diagram is provided for the embodiment of the present application;
[0043] Figure 4a A comparison diagram of the normal of the Gaussian sphere and the road direction is provided for the embodiment of the present application;
[0044] Figure 4b A road Gaussian sphere schematic diagram is provided for the embodiment of the present application;
[0045] Figure 5 A hybrid space-time coding schematic diagram is provided for the embodiment of the present application;
[0046] Figure 6 A rendering effect comparison diagram for a static scene is shown;
[0047] Figure 7 The schematic diagram before and after the target object in the point cloud is removed is shown;
[0048] Figure 8 The point cloud diagram for different perspectives of a road scene is shown;
[0049] Figure 9 The simulation schematic diagram of different types of laser point clouds is shown;
[0050] Figure 10 The structural schematic diagram of a training device provided by an embodiment of the present application is shown;
[0051] Figure 11 The structural schematic diagram of another training device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0052] The terms "first", "second", "third", and "fourth" and the like in the specification and claims of the present application and the drawings are used to distinguish different objects, and are not used to describe a specific order.
[0053] "Multiple" means two or more. "And / or" describes the association relationship of the associated objects, indicating that there are three relationships, for example, A and / or B means that there are three cases of A alone, A and B together, and B alone. The character " / " generally represents that the front and rear associated objects are in an "or" relationship.
[0054] Embodiments of the present application are described below with reference to the accompanying drawings.
[0055] Reference is made to Figure 1 , Figure 1 The system architecture schematic diagram provided by an embodiment of the present application is shown. As shown in Figure 1 , the system architecture includes a point cloud simulation device 101, a plurality of point cloud data acquisition devices 102, and a training device 103. Among them, the point cloud simulation device 101 has one or more.
[0056] The point cloud simulation device 101 can be a server, such as a cloud server, a distributed server, a cabinet server, a blade server, a tower server, etc.
[0057] The point cloud data acquisition device 102 can be a vehicle, a terminal device, etc. The terminal device can be a smart phone, a smart watch, a smart bracelet, a desktop computer, a notebook computer, a tablet, etc. The vehicle can be a bicycle, a tricycle, a small car, a medium-sized car, etc.
[0058] The training device 103 is a device capable of data processing, such as a server, a computer, etc.
[0059] The training device 103 and the point cloud simulation device 101 can be the same device or different devices.
[0060] A point cloud data acquisition device (such as a lidar) 102 acquires multiple first point cloud data at multiple locations in the target scene and transmits these multiple first point cloud data to a training device 103. The training device 103 trains a target point cloud simulation model based on the multiple first point cloud data.
[0061] In a specific example, such as Figure 2 As shown, the training device 103 converts multiple first point cloud data into multiple first depth maps based on the intrinsic parameters of the first camera. The training device 103 performs target detection and tracking on the multiple first point cloud data using a target detector and a target tracker to obtain the point cloud data corresponding to the target objects in the multiple first point cloud data and multiple pose information corresponding to the target objects in the multiple first point cloud data. There can be multiple target objects, such as three. The training device 103 removes the point cloud data corresponding to the target objects from the multiple first point cloud data to obtain multiple second point cloud data. The training device 103 constructs a point cloud map of the target scene based on the multiple second point cloud data, and determines the initial parameters of the background point cloud and the road point cloud based on the point cloud map of the target scene. The training device 103 obtains the mesh data of the target objects based on the point cloud data of the target objects. In one example, the training device processes the point cloud data of the target objects based on the MV-deepSDF model to obtain the mesh data of the target objects. The mesh data of the target object is downsampled to obtain the initial parameters of the target object point cloud. The initial parameters of the background point cloud include the parameters of the Gaussian sphere corresponding to all points in the background point cloud, the target object point cloud, and the road point cloud. The principal axis normal of the Gaussian sphere corresponding to all points in the road point cloud is consistent with the direction of the road in the target scene. The parameters of the Gaussian sphere include position parameters, rotation parameters, opacity, scale parameters, reflection intensity, loss rate, and material feature parameters, which are used as learning parameters. The training device 103 trains the initial parameters of the background point cloud, the target object point cloud, and the road point cloud based on multiple pose information, multiple first point cloud data, and multiple first depth maps to obtain the target parameters of the background point cloud, the target object point cloud, and the road point cloud. The parameters of the trained point cloud simulation model include the target parameters of the target background point cloud, the background point cloud, and the road point cloud. The loss values used during training include depth loss, reflection intensity loss, and loss rate loss.
[0062] The point cloud simulation device 101 obtains pose information of a target object and camera intrinsic parameters; the point cloud simulation device 101 processes based on a point cloud simulation model and the pose information of the target object to obtain a third depth map corresponding to a second time; and the point cloud simulation device 101 processes the third depth map corresponding to the second time based on the camera intrinsic parameters to obtain a point cloud of the target scene at the second time.
[0063] It should be noted that the target point cloud simulation model is a depth map corresponding to a static point cloud of a target scene, and the point cloud simulation device adjusts the depth information corresponding to the target object point cloud in the depth map corresponding to the static point cloud of the target scene based on the input pose, so that the pose of the target object in the point cloud of the target scene obtained by reflecting the adjusted depth map based on the input camera intrinsic parameters is consistent with the input pose, in other words, the point cloud simulation device can dynamically adjust the pose of the target object in the point cloud of the target scene, thereby realizing dynamic and real-time simulation of the point cloud of the target scene.
[0064] It can be seen that in the scheme of the embodiment, during training, the plurality of first point cloud data sequences are virtually converted into a plurality of first depth maps, and then the initial point cloud simulation model is trained using the first depth maps, thereby realizing joint optimization of the depth information, reflection intensity and loss rate of the point cloud, and when the trained point cloud simulation model is used for rendering of the point cloud, real-time point cloud simulation is facilitated, and the simulation accuracy of the point cloud is improved. Moreover, the conversion of the first point cloud data into the first depth maps used for training of the initial point cloud simulation model is based on the camera intrinsic parameters, so that the trained point cloud simulation model can be applied to simulation of point clouds of different types of laser radars. By decoupling the point cloud simulation model into a background point cloud part, a road point cloud part and a target object point cloud part, subsequent training of the road point cloud part and the target object point cloud part is facilitated, and the accuracy of the point cloud simulation model is improved.
[0065] Referring to Figure 3 , Figure 3 A flowchart of a point cloud simulation model training method provided by an embodiment of the present application is shown in FIG. 1. Figure 3 As shown in FIG. 1, the method comprises the following steps.
[0066] S301, a training device obtains a plurality of first point cloud data at a plurality of different positions in a target scene.
[0067] Optionally, the first point cloud data can be obtained by a laser radar, or can be obtained by other means, for example, the point cloud data is obtained by a laser radar in another device, and then the point cloud data is obtained from the other device.
[0068] S302, the training device constructs an initial point cloud simulation model based on the plurality of first point cloud data and determines a plurality of pose information of a target object.
[0069] The target object is an object in the plurality of first point cloud data, and further, the target object is a moving object in the plurality of first point cloud data, such as a person, a vehicle, an animal, etc. The parameters of the initial point cloud simulation model include the reflection intensity and the loss rate of the Gaussian sphere corresponding to the points in the point cloud.
[0070] In a feasible implementation, the training device determines the point cloud data of the target object in the plurality of first point cloud data, removes the point cloud data of the target object from each first point cloud data to obtain a plurality of second point cloud data, and determines the pose information of the target object in each first point cloud data, that is, the training device determines a plurality of pose information of the target object, and the plurality of pose information correspond to the plurality of first point cloud data; the training device constructs a point cloud map of the target scene based on the plurality of second point cloud data; the training device determines the initial parameters of the background point cloud and the initial parameters of the road point cloud based on the point cloud map of the target scene; the training device obtains mesh data of the target object based on the point cloud data of the target object in the plurality of first point cloud data; the training device performs down-sampling on the mesh data of the target object to obtain the initial parameters of the target object point cloud; and the initial point cloud simulation model includes the initial parameters of the background point cloud, the initial parameters of the target object point cloud, and the initial parameters of the road point cloud.
[0071] Exemplarily, the training device processes the plurality of first point cloud data by a target detection-based tracking algorithm to determine the point cloud data of the target object in the plurality of first point cloud data. Optionally, the target detection-based tracking algorithm can be a joint detection and embedding (JDE) algorithm or a simple online and real-time tracking (SORT) algorithm, and of course can also be other algorithms, which are not limited herein.
[0072] For example, the training device detects target objects in the plurality of first point cloud data based on a target detection algorithm to determine the target objects in the first point cloud data, and tracks the target objects in the plurality of first point cloud data based on a target tracking algorithm to determine pose information of the target objects in the plurality of first point cloud data. Optionally, the target detection algorithm can be a YOLO algorithm, a Single Shot Multibox Detector (SSD) algorithm, or an R-CNN algorithm, and can also be other algorithms, which are not limited herein. Optionally, the target tracking algorithm can be a SORT algorithm based on deep learning or a BoT-SORT algorithm, and can also be other algorithms, which are not limited herein.
[0073] Optionally, the training device respectively down-samples the plurality of second point cloud data to obtain a plurality of fourth point cloud data, and constructs a point cloud map of the target scene based on the plurality of fourth point cloud data.
[0074] By down-sampling the second point cloud data, the data amount of the point cloud data can be reduced, which is conducive to improving the efficiency of constructing the point cloud map of the target scene and reducing storage costs.
[0075] In a feasible implementation, the initial parameters of the background point cloud include parameters of Gaussian spheres corresponding to all points in the background point cloud, the initial parameters of the target object point cloud include parameters of Gaussian spheres corresponding to all points in the target object point cloud, and the main axis normal of the Gaussian sphere corresponding to all points in the road point cloud is consistent with the direction of the road in the target scene.
[0076] For the Gaussian spheres corresponding to all points in the background point cloud, the Gaussian spheres corresponding to all points in the target object point cloud, and the Gaussian spheres corresponding to all points in the road point cloud, the parameters of each Gaussian sphere include a location parameter, a rotation parameter, an opacity, a scale parameter, an intensity, a raydrop, and a materialvec parameter.
[0077] It should be noted that the reflection intensity parameter of the Gaussian sphere corresponding to the point in the point cloud is related to the surface material of the object, and the incident direction and distance of the pixel ray, and the training device can determine the reflection intensity of the Gaussian sphere corresponding to the point in the point cloud based on the surface material of the object, and the incident direction and distance of the pixel ray. Specifically, the training device processes the features obtained by concatenating d, γ(x) and materialvec using a multilayer perceptron (MLP) to obtain the reflection intensity of the Gaussian sphere wherein d is the incident direction of the ray, γ(x) is the distance of the ray, and materialvec is used to represent the surface material characteristics of the object.
[0078] It should be noted that the points in the point cloud correspond to the pixels in the depth map. The depth map in this embodiment can be regarded as a three-channel image, and the three channels respectively represent depth, reflection intensity and loss rate.
[0079] In one example, the reflection intensity parameter of the Gaussian sphere corresponding to the i-th point in the point cloud (including the background point cloud, the target object point cloud or the road point cloud) can be represented as:
[0080]
[0081] wherein α i is the opacity of the Gaussian sphere corresponding to the i-th point, α j is the opacity of the Gaussian sphere corresponding to the j-th point, is the reflection intensity value of the pixel corresponding to the i-th point obtained by integration, T i is the weight of the reflection intensity of the Gaussian sphere corresponding to the i-th point. The reflection intensity attribute of each point in the point cloud is obtained by Gaussian rasterization based on Gaussian splashing, that is, the reflection intensity of each point in the point cloud is obtained by accumulating the values of the reflection intensity parameters of the Gaussian spheres through which the rays of the pixel corresponding to the point pass.
[0082] The loss rate of the Gaussian sphere corresponding to each point in the point cloud is related to the incident direction and reflection intensity of the ray of the pixel corresponding to the point, and the training device determines the loss rate of the Gaussian sphere corresponding to each point in the point cloud based on the incident direction and reflection intensity of the ray. Specifically, the training device processes the features obtained by concatenating the incident direction and reflection intensity of the ray using an MLP to obtain the loss rate of the Gaussian sphere corresponding to each point.
[0083] In one example, the loss rate of the Gaussian sphere corresponding to the i-th point in the point cloud can be represented as:
[0084]
[0085] The loss rate attribute of each point in the point cloud is obtained by Gaussian rasterization based on Gaussian splashing, that is, the loss rate of each point in the point cloud is obtained by accumulating the loss rate parameter values of the Gaussian sphere through which the rays of the pixel corresponding to the point pass. Among them, in order to make the distribution of the simulated point cloud more in line with the actual situation, the points in the point cloud whose loss rate is lower than the preset loss rate are discarded.
[0086] It should be understood that the reflection intensity of the Gaussian sphere corresponding to the i-th point refers to the signal intensity reflected back by the Gaussian sphere through which the rays of the pixel corresponding to the i-th point pass, and the loss rate of the Gaussian sphere corresponding to the i-th point refers to the probability that the signal rays reflected back by the rays passing through the object through the Gaussian sphere do not pass through the Gaussian sphere.
[0087] For the Gaussian sphere corresponding to each point in the road point cloud, in addition to setting the reflection intensity and loss rate of the Gaussian sphere in the above manner, the principal axis normal of the Gaussian sphere corresponding to all points in the road point cloud is made consistent with the direction of the road in the target scene, Figure 4a The left drawing of FIG. 1 schematically shows the case where the normal of the Gaussian sphere is inconsistent with the direction of the road, Figure 4a The right drawing of FIG. 1 schematically shows the case where the normal of the Gaussian sphere is consistent with the direction of the road, and by constraining the size of the Gaussian sphere, the road is reconstructed as a flat shell, as Figure 4b shown, ensuring the accuracy of the road perspective geometry.
[0088] S303, the training device converts the plurality of first point cloud data into corresponding plurality of first depth images based on the camera intrinsic parameters.
[0089] In a feasible implementation manner, the point cloud simulation device will project the first point cloud data head into the first depth image based on the camera intrinsic parameters. In an example, the camera intrinsic parameters are (k x , k y , W, H), where (k x , k y ) is the focal length of the camera, and W and H are the size of the first depth image. It should be noted that in order to retain as much point cloud data information as possible in the first depth image, appropriate camera intrinsic parameters need to be used. In order to obtain appropriate camera intrinsic parameters, the camera intrinsic parameters are taken as optimization parameters, and the optimization target is that the number of points in the point cloud data projected onto a unique pixel is maximum. Based on this idea, the appropriate camera intrinsic parameters can be determined by a nonlinear optimization method.
[0090] In an example, the nonlinear optimization method can be represented as:
[0091] P 2D = F(k x , k y , W, H, P 3D ),
[0092] P 2D argmax(count(P 2D ))
[0093] wherein P 3D is a three-dimensional coordinate of a point in the point cloud data, and P 2D is a two-dimensional coordinate of a pixel in the depth map.
[0094] It can be seen that the point cloud data is projected into the depth map by the camera intrinsic parameter, and in the projection process, the number of points in the point cloud data projected onto a unique pixel is maximized by the set camera intrinsic parameter, so as to retain as much point cloud data information as possible. Compared with directly projecting the point cloud data into a distance image, the information loss of the point cloud data is reduced, and the application is also more compatible with the Gaussian rasterization point cloud simulation framework.
[0095] S304, the training device trains an initial point cloud simulation model based on the plurality of pose information and the plurality of first depth maps to obtain a target point cloud simulation model.
[0096] wherein the parameters of the target point cloud simulation model include target parameters of the target background point cloud, target parameters of the background point cloud, and target parameters of the road point cloud.
[0097] In a feasible implementation manner, the training device performs hybrid space-time coding based on the plurality of pose information, the plurality of first time instants, and the initial parameters of the target object point cloud to obtain the reflection intensity offset of the plurality of first time instants and the loss rate offset of the plurality of first time instants. The plurality of first time instants are respectively the collection time instants of the plurality of first point cloud data.
[0098] In an example, the training device encodes the plurality of pose information to obtain pose features, and encodes the plurality of first time instants to obtain time features. The training device extracts three triplane features based on the plurality of first time instants and the plurality of pose information. The training device processes the pose features, the time features, and the three triplane features to obtain the reflection intensity offset of the plurality of first time instants and the loss rate offset of the plurality of first time instants.
[0099] Specifically, as Figure 5As shown, the training device encodes the plurality of pose information (i.e., pose information) through the pose encoding layer to obtain the pose feature; encodes the plurality of first time (i.e., time information) through the time encoding layer to obtain the time feature; the training device encodes the initial parameters of the target object point cloud and the plurality of first time through the space-time embedding layer to obtain three three-dimensional plane features of the target object; the three three-dimensional plane features can be represented as (x, y, t), (y, z, t) and (x, z, t) respectively; wherein (x, y, z) represents the spatial coordinates, and t represents the time. Wherein, the target object can be N, N is an integer greater than 1. The training device concatenates the three three-dimensional plane features of the target object, the time feature and the pose feature to obtain the concatenated feature; the training device processes the concatenated feature using a multilayer perception reader to obtain the reflection intensity offset of the plurality of first time and the loss rate offset of the plurality of time.
[0100] The training device obtains the first parameters of the target object point cloud at the plurality of first time based on the reflection intensity offset of the plurality of first time, the loss rate offset of the plurality of first time, and the initial parameters of the target object point cloud, wherein the first parameters include the first reflection intensity and the first loss rate, the first reflection intensity is determined based on the reflection intensity in the initial parameters of the target point cloud object and the reflection intensity offset, the first time corresponding to the first reflection intensity is the same as the first time corresponding to the reflection intensity offset, the first loss rate is determined based on the loss rate in the initial parameters of the target point cloud object and the loss rate offset, and the first time corresponding to the first loss rate is the same as the first time corresponding to the loss rate offset.
[0101] The training device performs Gaussian rasterization processing on the plurality of third point cloud data to obtain a plurality of second depth maps; the plurality of third point cloud data corresponds to the plurality of first time, and each third point cloud data includes the initial parameters of the background point cloud, the initial parameters of the road point cloud, the second parameters and the first parameters of the same time as the first time corresponding to the third point cloud data; the second parameters include parameters in the initial parameters of the target point cloud object except the reflection intensity and the loss rate.
[0102] For example, the initial parameters of the target object point cloud include initial reflection intensity, initial loss rate and other initial parameters; the reflection intensity offsets of the three time points (including the reflection intensity offset of the t1 time point, the reflection intensity offset of the t2 time point and the reflection intensity offset of the t3 time point) and the loss rate offsets of the three time points (including the loss rate offset of the t1 time point, the loss rate offset of the t2 time point and the loss rate offset of the t3 time point) are obtained in the above manner, and the training device determines the reflection intensity of the t1 time point, the reflection intensity of the t2 time point and the reflection intensity of the t3 time point, i.e. the first reflection intensity of the three time points, based on the initial reflection intensity in the initial parameters of the target object point cloud and the reflection intensity offsets of the t1 time point, the t2 time point and the t3 time point respectively; the training device determines the loss rate of the t1 time point, the loss rate of the t2 time point and the loss rate of the t3 time point, i.e. the first loss rate of the three time points, based on the initial loss rate in the initial parameters of the target object point cloud and the loss rate offsets of the t1 time point, the t2 time point and the t3 time point respectively.
[0103] The training device obtains three third point cloud data based on the first reflection intensity of the three time points, the first loss rate of the three time points and the second parameter, the initial parameters of the background point cloud and the initial parameters of the road point cloud, and the three third point cloud data correspond to the three time points, wherein the third point cloud data of the three time points all include the second parameter, the initial parameters of the background point cloud and the initial parameters of the road point cloud, and the difference is that the first reflection intensity and the first loss rate of the corresponding time point are further included in the third point cloud data of the three time points respectively.
[0104] The training device calculates the depth loss value, the reflection intensity loss value and the loss rate loss value based on the plurality of first depth maps and the plurality of second depth maps. The training device adjusts the initial parameters of the background point cloud, the initial parameters of the target object point cloud and the initial parameters of the road point cloud based on the depth loss value, the reflection intensity loss value and the loss rate loss value to obtain the target parameters of the background point cloud, the target parameters of the target object point cloud and the target parameters of the road point cloud.
[0105] Optionally, the training device calculates a comprehensive loss value based on the depth loss value, the reflection intensity loss value and the loss rate loss value, such as weighted summation, etc., and adjusts the initial parameters of the background point cloud, the initial parameters of the target object point cloud and the initial parameters of the road point cloud based on the comprehensive loss value to obtain the target parameters of the background point cloud, the target parameters of the target object point cloud and the target parameters of the road point cloud.
[0106] It should be pointed out here that the above training process can be performed multiple times in an iterative manner until the number of training times reaches a number threshold or the above loss values converge.
[0107] It can be seen that, during training, by virtually converting the plurality of first point cloud data sequences into a plurality of first depth maps, and then training the initial point cloud simulation model using the first depth maps, the joint optimization of the depth information, reflection intensity and loss rate of the point cloud is realized. When rendering the point cloud using the trained point cloud simulation model, it is beneficial to realize real-time point cloud simulation and improve the simulation accuracy of the point cloud. Moreover, the conversion of the first point cloud data into the first depth map used to train the initial point cloud simulation model is based on the camera intrinsic parameters, so that the trained point cloud simulation model can be applied to the simulation of point clouds of different models of laser radars. By decoupling the point cloud simulation model into a background point cloud part, a road point cloud part and a target object point cloud part, it is convenient to subsequently optimize the road point cloud part and the target object point cloud part respectively during training, thereby facilitating the improvement of the accuracy of the point cloud simulation model. Moreover, based on the mesh data of the target object obtained from the point cloud data of the target object in the plurality of first point cloud data, the point cloud parameters of the target object are initialized using the mesh data of the target object, thereby solving the problem of inaccurate geometric reconstruction and penetration of the target object point cloud due to the sparsity of the target object point cloud in the first point cloud data. By introducing the offset of the reflection intensity and the loss rate of the target object point cloud corresponding to the reflection intensity and the loss rate of the target object point cloud at different time points through hybrid space-time coding, and adjusting the reflection intensity and the loss rate in the initial parameters of the target object point cloud based on the offset of the reflection intensity and the offset of the loss rate, the problem of inconsistent reflection intensity and loss rate of the target object point cloud due to the different positions / locations of the target object at different time points in the continuous frames (i.e. the plurality of first point cloud data) is overcome, thereby ensuring the simulation accuracy of the target object point cloud in a dynamic scene. During training of the point cloud simulation model, the training device adjusts the parameters of the initial point cloud simulation model based on the depth loss value, the reflection intensity loss value and the loss rate loss value, thereby facilitating the improvement of the accuracy of the point cloud simulation model.
[0108] The inference process based on the point cloud simulation model is introduced below.
[0109] The point cloud simulation device obtains the pose information of the target object and the camera intrinsic parameters; and processes based on the point cloud simulation model and the pose information of the target object to obtain a third depth map corresponding to the second time; and processes the third depth map corresponding to the second time based on the camera intrinsic parameters to obtain the point cloud of the target scene at the second time.
[0110] It should be noted that the target point cloud simulation model is a depth map corresponding to the static point cloud of the target scene, and the point cloud simulation device adjusts the depth information corresponding to the target object point cloud in the depth map corresponding to the static point cloud of the target scene based on the input pose, so that the pose of the target object in the point cloud of the target scene obtained by reflecting the adjusted depth map based on the input camera intrinsic parameter is consistent with the input pose, in other words, the point cloud simulation device can dynamically adjust the pose of the target object in the point cloud of the target scene, or input a pose sequence, thereby realizing dynamic real-time simulation of the point cloud of the target scene. And by changing the input pose or pose sequence, point cloud simulation at different angles can be realized. During training and simulation, the camera intrinsic parameters used are different for point cloud data corresponding to different types of laser devices, so simulation of point cloud data corresponding to different types of laser devices can be realized by changing the camera intrinsic parameters.
[0111] It should be noted that the point cloud data in the present application can be point cloud data corresponding to a laser radar, or point cloud data corresponding to a millimeter wave radar. The scheme of the present application can be applied to fields that need to generate point clouds, such as the field of autonomous driving, the field of AR\VR, the field of high-precision surveying and mapping, and the field of three-dimensional scene reconstruction.
[0112] The beneficial effects produced by the present application will be described below in conjunction with the accompanying drawings.
[0113] Figure 6 A rendering effect comparison diagram for a static scene is shown. Figure 6 a diagram in the figure is a point cloud diagram corresponding to first point cloud data, Figure 6 b diagram in the figure is a point cloud diagram corresponding to point data obtained by using the scheme of the present application, Figure 6 c diagram in the figure is a point cloud diagram corresponding to point cloud obtained by using a NeRF method, Figure 6 d diagram in the figure is a point cloud diagram corresponding to point cloud obtained by a point cloud generator (PCGen), Figure 6 a diagram in the figure is a reference, and the performance of the scheme of the present application, the NeRF method and the PCGen is shown in Table 1 as follows:
[0114] The scheme of the present application NeRF approach PCGen Chamfer distance 0.12m 0.25m 0.34m Frame rate 86.5 1.0 8.3
[0115] Table 1
[0116] It should be noted that the chamfer distance is used to represent the similarity of two point cloud diagrams, and the smaller the chamfer distance, the higher the similarity of the two point cloud diagrams. As shown in Table 1, the point cloud diagram obtained by using the scheme of the present application is most similar to the reference point cloud diagram, and the rendering frame rate of the scheme of the present application is the highest. It can be seen that in terms of point cloud rendering precision and rendering efficiency, the scheme of the present application is superior to the other two schemes.
[0117] Figure 7 Fig. 1 shows a schematic diagram of a point cloud before and after removing a target object. Figure 7 Fig. 1a and Fig. 1c are point cloud diagrams before removing the target object, Figure 7 Fig. 1b and Fig. 1d are point cloud diagrams after removing the target object, which can be achieved by using the scheme of the present application. As can be seen, the scheme of the present application can remove the target object from the point cloud data.
[0118] Figure 8 Fig. 2 shows point cloud diagrams of different perspectives of a road scene. Among them, Figure 8 Fig. 2a is a point cloud diagram of the original perspective of the road scene rendered according to the scheme of the present application, Figure 8 Fig. 2b and Fig. 2c are respectively the left lane change and right lane change point cloud diagrams corresponding to Fig. 2a rendered according to the scheme of the present application. As can be seen, the scheme of the present application can render point clouds of different perspectives for the same scene.
[0119] Figure 9 Fig. 3 shows a point cloud simulation diagram corresponding to different types of laser devices. Among them, Figure 9 Fig. 3a is a point cloud diagram of the original perspective of the road scene, Fig. 3b is a point cloud diagram corresponding to the simulated M2 laser, and Fig. 3c is a right lane change point cloud diagram corresponding to the simulated M2 laser.
[0120] Referring to Fig. 4, it is a structural schematic diagram of a training device provided by an embodiment of the present application. As shown in Fig. 4, Figure 10 As shown in Fig. 4, the training device 1000 comprises: Figure 10
[0121] An acquisition unit 1001 is configured to acquire a plurality of first point cloud data of a plurality of different positions in a target scene.
[0122] A determination unit 1002 is configured to construct an initial point cloud simulation model based on the plurality of first point cloud data and determine a plurality of pose information of a target object. The parameters of the point cloud simulation model include the reflection intensity and loss rate of the Gaussian sphere corresponding to the points in the point cloud. The target object is an object in the plurality of first point cloud data. The plurality of first point cloud data are converted into a plurality of first depth maps based on the camera intrinsic parameters.
[0123] A training unit 1003 is configured to train the initial point cloud simulation model based on the plurality of pose information and the plurality of first depth maps to obtain a target point cloud simulation model.
[0124] In a feasible implementation manner, the parameters of the initial point cloud simulation model include initial parameters of background point cloud, initial parameters of target object point cloud and initial parameters of road point cloud. In terms of constructing the initial point cloud simulation model based on the plurality of first point cloud data, the determination unit 1002 is configured to:
[0125] The point cloud data of the target object in the plurality of first point cloud data is determined, the point cloud data of the target object is removed from each first point cloud data to obtain a plurality of second point cloud data; a point cloud map of the target scene is constructed based on the plurality of second point cloud data; initial parameters of the background point cloud and initial parameters of the road point cloud are determined based on the point cloud map of the target scene; mesh data of the target object is obtained based on the point cloud data of the target object in the plurality of first point cloud data; and the mesh data of the target object is down-sampled to obtain initial parameters of the target object point cloud.
[0126] In a feasible implementation, the initial parameters of the background point cloud include parameters of Gaussian spheres corresponding to all points in the background point cloud, the initial parameters of the target object point cloud include parameters of Gaussian spheres corresponding to all points in the target object point cloud, and the initial parameters of the road point cloud include parameters of Gaussian spheres corresponding to all points in the road point cloud, and a principal axis normal of the Gaussian sphere corresponding to all points in the road point cloud is consistent with a direction of the road in the target scene; the parameters of the Gaussian sphere include a position parameter, a rotation parameter, an opacity, a scale parameter, a reflection intensity, a loss rate, and a material characteristic parameter.
[0127] In a feasible implementation, the training unit 1003 is configured to:
[0128] The initial parameters of the target object point cloud are mixed based on the plurality of pose information, the plurality of time points and the initial parameters of the target object point cloud to obtain a plurality of time point reflection intensity offsets and a plurality of time point loss rate offsets; the plurality of time points are respectively collection time points of the plurality of first point cloud data; the first parameters of the target object point cloud at the plurality of time points are obtained based on the plurality of time point reflection intensity offsets, the plurality of time point loss rate offsets and the initial parameters of the target object point cloud, wherein the first parameters include a first reflection intensity and a first loss rate, the first reflection intensity is determined based on the reflection intensity in the initial parameters of the target object point cloud and the reflection intensity offset, the time point corresponding to the first reflection intensity is the same as the time point corresponding to the reflection intensity offset, the first loss rate is determined based on the loss rate in the initial parameters of the target object point cloud and the loss rate offset, and the time point corresponding to the first loss rate is the same as the time point corresponding to the loss rate offset; the plurality of third point cloud data are subjected to Gaussian rasterization processing to obtain a plurality of second depth maps; the plurality of third point cloud data correspond to the plurality of time points, and each of the third point cloud data includes the initial parameters of the background point cloud, the initial parameters of the road point cloud, the second parameters, and the first parameters of the same time point as the time point corresponding to the third point cloud data; the second parameters include parameters other than the reflection intensity and the loss rate in the initial parameters of the target object point cloud; the depth loss value, the reflection intensity loss value and the loss rate loss value are calculated based on the plurality of first depth maps and the plurality of second depth maps; and the initial parameters of the background point cloud, the initial parameters of the target object point cloud and the initial parameters of the road point cloud are adjusted based on the depth loss value, the reflection intensity loss value and the loss rate loss value to obtain target parameters of the background point cloud, target parameters of the target object point cloud and target parameters of the road point cloud, and the parameters of the target point cloud simulation model include the target parameters of the background point cloud, the target parameters of the target object point cloud and the target parameters of the road point cloud.
[0129] In a feasible implementation, in the aspect of mixing space-time coding based on the plurality of pose information, the plurality of time points and the initial parameters of the target object point cloud to obtain the plurality of time point reflection intensity offsets and the plurality of time point loss rate offsets, the training unit 1003 is specifically configured to:
[0130] The plurality of pose information is coded to obtain pose features; the plurality of time points is coded to obtain time features; three three-dimensional plane features are extracted based on the initial parameters of the target object point cloud, the plurality of time points and the plurality of pose information; and the plurality of time point reflection intensity offsets and the plurality of time point loss rate offsets are obtained by processing the pose features, the time features and the three three-dimensional plane features.
[0131] In a feasible implementation, in the aspect of converting the plurality of first point cloud data into the plurality of first depth maps based on the camera intrinsic parameters, the determination unit 1002 is specifically configured to:
[0132] Project the first point cloud data into a first depth map based on the camera intrinsic parameters, wherein the number of points in the first point cloud data that project onto a unique pixel is maximum.
[0133] It is worth pointing out that the specific implementation of the training device 1000 is described above in relation to the point cloud simulation method, such as the acquisition unit 1001 for performing the related content of S301, the determination unit 1002 for performing the related content of S302 and S303, and the training unit 1003 for performing the related content of S304. The various units or modules in the training device 1000 can be combined into one or several other units or modules respectively or all, or some of the units or modules therein can be further split into a plurality of units or modules with smaller functions to constitute, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of the present application. The above-mentioned units or modules are divided based on logical functions. In actual application, the function of one unit (or module) is realized by a plurality of units (or modules), or the functions of a plurality of units (or modules) are realized by one unit (or module).
[0134] Based on the description of the above method embodiments and related device embodiments, please refer to Figure 11 The embodiments of the present application also provide a structure diagram of a training device 1100. Figure 11 The training device 1100 shown includes a memory 1101, a processor 1102, a communication interface 1103, and a bus 1104. Among them, the memory 1101, the processor 1102, and the communication interface 1103 are communicatively connected to each other through the bus 1104.
[0135] Optionally, the memory 1101 is a ROM, a static storage device, a dynamic storage device, or a RAM.
[0136] The memory 1101 can store programs, and when the programs stored in the memory 1101 are executed by the processor 1102, the processor 1102 and the communication interface 1103 are used to execute Figure 3 The various steps of the point cloud simulation model training method of the embodiment shown.
[0137] The processor 1102 adopts a general-purpose CPU, a microprocessor, an application-specific integrated circuit ASIC, a GPU, or one or more integrated circuits, and is used to execute related programs to implement Figure 3 The training method of the point cloud simulation model of the embodiment shown.
[0138] The processor 1102 can also be an integrated circuit chip having a processing capability for signals. In implementation, each step of the traffic scheduling of the present application can be completed by integrated logic circuitry of hardware or instructions in the form of software in the processor 1102. Alternatively, the processor 1102 is a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The processor can implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the present application. The general-purpose processor is a microprocessor or any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware coding processor for execution, or a combination of hardware and software modules in the coding processor for execution. The software module is located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, or other mature storage medium in the art. The storage medium is located in the memory 1101, and the processor 1102 reads information in the memory 1101, and combines hardware to complete the functions required by the units included in the training device 1000 in the embodiments of the present application, or executes the training method of the point cloud simulation model in the method embodiments of the present application.
[0139] The communication interface 1103 uses a transceiver-related device such as but not limited to a transceiver to realize communication between the training device 1100 and other devices or communication networks.
[0140] The bus 1104 can include a path for transmitting information between various components (e.g., the memory 1101, the processor 1102, the communication interface 1103) of the training device 1100.
[0141] It should be noted that although Figure 11 The training device 1100 shown only shows the memory, the processor, and the communication interface, but in specific implementation, those skilled in the art should understand that the training device 1100 also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the training device 1100 can also include hardware devices for realizing other additional functions. In addition, those skilled in the art should understand that the training device 1100 can also only include devices necessary for the embodiments of the present application, and does not have to include all the devices shown in the above. Figure 7 At the same time, those skilled in the art should understand that the training device 1100 can also include other devices necessary for the embodiments of the present application.
[0142] The embodiments of the present application also provide a chip, which includes a processor and a data interface, and the processor reads instructions stored on the memory through the data interface to realize the point cloud simulation model training method of the embodiments of the present application.
[0143] Optionally, as an implementation manner, the chip further comprises a memory, and the memory stores instructions, and the processor is configured to execute the instructions stored in the memory, and when the instructions are executed, the processor is configured to execute the point cloud simulation model training method.
[0144] The embodiments of the present application further provide a computer readable storage medium, which stores instructions, and when the instructions are executed on a computer or a processor, the computer or the processor executes one or more steps in any one of the above methods.
[0145] The embodiments of the present application further provide a computer program product comprising instructions, and when the computer program product is executed on a computer or a processor, the computer or the processor executes one or more steps in any one of the above methods.
[0146] Those skilled in the art will appreciate that the functions described with respect to the various illustrative logical blocks, modules, and algorithm steps described in this specification can be implemented by hardware, software, firmware, or any combination thereof. If implemented in software, the functions described with respect to the various illustrative logical blocks, modules, and steps described in this specification can be stored as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of the computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally can correspond to (1) tangible computer- readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this specification. A computer program product can include a computer-readable medium.
[0147] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code means in the form of instructions or data structures and that can be accessed by a computer. Also, any
[0148] Instructions can be executed by one or more processors, such as one or more DSPs, general purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein can refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0149] In several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method can be implemented in other ways. For example, the division of the units is only a logical function division. In actual implementation, there can be another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. Alternatively, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices, or units, such as electrical, mechanical, or other forms.
[0150] Optionally, the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., located in one place, or distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0151] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, all or part can be realized in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed on a computer, the computer program instructions produce all or part of the processes or functions according to the embodiments of the present application.
[0152] The above is only a specific implementation of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto, any change or replacement within the technical scope disclosed by the embodiments of the present application should be covered within the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.
Claims
1. A point cloud simulation model training method, characterized in that, The method comprises: acquiring a plurality of first point cloud data of a plurality of different positions in a target scene; constructing an initial point cloud simulation model based on the plurality of first point cloud data and determining a plurality of pose information of a target object, parameters of the point cloud simulation model including reflection intensity and loss rate of a Gaussian sphere corresponding to a point in the point cloud; the target object being an object in the plurality of first point cloud data; converting the plurality of first point cloud data into a plurality of first depth maps based on camera intrinsic parameters; training the initial point cloud simulation model based on the plurality of pose information and the plurality of first depth maps to obtain a target point cloud simulation model.
2. The method of claim 1, wherein, The parameters of the initial point cloud simulation model include initial parameters of a background point cloud, initial parameters of a target object point cloud, and initial parameters of a road point cloud; The method comprises: determining point cloud data of the target object in the plurality of first point cloud data, and removing the point cloud data of the target object from each first point cloud data to obtain a plurality of second point cloud data; constructing a point cloud map of the target scene based on the plurality of second point cloud data, and determining the initial parameters of the background point cloud and the initial parameters of the road point cloud based on the point cloud map of the target scene; obtaining mesh data of the target object based on the point cloud data of the target object in the plurality of first point cloud data; down-sampling the mesh data of the target object to obtain the initial parameters of the target object point cloud.
3. The method of claim 2, wherein, The initial parameters of the background point cloud include parameters of a Gaussian sphere corresponding to all points in the background point cloud, the initial parameters of the target object point cloud include parameters of a Gaussian sphere corresponding to all points in the target object point cloud, and the initial parameters of the road point cloud include parameters of a Gaussian sphere corresponding to all points in the road point cloud, a principal axis normal of the Gaussian sphere corresponding to all points in the road point cloud being consistent with a direction of a road in the target scene; the parameters of the Gaussian sphere including position parameters, rotation parameters, opacity, scale parameters, reflection intensity, loss rate, and material characteristic parameters.
4. The method of claim 3, wherein, The method comprises: performing mixed space-time coding based on the plurality of pose information, a plurality of time points, and the initial parameters of the target object point cloud to obtain reflection intensity offsets of the plurality of time points and loss rate offsets of the plurality of time points; the plurality of time points being acquisition time points of the plurality of first point cloud data, respectively. The first parameters of the target object point cloud at multiple time points are obtained based on the reflection intensity offset at the multiple time points, the loss rate offset at the multiple time points, and the initial parameters of the target object point cloud, wherein the first parameters include a first reflection intensity and a first loss rate, the first reflection intensity is determined based on the reflection intensity in the initial parameters of the target object point cloud and the reflection intensity offset, the time point corresponding to the first reflection intensity is the same as the time point corresponding to the reflection intensity offset, the first loss rate is determined based on the loss rate in the initial parameters of the target object point cloud and the loss rate offset, and the time point corresponding to the first loss rate is the same as the time point corresponding to the loss rate offset; The third point cloud data are subjected to Gaussian rasterization processing to obtain a plurality of second depth maps; the third point cloud data correspond to the multiple time points, and each of the third point cloud data includes the initial parameters of the background point cloud, the initial parameters of the road point cloud, second parameters, and the first parameters of the same time point as the third point cloud data; the second parameters include parameters in the initial parameters of the target object point cloud except the reflection intensity and the loss rate; The depth loss value, the reflection intensity loss value, and the loss rate loss value are calculated based on the multiple first depth maps and the multiple second depth maps; The initial parameters of the background point cloud, the initial parameters of the target object point cloud, and the initial parameters of the road point cloud are adjusted based on the depth loss value, the reflection intensity loss value, and the loss rate loss value to obtain target parameters of the background point cloud, target parameters of the target object point cloud, and target parameters of the road point cloud, and the parameters of the target point cloud simulation model include the target parameters of the background point cloud, the target parameters of the target object point cloud, and the target parameters of the road point cloud.
5. The method of claim 4, wherein, The reflection intensity offset at the multiple time points and the loss rate offset at the multiple time points are obtained based on the multiple pose information, the multiple time points, and the initial parameters of the target object point cloud, and the reflection intensity offset at the multiple time points and the loss rate offset at the multiple time points are obtained by mixing space-time coding, including: The multiple pose information is coded to obtain pose features, and the multiple time points are coded to obtain time features; Three three-dimensional plane features are extracted from the initial parameters of the target object point cloud based on the multiple time points and the multiple pose information; The reflection intensity offset at the multiple time points and the loss rate offset at the multiple time points are obtained by processing the pose features, the time features, and the three three-dimensional plane features.
6. The method according to any one of claims 1 to 5, characterized in that, The multiple first point cloud data are converted into multiple first depth maps based on the camera intrinsic parameters, including: The first point cloud data are projected into the first depth map based on the camera intrinsic parameters, wherein the number of points in the first point cloud data that are projected onto a unique pixel is the largest.
7. A training device, characterized in that including: An acquisition unit is configured to acquire a plurality of first point cloud data at multiple different positions in a target scene. The determining unit is configured to construct an initial point cloud simulation model based on the plurality of first point cloud data and determine a plurality of pose information of a target object, parameters of the point cloud simulation model including reflection intensity and loss rate of a Gaussian sphere corresponding to a point in the point cloud; the target object is an object in the plurality of first point cloud data; and the plurality of first point cloud data is converted into a plurality of first depth maps based on camera intrinsic parameters. The training unit is configured to train the initial point cloud simulation model based on the plurality of pose information and the plurality of first depth maps to obtain a target point cloud simulation model.
8. The training device of claim 7, wherein, The parameters of the initial point cloud simulation model include initial parameters of a background point cloud, initial parameters of a target object point cloud, and initial parameters of a road point cloud; in the aspect of constructing the initial point cloud simulation model based on the plurality of first point cloud data, the determining unit is configured to: determine point cloud data of the target object in the plurality of first point cloud data, and remove the point cloud data of the target object from each first point cloud data to obtain a plurality of second point cloud data; construct a point cloud map of the target scene based on the plurality of second point cloud data; determine the initial parameters of the background point cloud and the initial parameters of the road point cloud based on the point cloud map of the target scene; obtain mesh data of the target object based on the point cloud data of the target object in the plurality of first point cloud data; down-sample the mesh data of the target object to obtain the initial parameters of the target object point cloud.
9. Training device according to claim 7 or 8, characterized in that, The initial parameters of the background point cloud include parameters of a Gaussian sphere corresponding to all points in the background point cloud, the initial parameters of the target object point cloud include parameters of a Gaussian sphere corresponding to all points in the target object point cloud, and the initial parameters of the road point cloud include parameters of a Gaussian sphere corresponding to all points in the road point cloud, a principal axis normal of the Gaussian sphere corresponding to all points in the road point cloud is consistent with a direction of a road in the target scene; the parameters of the Gaussian sphere include a position parameter, a rotation parameter, an opacity, a scale parameter, a reflection intensity, a loss rate, and a material characteristic parameter.
10. The training device of claim 9, wherein, The training unit is configured to: perform mixed space-time coding based on the plurality of pose information, a plurality of time instants, and the initial parameters of the target object point cloud to obtain reflection intensity offsets of the plurality of time instants and loss rate offsets of the plurality of time instants; the plurality of time instants are respectively collection time instants of the plurality of first point cloud data; obtain first parameters of the target object point cloud at the plurality of time instants based on the reflection intensity offsets of the plurality of time instants, the loss rate offsets of the plurality of time instants, and the initial parameters of the target object point cloud, wherein the first parameters include a first reflection intensity and a first loss rate, the first reflection intensity is determined based on the reflection intensity in the initial parameters of the target object point cloud and the reflection intensity offsets, the time instant corresponding to the first reflection intensity is the same as the time instant corresponding to the reflection intensity offsets, the first loss rate is determined based on the loss rate in the initial parameters of the target object point cloud and the loss rate offsets, and the time instant corresponding to the first loss rate is the same as the time instant corresponding to the loss rate offsets. Gaussian rasterization is performed on a plurality of third point cloud data to obtain a plurality of second depth maps; the plurality of third point cloud data correspond to the plurality of time instants, and each of the third point cloud data includes initial parameters of the background point cloud, initial parameters of the road point cloud, second parameters, and the first parameters of the same time instant as the third point cloud data; the second parameters include parameters other than the reflection intensity and the loss rate in the initial parameters of the target object point cloud; Based on the plurality of first depth maps and the plurality of second depth maps, a depth loss value, a reflection intensity loss value, and a loss rate loss value are calculated; Based on the depth loss value, the reflection intensity loss value, and the loss rate loss value, the initial parameters of the background point cloud, the initial parameters of the target object point cloud, and the initial parameters of the road point cloud are adjusted to obtain target parameters of the background point cloud, target parameters of the target object point cloud, and target parameters of the road point cloud; the parameters of the target point cloud simulation model include the target parameters of the background point cloud, the target parameters of the target object point cloud, and the target parameters of the road point cloud.
11. The training device of claim 10, wherein, In the aspect of mixing space-time coding based on the plurality of pose information, the plurality of time instants, and the initial parameters of the target object point cloud to obtain reflection intensity offsets of the plurality of time instants and loss rate offsets of the plurality of time instants, the training unit is specifically configured to: Encode the plurality of pose information to obtain pose features; Encode the plurality of time instants to obtain time features; Extract three three-dimensional plane features based on the plurality of time instants and the initial parameters of the target object point cloud; Process the pose features, the time features, and the three three-dimensional plane features to obtain the reflection intensity offsets of the plurality of time instants and the loss rate offsets of the plurality of time instants.
12. Training apparatus according to any of claims 7-11, characterized in that, In the aspect of converting the plurality of first point cloud data into the plurality of first depth maps based on the camera intrinsic parameters, the determination unit is specifically configured to: Project the first point cloud data into the first depth map based on the camera intrinsic parameters, wherein the number of points in the first point cloud data that are projected onto a unique pixel is the largest.
13. A training device, characterized by A processor and a memory are included, wherein the memory is configured to store program code, and the processor is configured to execute the program code to implement the method of any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1 to 6.
15. A computer program product, when the computer program product is run on a computer, causes the computer to execute the method of any one of claims 1 to 6.