Point cloud acquisition method and related equipment thereof
By utilizing camera pose and neural network models to process images, simulate 3D point clouds, and fuse them, the high cost problem caused by reliance on LiDAR is solved, achieving high-quality 4D point cloud acquisition, which is suitable for cost-constrained scenarios.
Patent Information
- Application Number
- CN202410465284.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-17
- Publication Date
- 2025-10-24
AI Technical Summary
Existing technologies for acquiring four-dimensional point clouds rely on devices such as LiDAR and IMU, resulting in excessively high hardware costs and making them difficult to implement in cost-constrained scenarios. Furthermore, LiDAR fails in extreme weather or with objects of low reflectivity, making it impossible to successfully acquire three-dimensional point clouds.
By acquiring different camera poses, a neural network model is used to process camera images to simulate 3D point clouds at different times, and then the images are fused to obtain 4D point clouds, thus avoiding dependence on LiDAR.
It reduces the hardware cost of point cloud acquisition, improves feasibility in cost-constrained scenarios, and enhances the quality of four-dimensional point clouds through neural network models.
Smart Images

Figure CN120833435A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of artificial intelligence (AI) technology, and in particular, to a point cloud acquisition method and a related device thereof. BACKGROUND
[0002] In the field of autonomous driving, a vehicle usually needs to utilize a general object detection (GOD) model to perceive and understand the environment around the vehicle, so as to improve the safety of autonomous driving of the vehicle. Four-dimensional point cloud (also referred to as 4D Semantic Occupancy, 4D-SO) as training data of the GOD model plays a crucial role in the performance of the GOD model.
[0003] In the related art, in order to acquire the four-dimensional point cloud, a three-dimensional point cloud is acquired by using a device such as a laser radar, and then the four-dimensional point cloud that can be used to train the GOD model is acquired by using the three-dimensional point cloud. Specifically, the laser radar can be used to acquire data of the surrounding environment at different times, so as to obtain three-dimensional point clouds at different times. Then, the three-dimensional point clouds at different times can be fused to obtain the four-dimensional point cloud. In this way, the GOD model can be trained by using the obtained four-dimensional point cloud.
[0004] Since the above-mentioned method of acquiring the four-dimensional point cloud depends on the laser radar and needs the assistance of devices such as an inertial measurement unit (IMU), the cost of the devices required by the method is too high, which is not conducive to implementation in some cost-limited model training scenarios. SUMMARY
[0005] Embodiments of the present application provide a point cloud acquisition method and a related device thereof, which can reduce the hardware cost required for acquiring the point cloud and are conducive to implementation in some cost-limited scenarios.
[0006] A first aspect of embodiments of the present application provides a point cloud acquisition method, which comprises:
[0007] When it is necessary to acquire a four-dimensional point cloud, a first pose of a camera at a first time and a second pose of the camera at a second time can be acquired. The first time and the second time are two different times, and therefore the first pose and the second pose are two poses associated in time.
[0008] Then, the first pose can be input to the first neural network model to process the first pose by the first neural network model, so as to simulate a first image captured by the camera in the first pose. Then, the first image captured by the camera can be converted to obtain a first three-dimensional point cloud.
[0009] Then, the second pose can be input to the first neural network model to process the second pose by the first neural network model, so as to simulate a second image captured by the camera in the second pose. Then, the second image captured by the camera can be converted to obtain a second three-dimensional point cloud.
[0010] Finally, the first three-dimensional point cloud and the second three-dimensional point cloud can be fused to obtain a first four-dimensional point cloud. Then, the first four-dimensional point cloud can be used as data for subsequent model training.
[0011] As can be seen from the above method, when a four-dimensional point cloud needs to be obtained, a first pose of the camera and a second pose of the camera can be obtained first, the first pose and the second pose are associated in time. Then, the first pose can be input to the first neural network model to process the first pose by the first neural network model to obtain a first image captured by the camera, and the first image can be converted to obtain a first three-dimensional point cloud. Then, the second pose can also be input to the first neural network model to process the second pose by the first neural network model to obtain a second image captured by the camera, and the second image can be converted to obtain a second three-dimensional point cloud. Finally, the first three-dimensional point cloud and the second three-dimensional point cloud can be fused to obtain a first four-dimensional point cloud. In the foregoing process, the neural network model (i.e., the first neural network model mentioned above, for example, a neural implicit field) can be used to process different poses (i.e., the first pose and the second pose mentioned above) of the camera at different times to obtain three-dimensional point clouds corresponding to different poses (which can also be understood as three-dimensional point clouds at different times, i.e., the first three-dimensional point cloud and the second three-dimensional point cloud mentioned above), and then the three-dimensional point clouds corresponding to different poses are used to obtain a four-dimensional point cloud (i.e., the first four-dimensional point cloud mentioned above). This way of obtaining a four-dimensional point cloud does not need to rely on a laser radar, but only needs a camera, which can reduce the hardware cost required for obtaining a point cloud, and is conducive to the implementation of this way of obtaining a four-dimensional point cloud in some cost-limited scenarios.
[0012] In a possible implementation manner, the first image can include a first red / green / blue (RGB) image captured by the camera, a first semantic map corresponding to the first RGB image, and a first depth map corresponding to the first RGB image, where the first RGB image includes a plurality of pixel points (i.e., the first RGB image includes colors of the plurality of pixel points), the first semantic map includes categories of each pixel point in the first RGB image, and the first depth map includes depths of each pixel point in the first RGB image.
[0013] In a possible implementation, the second image can include a second RGB image captured by the camera, a second semantic image corresponding to the second RGB image, and a second depth image corresponding to the second RGB image, where the second RGB image includes a plurality of pixel points (i.e., the second RGB image includes colors of the plurality of pixel points), the second semantic image includes categories of the pixel points in the second RGB image, and the second depth image includes depths of the pixel points in the second RGB image.
[0014] In a possible implementation, the first RGB image includes a first pixel point, the first pixel point is any one of the pixel points in the first RGB image, and the obtaining of the first image captured by the camera in the first pose by using the first neural network model includes: obtaining, by using the first neural network model, a first ray sent to the first pixel point by the camera in the first pose; sampling, by using the first neural network model, the first ray to obtain a plurality of first sampling points; processing, by using the first neural network model, the plurality of first sampling points to obtain colors of the plurality of first sampling points, categories of the plurality of first sampling points, depths of the plurality of first sampling points, and transparencies of the plurality of first sampling points; and obtaining, based on the colors of the plurality of first sampling points, the categories of the plurality of first sampling points, the depths of the plurality of first sampling points, and the transparencies of the plurality of first sampling points, a color of the first pixel point, a category of the first pixel point, and a depth of the first pixel point. In the foregoing implementation, for ease of illustration, any one of the pixel points in the first RGB image, i.e., the first pixel point, is used for illustrative introduction. After obtaining the first pose of the camera, the first pose can be input to the first neural network model. Then, the first neural network model can construct the first ray sent to the first pixel point by the camera in the first pose. Then, the first neural network model can sample the first ray to obtain the plurality of first sampling points on the first ray. After obtaining the plurality of first sampling points, the first neural network model can perform a series of processing on the plurality of first sampling points to obtain the colors of the plurality of first sampling points, the categories of the plurality of first sampling points, the depths of the plurality of first sampling points, and the transparencies of the plurality of first sampling points. Subsequently, the colors of the plurality of first sampling points, the categories of the plurality of first sampling points, the depths of the plurality of first sampling points, and the transparencies of the plurality of first sampling points can be calculated to obtain the color of the first pixel point, the category of the first pixel point, and the depth of the first pixel point. Then, the same operations performed on the first pixel point can also be performed on the remaining pixel points in the first RGB image, and therefore the colors of all the pixel points in the first RGB image, the categories of all the pixel points in the first RGB image, and the depths of all the pixel points in the first RGB image can be finally obtained, which is equivalent to obtaining the first RGB image, the first semantic image, and the first depth image.
[0015] In a possible implementation manner, the second RGB image includes a second pixel point, the second pixel point being any one of the pixels in the second RGB image, and the processing of the second pose by the first neural network model to obtain the second image captured by the camera includes: obtaining, by the first neural network model, a second ray sent to the second pixel point by the camera in the second pose; sampling, by the first neural network model, the second ray to obtain a plurality of second sampling points; processing, by the first neural network model, the plurality of second sampling points to obtain colors of the plurality of second sampling points, categories of the plurality of second sampling points, depths of the plurality of second sampling points, and transparencies of the plurality of second sampling points; and obtaining, based on the colors of the plurality of second sampling points, the categories of the plurality of second sampling points, the depths of the plurality of second sampling points, and the transparencies of the plurality of second sampling points, a color of the second pixel point, a category of the second pixel point, and a depth of the second pixel point. In the foregoing implementation manner, for the convenience of description, any one of the pixels in the second RGB image, i.e., the second pixel point, is used for illustrative introduction. After the second pose of the camera is obtained, the second pose can be input to the second neural network model. Then, the second neural network model can construct a second ray sent to the second pixel point by the camera in the second pose. Then, the second neural network model can sample the second ray to obtain a plurality of second sampling points on the second ray. After the plurality of second sampling points are obtained, the second neural network model can perform a series of processing on the plurality of second sampling points to obtain colors of the plurality of second sampling points, categories of the plurality of second sampling points, depths of the plurality of second sampling points, and transparencies of the plurality of second sampling points. Subsequently, the colors of the plurality of second sampling points, the categories of the plurality of second sampling points, the depths of the plurality of second sampling points, and the transparencies of the plurality of second sampling points can be calculated to obtain a color of the second pixel point, a category of the second pixel point, and a depth of the second pixel point. Then, for the remaining pixels in the second RGB image except the second pixel point, the same operations performed on the second pixel point can also be performed on the remaining pixels, and therefore, the colors of all the pixels in the second RGB image, the categories of all the pixels in the second RGB image, and the depths of all the pixels in the second RGB image can be finally obtained, which is equivalent to obtaining the second RGB image, the second semantic image, and the second depth image.
[0016] In a possible implementation, the color, the category, and the depth of the first pixel point are obtained based on the colors, the categories, the depths, and the transparencies of the first sampling points, including: calculating the colors and the transparencies of the first sampling points to obtain the color of the first pixel point; calculating the categories and the transparencies of the first sampling points to obtain the category of the first pixel point; and calculating the depths and the transparencies of the first sampling points to obtain the depth of the first pixel point. In the foregoing implementation, after the colors, the categories, the depths, and the transparencies of the first sampling points are obtained, the colors, the transparencies, and the distances between the first sampling points can be calculated to obtain the color of the first pixel point. Meanwhile, the categories, the transparencies, and the distances between the first sampling points can be calculated to obtain the category of the first pixel point. Meanwhile, the depths, the transparencies, and the distances between the first sampling points can be calculated to obtain the depth of the first pixel point.
[0017] In a possible implementation, the color, the category, and the depth of the second pixel point are obtained based on the colors, the categories, the depths, and the transparencies of the second sampling points, including: calculating the colors and the transparencies of the second sampling points to obtain the color of the second pixel point; calculating the categories and the transparencies of the second sampling points to obtain the category of the second pixel point; and calculating the depths and the transparencies of the second sampling points to obtain the depth of the second pixel point. In the foregoing implementation, after the colors, the categories, the depths, and the transparencies of the second sampling points are obtained, the colors, the transparencies, and the distances between the second sampling points can be calculated to obtain the color of the second pixel point. Meanwhile, the categories, the transparencies, and the distances between the second sampling points can be calculated to obtain the category of the second pixel point. Meanwhile, the depths, the transparencies, and the distances between the second sampling points can be calculated to obtain the depth of the second pixel point.
[0018] In a possible implementation manner, the method further comprises: calculating the depths of the plurality of first sampling points and the depth of the first pixel point to obtain an evaluation value of the first pixel point; calculating the depths of the plurality of second sampling points and the depth of the second pixel point to obtain an evaluation value of the second pixel point; and fusing the first three-dimensional point cloud and the second three-dimensional point cloud to obtain the first four-dimensional point cloud, comprising: taking the evaluation value of each pixel point in the first RGB image as the evaluation value of each spatial point in the first three-dimensional point cloud, and removing, from the first three-dimensional point cloud, spatial points with an evaluation value less than an evaluation threshold to obtain a denoised first three-dimensional point cloud; taking the evaluation value of each pixel point in the second RGB image as the evaluation value of each spatial point in the second three-dimensional point cloud, and removing, from the second three-dimensional point cloud, spatial points with an evaluation value less than the evaluation threshold to obtain a denoised second three-dimensional point cloud; and fusing the denoised first three-dimensional point cloud and the denoised second three-dimensional point cloud to obtain the first four-dimensional point cloud. In the foregoing implementation manner, after obtaining the depth of the first pixel point, the depths of the plurality of first sampling points and the depth of the first pixel point are calculated to obtain the evaluation value of the first pixel point, and the evaluation value of the first pixel point is used to indicate the quality of the first pixel point. Similarly, for the remaining pixel points of the first RGB image except the first pixel point, the same operation as performed on the first pixel point can also be performed on the remaining pixel points, and therefore the evaluation values of the pixel points in the first RGB image can be finally obtained. Similarly, the evaluation values of the pixel points in the second RGB image can also be obtained. Then, after projecting the first RGB image, the first semantic image and the first depth image to obtain the first three-dimensional point cloud, since each pixel point in the first RGB image is in one-to-one correspondence with each spatial point in the first three-dimensional point cloud, the evaluation values of the pixel points in the first RGB image can be taken as the evaluation values of the spatial points in the first three-dimensional point cloud, and spatial points with an evaluation value lower than the evaluation threshold can be removed from the first three-dimensional point cloud to obtain the denoised first three-dimensional point cloud. Similarly, the denoised second three-dimensional point cloud can also be obtained. In this way, the denoised first three-dimensional point cloud and the denoised second three-dimensional point cloud can be fused to obtain the first four-dimensional point cloud. As can be seen, by obtaining the evaluation values of the pixel points in the RGB image and taking the evaluation values of the pixel points as the evaluation values of the spatial points in the three-dimensional point cloud, the noise points in the three-dimensional point cloud can be effectively removed, and therefore the four-dimensional point cloud obtained based on the denoised three-dimensional point cloud has better quality.
[0019] In a possible implementation, the method further includes: processing the first image and the second image by a second neural network model to obtain a second four-dimensional point cloud; and optimizing the first four-dimensional point cloud based on the second four-dimensional point cloud to obtain an optimized first four-dimensional point cloud. In the foregoing implementation, after the first image and the second image are obtained, the first image and the second image can be input into the second neural network model, so that the first image and the second image are processed by the second neural network model in a series of processes, thereby obtaining the second four-dimensional point cloud. After the second four-dimensional point cloud is obtained, the first four-dimensional point cloud and the second four-dimensional point cloud can be fused, thereby obtaining the optimized first four-dimensional point cloud. As can be seen, for the four-dimensional point cloud obtained based on the three-dimensional point cloud, the model-predicted four-dimensional point cloud can also be obtained, and the model-predicted four-dimensional point cloud can be used to optimize the four-dimensional point cloud obtained based on the three-dimensional point cloud, thereby obtaining a four-dimensional point cloud with better quality.
[0020] In a possible implementation, the first neural network model is a neural implicit field.
[0021] The second aspect of the embodiment of the present application provides a model training method, which includes: processing a pose of a camera by a to-be-trained model to obtain an image captured by the camera, the image being used to obtain a three-dimensional point cloud, and the three-dimensional point cloud being used to obtain a four-dimensional point cloud; obtaining a target loss based on the image and a real image corresponding to the image, the target loss being used to indicate a difference between the image and the real image; updating parameters of the to-be-trained model based on the target loss until a model training condition is met, and obtaining a first neural network model.
[0022] In a possible implementation, the image includes an RGB image captured by the camera, a semantic image corresponding to the RGB image, and a depth image corresponding to the RGB image, where the semantic image includes categories of each pixel point in the RGB image, and the depth image includes depths of each pixel point in the RGB image.
[0023] In a possible implementation, the RGB image includes a target pixel point, the target pixel point being any one pixel point in the RGB image, and processing the pose by the first neural network model to obtain the image captured by the camera includes: obtaining, by the to-be-trained model, a ray sent to the target pixel point by the camera in the pose; sampling, by the to-be-trained model, the ray to obtain a plurality of sampling points; processing, by the to-be-trained model, the plurality of sampling points to obtain colors of the plurality of sampling points, categories of the plurality of sampling points, depths of the plurality of sampling points, and transparencies of the plurality of sampling points; and obtaining, based on the colors of the plurality of sampling points, the categories of the plurality of sampling points, the depths of the plurality of sampling points, and the transparencies of the plurality of sampling points, a color of the target pixel point, a category of the target pixel point, and a depth of the target pixel point.
[0024] In a possible implementation, the color, the category, and the depth of the target pixel point are obtained based on the colors of the plurality of sampling points, the categories of the plurality of sampling points, the depths of the plurality of sampling points, and the transparencies of the plurality of sampling points, including: calculating the colors of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the color of the target pixel point; calculating the categories of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the category of the target pixel point; and calculating the depths of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the depth of the target pixel point.
[0025] In a possible implementation, the real image includes a real RGB image, a real semantic image, and a real depth image, the real RGB image includes real colors of the target pixel points, the real semantic image includes real categories of the target pixel points, and the real depth image includes real depths of the target pixel points; and the target loss is obtained based on the image and the real image corresponding to the image, including: obtaining a first loss based on the color of the target pixel point and the real color of the target pixel point, the first loss being used to indicate a difference between the color of the target pixel point and the real color of the target pixel point; obtaining a second loss based on the category of the target pixel point and the real category of the target pixel point, the second loss being used to indicate a difference between the category of the target pixel point and the real category of the target pixel point; obtaining a third loss based on the depth of the target pixel point and the real depth of the target pixel point, the third loss being used to indicate a difference between the depth of the target pixel point and the real depth of the target pixel point; and obtaining the target loss based on the first loss, the second loss, and the third loss.
[0026] In a possible implementation, the target loss is obtained based on the image and the real image corresponding to the image, further including: obtaining a fourth loss based on the transparency distribution of the plurality of sampling points and a real transparency distribution of the plurality of sampling points; and the target loss is obtained based on the first loss, the second loss, and the third loss, including: obtaining the target loss based on the first loss, the second loss, the third loss, and the fourth loss.
[0027] In a possible implementation, the first neural network model is a neural implicit field.
[0028] A third aspect of an embodiment of the present application provides a point cloud acquisition device, which includes: an acquisition module for acquiring a first pose of a camera and a second pose of the camera, the first pose and the second pose being temporally associated; a first processing module for processing the first pose through a first neural network model to obtain a first image taken by the camera, and converting the first image to obtain a first three-dimensional point cloud; a second processing module for processing the second pose through the first neural network model to obtain a second image taken by the camera, and converting the second image to obtain a second three-dimensional point cloud; and a fusion module for fusing the first three-dimensional point cloud and the second three-dimensional point cloud to obtain a first four-dimensional point cloud.
[0029] In one possible implementation, the first image includes a first RGB image captured by a camera, a first semantic image corresponding to the first RGB image, and a first depth image corresponding to the first RGB image, wherein the first semantic image includes the category of each pixel in the first RGB image, and the first depth map includes the depth of each pixel in the first RGB image.
[0030] In one possible implementation, the second image includes a second RGB image captured by the camera, a second semantic image corresponding to the second RGB image, and a second depth image corresponding to the second RGB image, wherein the second semantic image includes the category of each pixel in the second RGB image, and the second depth map includes the depth of each pixel in the second RGB image.
[0031] In one possible implementation, the first RGB image includes a first pixel point, which is any pixel point in the first RGB image. The first processing module is used to: obtain a first ray sent by the camera in the first pose to the first pixel point through a first neural network model; sample the first ray through the first neural network model to obtain multiple first sampling points; process the multiple first sampling points through the first neural network model to obtain the colors of the multiple first sampling points, the categories of the multiple first sampling points, the depths of the multiple first sampling points, and the transparency of the multiple first sampling points; obtain the color of the first pixel point, the category of the multiple first sampling points, and the transparency of the multiple first sampling points based on the colors of the multiple first sampling points, the categories of the multiple first sampling points, the depths of the multiple first sampling points, and the transparency of the multiple first sampling points.
[0032] In a possible implementation, the second RGB image includes a second pixel point, the second pixel point being any one of the pixel points in the second RGB image, and the second processing module is configured to: acquire, by using the first neural network model, a second ray sent to the second pixel point by the camera in the second pose; sample, by using the first neural network model, the second ray to obtain a plurality of second sampling points; process, by using the first neural network model, the plurality of second sampling points to obtain colors of the plurality of second sampling points, categories of the plurality of second sampling points, depths of the plurality of second sampling points, and transparencies of the plurality of second sampling points; and acquire, based on the colors of the plurality of second sampling points, the categories of the plurality of second sampling points, the depths of the plurality of second sampling points, and the transparencies of the plurality of second sampling points, a color of the second pixel point, a category of the second pixel point, and a depth of the second pixel point.
[0033] In a possible implementation, based on the colors of the plurality of first sampling points, the categories of the plurality of first sampling points, the depths of the plurality of first sampling points, and the transparencies of the plurality of first sampling points, the first processing module is configured to: calculate the colors of the plurality of first sampling points and the transparencies of the plurality of first sampling points to obtain a color of the first pixel point; calculate the categories of the plurality of first sampling points and the transparencies of the plurality of first sampling points to obtain a category of the first pixel point; and calculate the depths of the plurality of first sampling points and the transparencies of the plurality of first sampling points to obtain a depth of the first pixel point.
[0034] In a possible implementation, based on the colors of the plurality of second sampling points, the categories of the plurality of second sampling points, the depths of the plurality of second sampling points, and the transparencies of the plurality of second sampling points, the second processing module is configured to: calculate the colors of the plurality of second sampling points and the transparencies of the plurality of second sampling points to obtain a color of the second pixel point; calculate the categories of the plurality of second sampling points and the transparencies of the plurality of second sampling points to obtain a category of the second pixel point; and calculate the depths of the plurality of second sampling points and the transparencies of the plurality of second sampling points to obtain a depth of the second pixel point.
[0035] In a possible implementation, the apparatus further includes: a first evaluation module, configured to calculate the depth of the plurality of first sampling points and the depth of the first pixel point to obtain an evaluation value of the first pixel point; a second evaluation module, configured to calculate the depth of the plurality of second sampling points and the depth of the second pixel point to obtain an evaluation value of the second pixel point; and a fusion module, configured to: take the evaluation value of each pixel point in the first RGB image as an evaluation value of each spatial point in the first three-dimensional point cloud, remove spatial points with an evaluation value less than an evaluation threshold from the first three-dimensional point cloud to obtain a first denoised three-dimensional point cloud; take the evaluation value of each pixel point in the second RGB image as an evaluation value of each spatial point in the second three-dimensional point cloud, remove spatial points with an evaluation value less than the evaluation threshold from the second three-dimensional point cloud to obtain a second denoised three-dimensional point cloud; and fuse the first denoised three-dimensional point cloud and the second denoised three-dimensional point cloud to obtain the first four-dimensional point cloud.
[0036] In a possible implementation, the apparatus further includes an optimization module, configured to: process the first image and the second image by using the second neural network model to obtain a second four-dimensional point cloud; and optimize the first four-dimensional point cloud based on the second four-dimensional point cloud to obtain an optimized first four-dimensional point cloud.
[0037] In a possible implementation, the first neural network model is a neural implicit field.
[0038] A fourth aspect of the embodiment of the present application provides a model training apparatus, which is characterized in that the apparatus includes: a processing module, configured to process a pose of a camera by using a to-be-trained model to obtain an image captured by the camera, the image being used to obtain a three-dimensional point cloud, and the three-dimensional point cloud being used to obtain a four-dimensional point cloud; an obtaining module, configured to obtain a target loss based on the image and a real image corresponding to the image, the target loss being used to indicate a difference between the image and the real image; and a training module, configured to update a parameter of the to-be-trained model based on the target loss until a model training condition is met to obtain a first neural network model.
[0039] In a possible implementation, the image includes an RGB image captured by the camera, a semantic image corresponding to the RGB image, and a depth image corresponding to the RGB image, where the semantic image includes a category of each pixel point in the RGB image, and the depth image includes a depth of each pixel point in the RGB image.
[0040] In a possible implementation, the RGB image contains a target pixel point, the target pixel point being any one of the pixels in the RGB image, and the processing module is configured to: acquire, by using the to-be-trained model, a ray sent by the camera in the pose to the target pixel point; sample, by using the to-be-trained model, the ray to obtain a plurality of sampling points; process, by using the to-be-trained model, the plurality of sampling points to obtain colors of the plurality of sampling points, categories of the plurality of sampling points, depths of the plurality of sampling points, and transparencies of the plurality of sampling points; and acquire, based on the colors of the plurality of sampling points, the categories of the plurality of sampling points, the depths of the plurality of sampling points, and the transparencies of the plurality of sampling points, a color of the target pixel point, a category of the target pixel point, and a depth of the target pixel point.
[0041] In a possible implementation, the processing module is configured to: calculate the colors of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the color of the target pixel point; calculate the categories of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the category of the target pixel point; and calculate the depths of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the depth of the target pixel point.
[0042] In a possible implementation, the real image contains a real RGB image, a real semantic image, and a real depth image, the real RGB image containing a real color of the target pixel point, the real semantic image containing a real category of the target pixel point, and the real depth image containing a real depth of the target pixel point; and the acquisition module is configured to: acquire, based on the color of the target pixel point and the real color of the target pixel point, a first loss, the first loss being used to indicate a difference between the color of the target pixel point and the real color of the target pixel point; acquire, based on the category of the target pixel point and the real category of the target pixel point, a second loss, the second loss being used to indicate a difference between the category of the target pixel point and the real category of the target pixel point; acquire, based on the depth of the target pixel point and the real depth of the target pixel point, a third loss, the third loss being used to indicate a difference between the depth of the target pixel point and the real depth of the target pixel point; and acquire, based on the first loss, the second loss, and the third loss, a target loss.
[0043] In a possible implementation, the acquisition module is further configured to acquire, based on a transparency distribution of the plurality of sampling points and a real transparency distribution of the plurality of sampling points, a fourth loss; and the acquisition module is configured to acquire, based on the first loss, the second loss, the third loss, and the fourth loss, the target loss.
[0044] In a possible implementation, the first neural network model is a neural implicit field.
[0045] In a fifth aspect, an embodiment of the present application provides a point cloud acquisition apparatus, the apparatus comprising a memory and a processor; the memory stores codes, and the processor is configured to execute the codes, when the codes are executed, the point cloud acquisition apparatus performs the method in the first aspect or any possible implementation manner of the first aspect.
[0046] In a sixth aspect, an embodiment of the present application provides a model training apparatus, the apparatus comprising a memory and a processor; the memory stores codes, and the processor is configured to execute the codes, when the codes are executed, the model training apparatus performs the method in the second aspect or any possible implementation manner of the second aspect.
[0047] In a seventh aspect, an embodiment of the present application provides circuitry, the circuitry comprising processing circuitry configured to perform the method in the first aspect, any possible implementation manner of the first aspect, the second aspect or any possible implementation manner of the second aspect.
[0048] In an eighth aspect, an embodiment of the present application provides a chip system, the chip system comprising a processor, and the processor is configured to invoke a computer program or computer instructions stored in a memory, so that the processor performs the method in the first aspect, any possible implementation manner of the first aspect, the second aspect or any possible implementation manner of the second aspect.
[0049] In a possible implementation manner, the processor is coupled with the memory through an interface.
[0050] In a possible implementation manner, the chip system further comprises the memory, and the memory stores the computer program or computer instructions.
[0051] In a ninth aspect, an embodiment of the present application provides a computer storage medium, the computer storage medium stores a computer program, and the program, when executed by a computer, causes the computer to implement the method in the first aspect, any possible implementation manner of the first aspect, the second aspect or any possible implementation manner of the second aspect.
[0052] In a tenth aspect, an embodiment of the present application provides a computer program product, the computer program product stores instructions, and the instructions, when executed by a computer, causes the computer to implement the method in the first aspect, any possible implementation manner of the first aspect, the second aspect or any possible implementation manner of the second aspect.
[0053] In the embodiments of the present application, when a four-dimensional point cloud needs to be acquired, a first pose of the camera and a second pose of the camera can be acquired first, and the first pose and the second pose are associated in time. Then, the first pose can be input into the first neural network model to process the first pose through the first neural network model, to obtain a first image captured by the camera, and to convert the first image to obtain a first three-dimensional point cloud. Then, the second pose can also be input into the first neural network model to process the second pose through the first neural network model, to obtain a second image captured by the camera, and to convert the second image to obtain a second three-dimensional point cloud. Finally, the first three-dimensional point cloud and the second three-dimensional point cloud can be fused to obtain a first four-dimensional point cloud. In the foregoing process, the neural network model (i.e., the first neural network model described above, for example, a neural implicit field) can be used to process different poses (i.e., the first pose and the second pose described above) of the camera at different times to obtain three-dimensional point clouds corresponding to different poses (which can also be understood as three-dimensional point clouds at different times, i.e., the first three-dimensional point cloud and the second three-dimensional point cloud described above), and then the three-dimensional point clouds corresponding to different poses are used to acquire a four-dimensional point cloud (i.e., the first four-dimensional point cloud described above). This way of acquiring a four-dimensional point cloud does not need to rely on a laser radar, but only needs a camera, which can reduce the hardware cost required for acquiring a point cloud, and is conducive to the implementation of this way of acquiring a four-dimensional point cloud in some cost-limited scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 A structural schematic diagram of an artificial intelligence subject framework;
[0055] Figure 2a A structural schematic diagram of a fault prediction system provided by an embodiment of the present application;
[0056] Figure 2b Another structural schematic diagram of a fault prediction system provided by an embodiment of the present application;
[0057] Figure 2c A schematic diagram of a device related to fault prediction provided by an embodiment of the present application;
[0058] Figure 3 A schematic diagram of a system 100 architecture provided by an embodiment of the present application;
[0059] Figure 4 A flowchart of a point cloud acquisition method provided by an embodiment of the present application;
[0060] Figure 5 A schematic diagram of image acquisition provided by an embodiment of the present application;
[0061] Figure 6 A schematic diagram of point cloud acquisition provided by an embodiment of the present application;
[0062] Figure 7 A schematic diagram of point cloud optimization provided by an embodiment of the present application;
[0063] Figure 8 A schematic diagram of images in the field of autonomous driving provided by an embodiment of the present application;
[0064] Figure 9 A schematic diagram of three-dimensional point cloud provided by an embodiment of the present application;
[0065] Figure 10 A schematic diagram of four-dimensional point cloud provided by an embodiment of the present application;
[0066] Figure 11 A schematic diagram of a model training method provided by an embodiment of the present application;
[0067] Figure 12 A schematic diagram of a point cloud acquisition device provided by an embodiment of the present application;
[0068] Figure 13 A schematic diagram of a model training device provided by an embodiment of the present application;
[0069] Figure 14 A schematic diagram of an execution device provided by an embodiment of the present application;
[0070] Figure 15 A schematic diagram of a training device provided by an embodiment of the present application;
[0071] Figure 16 A schematic diagram of a chip provided by an embodiment of the present application. DETAILED DESCRIPTION
[0072] Embodiments of the present application provide a point cloud acquisition method and related equipment, which can reduce the hardware cost required for point cloud acquisition, and is conducive to implementation in some cost-limited scenarios.
[0073] The terms "first", "second", etc. in the specification and claims of the present application and in the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way adopted in the description of the embodiments of the present application for the objects with the same attribute in the description. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or equipment containing a series of units do not have to be limited to those units, but can include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0074] In the field of autonomous driving, vehicles usually need to perceive and understand the environment around the vehicle by using a GOD model to improve the safety of vehicle autonomous driving. Four-dimensional point cloud as the training data of the GOD model plays a crucial role in the performance of the GOD model.
[0075] In the related art, in order to collect four-dimensional point cloud, three-dimensional point cloud is collected by using devices such as lidar, and then the four-dimensional point cloud that can be used to train the GOD model is obtained by using the three-dimensional point cloud. Specifically, the surrounding environment can be data collected at different times by using a lidar, thereby obtaining three-dimensional point cloud at different times. Then, the three-dimensional point cloud at different times can be fused to obtain four-dimensional point cloud. In this way, the obtained four-dimensional point cloud can be used to train the GOD model.
[0076] Since the above-mentioned method of obtaining four-dimensional point cloud depends on lidar and also needs the assistance of devices such as IMU, the device cost required by this method is too high, which is not conducive to the implementation in some cost-limited model training scenarios.
[0077] Further, the lidar will fail when dealing with extreme weather or low reflection intensity objects, it cannot successfully collect three-dimensional point cloud, and thus cannot successfully obtain four-dimensional point cloud. It can be seen that this method is also difficult to implement in some scenarios with special conditions.
[0078] In order to solve the above problems, the embodiment of the present application provides a scene perception method, which can be implemented in combination with artificial intelligence (artificial intelligence, AI) technology. AI technology is a technology subject that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence. AI technology obtains the best results by perceiving the environment, acquiring knowledge and using knowledge. In other words, artificial intelligence technology is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Data processing using artificial intelligence is a common application of artificial intelligence.
[0079] First, the overall workflow of the artificial intelligence system is described, please refer to Figure 1 , Figure 1As a structural diagram of the artificial intelligence subject framework, the following elaborates the above-mentioned artificial intelligence subject framework from two dimensions of "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a condensation process of "data-information-knowledge-wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of human intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0080] (1) Infrastructure
[0081] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and realizes support through the underlying platform. Communication with the outside world through sensors; computing power is provided by intelligent chips (CPU, NPU, GPU, ASIC, FPGA, etc. Hardware acceleration chips); the underlying platform includes distributed computing framework and network-related platform support and support, which can include cloud storage and computing, interconnection network, etc. For example, sensors and external communication obtain data, which are provided to intelligent chips in the distributed computing system provided by the underlying platform for calculation.
[0082] (2) Data
[0083] The data on the upper layer of the infrastructure is used to represent the data source in the field of artificial intelligence. Data involves graphics, images, speech, text, and also involves Internet of Things data of traditional devices, including business data of existing systems and sensing data such as force, displacement, liquid level, temperature, humidity, etc.
[0084] (3) Data processing
[0085] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0086] Among them, machine learning and deep learning can symbolize and formalize intelligent information modeling, extraction, preprocessing, training, etc.
[0087] Reasoning refers to the process of simulating human intelligent reasoning methods in computers or intelligent systems, using formalized information to perform machine thinking and solve problems according to reasoning control strategies, and the typical function is search and matching.
[0088] Decision-making refers to the process of decision-making after intelligent information reasoning, which usually provides functions such as classification, sorting, prediction, etc.
[0089] (4) General capabilities
[0090] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0091] (5) Smart products and industry applications
[0092] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.
[0093] Next, several application scenarios of this application are introduced.
[0094] Figure 2a This is a schematic diagram of the structure of the point cloud acquisition system provided in an embodiment of the present application. The point cloud acquisition system includes a user device and a data processing device. The user device includes intelligent terminals such as a user's mobile phone and an onboard computer in a vehicle driven by the user. The user device is the initiator of point cloud acquisition, serving as the initiator of a point cloud acquisition request, typically initiated by the user through the user device.
[0095] The aforementioned data processing device can be a device or server with data processing capabilities, such as a cloud server, network server, application server, or management server. The data processing device receives point cloud acquisition requests from intelligent terminals via an interactive interface. It then uses its memory and data processing processor to perform point cloud acquisition through machine learning, deep learning, search, reasoning, and decision-making. The memory in the data processing device is a general term that includes local storage and a database that stores historical data. The database can be located on the data processing device or on other network servers.
[0096] exist Figure 2a In the point cloud acquisition system shown, the user device can obtain the position of the user device's camera at different times, and then initiate a request to the data processing device, so that the data processing device performs point cloud acquisition processing on the position of the user device's camera at different times, thereby obtaining a four-dimensional point cloud. For example, when the user triggers the user device, the user device can obtain the position of its own camera at different times, and then the user device can initiate a point cloud acquisition request to the data processing device, so that the data processing device performs a series of processing on the position of the user device's camera at different times based on the point cloud acquisition request, thereby obtaining a four-dimensional point cloud, that is, 4D semantic occupancy data, which can be used for subsequent neural network model (for example, GOD model, etc.) training.
[0097] In Figure 2a , the data processing device can perform the point cloud acquisition method of the embodiments of the present application.
[0098] Figure 2b Another structural schematic diagram of the point cloud acquisition system provided by the embodiments of the present application is shown in Figure 2b , the user device directly serves as the data processing device, which can directly acquire the input from the user and directly process by the hardware of the user device itself. The specific process is similar to Figure 2a , and reference can be made to the above description, which will not be repeated here.
[0099] In Figure 2b , when the user triggers the user device, the user device can acquire the poses of its camera at different times, and then the user device can perform a series of processing on the poses of its camera at different times, thereby obtaining four-dimensional point cloud, that is, 4D semantic occupancy data, which can be used for subsequent neural network model (for example, GOD model, etc.) training.
[0100] In Figure 2b , the user device itself can perform the point cloud acquisition method of the embodiments of the present application.
[0101] Figure 2c A schematic diagram of a related device for point cloud acquisition provided by the embodiments of the present application.
[0102] The user device in the above Figure 2a and Figure 2b may be specifically the local device 301 or the local device 302 in Figure 2c , the data processing device in Figure 2a may be specifically the execution device 210 in Figure 2c , wherein the data storage system 250 can store the to-be-processed data of the execution device 210, and the data storage system 250 can be integrated on the execution device 210, or can be set on the cloud or other network servers.
[0103] Figure 2a The processor in Figure 2b may perform data training / machine learning / deep learning through a neural network model or other model (for example, a model based on a support vector machine), and finally train or learn a model through the data, and then perform point cloud acquisition application on the image through the model, thereby obtaining a corresponding processing result.
[0104] Figure 3 A schematic diagram of the system 100 architecture provided by the embodiments of the present application is shown in Figure 3In the embodiment of the present application, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with an external device. The user can input data to the I / O interface 112 through the client device 140. The input data may include: various tasks to be scheduled, callable resources and other parameters.
[0105] When the execution device 110 preprocesses the input data, or when the computing module 111 of the execution device 110 performs calculations and other related processing (such as implementing the functions of the neural network in this application), the execution device 110 can call the data, code, etc. in the data storage system 150 for the corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.
[0106] Finally, the I / O interface 112 returns the processing result to the client device 140 so as to provide it to the user.
[0107] It is worth noting that the training device 120 can generate corresponding target models / rules based on different training data for different goals or tasks. The corresponding target models / rules can be used to achieve the above goals or complete the above tasks, thereby providing the user with the desired results. The training data can be stored in the database 130 and come from training samples collected by the data collection device 160.
[0108] exist Figure 3 In the case shown in FIG, the user can manually input data, which can be operated through the interface provided by I / O interface 112. In another case, client device 140 can automatically send input data to I / O interface 112. If the automatic transmission of input data by client device 140 requires user authorization, the user can set the corresponding permissions in client device 140. The user can view the results output by execution device 110 on client device 140, which can be presented in the form of display, sound, action, etc. Client device 140 can also serve as a data acquisition terminal, collecting input data input into I / O interface 112 and output results from I / O interface 112 as new sample data and storing them in database 130. Of course, the collection can also be performed without client device 140, and instead the input data input into I / O interface 112 and output results from I / O interface 112 as new sample data can be directly stored in database 130 by I / O interface 112.
[0109] It is worth noting that Figure 3 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example,Figure 3 In some implementations, the data storage system 150 is external to the execution device 110. In other implementations, the data storage system 150 can be incorporated into the execution device 110. Figure 3 As shown, the neural network can be trained by the training device 120.
[0110] The embodiments of the present application also provide a chip including a neural network processor NPU. The chip can be arranged in the execution device 110 as shown to complete the computing work of the computing module 111. The chip can also be arranged in the training device 120 as shown to complete the training work of the training device 120 and output the target model / rule. Figure 3 Figure 3
[0111] The neural network processor NPU is mounted on a host central processing unit (CPU) as a coprocessor, and tasks are allocated by the host CPU. The core part of the NPU is an operation circuit, and the controller controls the operation circuit to extract data in the memory (weight memory or input memory) and perform operations.
[0112] In some implementations, the operation circuit includes a plurality of processing units (PEs) inside. In some implementations, the operation circuit is a two-dimensional systolic array. The operation circuit can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit is a general matrix processor.
[0113] For example, assuming there are input matrix A, weight matrix B, and output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory and performs matrix operations with matrix B to obtain partial results or final results of the matrix, which are saved in the accumulator.
[0114] The vector calculation unit can further process the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit can be used for network calculation of non-convolution / non-FC layers in the neural network, such as pooling, batch normalization, local response normalization, etc.
[0115] In some implementations, the vector computation unit can store the processed output vector to the unified buffer. For example, the vector computation unit can apply a non-linear function to the output of the arithmetic circuit, e.g., a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit generates normalized values, merged values, or both. In some implementations, the processed output vector can be used as an activation input to the arithmetic circuit, e.g., for use in a subsequent layer in a neural network.
[0116] The unified memory is used to store input data and output data.
[0117] The weight data is transferred from the external memory to the input memory and / or the unified memory, from the external memory to the weight memory, and from the unified memory to the external memory by a direct memory access controller (DMAC).
[0118] A bus interface unit (BIU) is used to interact with the main CPU, the DMAC, and the instruction fetch buffer via a bus.
[0119] An instruction fetch buffer connected to the controller is used to store instructions used by the controller.
[0120] The controller is used to call the cached instructions in the instruction fetch buffer to control the working process of the arithmetic accelerator.
[0121] Generally, the unified memory, the input memory, the weight memory, and the instruction fetch buffer are on-chip memories, and the external memory is a memory external to the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or other readable and writable memories.
[0122] Since the embodiments of the present application involve a large number of neural network applications, in order to facilitate understanding, the related terms and concepts related to neural networks involved in the embodiments of the present application will be introduced first.
[0123] (1) Neural network
[0124] The neural network can be composed of neural units, and a neural unit can be an operation unit with xs and intercept 1 as inputs. The output of the operation unit can be:
[0125]
[0126] where s = 1, 2, … n, n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. The neural network is a network formed by connecting many single neural units described above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.
[0127] The work of each layer in the neural network can be described by the mathematical expression y = a(Wx + b): from the physical layer, the work of each layer in the neural network can be understood as completing the transformation of the input space (the set of input vectors) to the output space (i.e. the row space of the matrix to the column space) by five operations on the input space, which include: 1, dimensionality increase / dimensionality decrease; 2, magnification / reduction; 3, rotation; 4, translation; 5, "bending". Among them, the operations of 1, 2, and 3 are completed by Wx, the operation of 4 is completed by +b, and the operation of 5 is completed by a(). The reason why "space" is used here is that the objects to be classified are not single things, but a class of things, and the space refers to the set of all individuals of this class of things. Among them, W is a weight vector, each value in the vector represents the weight value of a neuron in the neural network of this layer. The vector W determines the spatial transformation of the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training the neural network is also to obtain the weight matrix of all layers of the trained neural network (the weight matrix formed by the vectors W of many layers). Therefore, the training process of the neural network is essentially learning the way to control the spatial transformation, more specifically, learning the weight matrix.
[0128] Because the output of the neural network is expected to be as close as possible to the value that is actually intended to be predicted, the weight vector of each layer of the neural network can be updated by comparing the predicted value of the current network with the target value that is actually intended to be predicted, and then according to the difference between the two, for example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and the adjustment is continuously made until the neural network can predict the target value that is actually intended to be predicted. Therefore, it is necessary to define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function, which is an important equation for measuring the difference between the predicted value and the target value. Taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, and then the training of the neural network becomes a process of trying to minimize the loss.
[0129] (2) Back propagation algorithm
[0130] The neural network can use the back propagation (BP) algorithm to correct the size of the parameters in the initial neural network model in the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will produce an error loss, and the initial neural network model parameters are updated by back propagating the error loss information, so as to make the error loss converge. The back propagation algorithm is a back propagation movement dominated by error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.
[0131] (3) Bird eye view
[0132] The bird eye view (BEV) graph is a natural and direct candidate view, which can be used as a unified representation. Compared with the front view or perspective view widely studied in the field of two-dimensional vision, the BEV representation has some inherent advantages. First, the BEV does not have the problem of occlusion and scale that exists universally in two-dimensional tasks. In the field of autonomous driving, the problem of identifying vehicles with occlusion or crossing traffic can be better solved by using BEV. In addition, using BEV to represent objects or road elements will be beneficial to the development and deployment of subsequent modules (such as planning and control).
[0133] The method provided in the present application is described below from the training side of the neural network and the application side of the neural network.
[0134] The model training method provided in the embodiments of the present application relates to processing of data sequences, and can be applied to data training, machine learning, deep learning and the like. The training data (for example, the pose of the camera in the model training method provided in the embodiments of the present application) is subjected to intelligent information modeling, extraction, preprocessing, training and the like in a symbolic and formal manner, and finally a trained neural network (for example, the first neural network model in the model training method provided in the embodiments of the present application) is obtained. In addition, the point cloud acquisition method provided in the embodiments of the present application can use the trained neural network, input the input data (for example, the first pose of the camera and the second pose of the camera in the point cloud acquisition method provided in the embodiments of the present application) into the trained neural network, and obtain the output data (for example, the first four-dimensional point cloud provided in the embodiments of the present application). It should be noted that the model training method and the point cloud acquisition method provided in the embodiments of the present application are based on the same concept and can be understood as two parts of a system or two stages of an overall process, for example, a model training stage and a model application stage.
[0135] The four-dimensional point cloud data obtained by the point cloud acquisition method provided in the embodiments of the present application can be applied not only to model training required by object detection and object segmentation in the field of autonomous driving, but also to model training required by more scenarios such as robot route planning in the field of logistics transportation, which is not limited here. In order to understand the process of the embodiments of the present application, the process is introduced below, Figure 4 Figure 4 A flowchart of the point cloud acquisition method provided in the embodiments of the present application is shown in Figure 4 The method comprises the following steps.
[0136] 401. Obtain a first pose of a camera and a second pose of the camera, the first pose being associated with the second pose in time.
[0137] In the embodiments of the present application, when a four-dimensional point cloud is needed, a first pose of a camera at a first time and a second pose of the camera at a second time can be obtained. Since the first time and the second time can be regarded as two different times (for example, the first time and the second time are adjacent times), the first pose and the second pose are two poses associated in time.
[0138] 402. Process the first pose by a first neural network model to obtain a first image captured by the camera, and convert the first image to obtain a first three-dimensional point cloud.
[0139] After obtaining the first pose, the first pose can be input into a first neural network model, which performs a series of processing on the first pose to simulate a first image captured by the camera in the first pose. The first image captured by the camera can then be transformed (e.g., using a three-dimensional projection) to obtain a first three-dimensional point cloud.
[0140] Specifically, the first neural network model can be a neural implicit field (that is, a fully connected neural network model) or other neural network models, which is not limited here.
[0141] More specifically, the first image may include a first red / green / blue (RGB) image captured by a camera, a first semantic image corresponding to the first RGB image, a first depth image corresponding to the first RGB image, and a first mask corresponding to the first RGB image. The first RGB image may include multiple pixels, the first semantic image may include the category of each pixel in the first RGB image, the first depth image may include the depth of each pixel in the first RGB image, and the first mask may include the evaluation value of each pixel in the first RGB image.
[0142] More specifically, the first image can be obtained by:
[0143] Since the first RGB image contains multiple pixels, the operations performed on each pixel in this embodiment are similar. For ease of explanation, the following is a schematic introduction using any pixel contained in the first RGB image, and this pixel is referred to as the first pixel. Then, the process of obtaining information of the first pixel includes:
[0144] (1) After obtaining the first pose of the camera, the first pose can be input into the first neural network model. Then, the first neural network model can construct a first ray sent from the camera in the first pose to the first pixel point, wherein the first ray can start from the camera in the first pose and end at the first pixel point, or the first ray can also start from the camera in the first pose and pass through the first pixel point, without limitation.
[0145] For example, Figure 5 As shown ( Figure 5 (a schematic diagram of image acquisition provided in an embodiment of the present application) After determining the position of the camera at a certain moment, the position can be input into the neural implicit field. The neural implicit field can determine a pixel point (u, v) in the RGB image (simulated) captured by the camera in the position, and construct the ray emitted by the camera in the position to (u, v). Where u is the horizontal coordinate of (u, v) in the RGB image, and v is the vertical coordinate of (u, v) in the RGB image.
[0146] (2) After obtaining the first ray, the first neural network model can sample the first ray, so as to obtain a plurality of first sampling points on the first ray.
[0147] Still as in the above example, after constructing the ray emitted by the camera in the pose to (u, v), the neural implicit field can sample the ray to obtain N sampling points
[0148] (3) After obtaining the plurality of first sampling points, the first neural network model can perform a series of processing on the plurality of first sampling points, so as to obtain the color of each first sampling point in the plurality of first sampling points, the category of each first sampling point in the plurality of first sampling points, the depth of each first sampling point in the plurality of first sampling points, and the transparency of each first sampling point in the plurality of first sampling points (which can also be referred to as the probability density of each first sampling point in the plurality of first sampling points).
[0149] Still as in the above example, among the N sampling points, for the i-th sampling point p i , the neural implicit field can process p i , so as to output the color c(p i ) of p i , the category s(p i ) of p i , the depth d(p i ) of p i , and the probability density σ(p i ) of p i . For the remaining sampling points in the N sampling points, the neural implicit field can also output the color of the remaining sampling points, the category of the remaining sampling points, the depth of the remaining sampling points, and the probability density of the remaining sampling points, which will not be described here. In this way, the color of each sampling point, the category of each sampling point, the depth of each sampling point, and the probability density of each sampling point in the N sampling points can be obtained.
[0150] (4) After obtaining the color of each first sampling point in the plurality of first sampling points, the category of each first sampling point in the plurality of first sampling points, the depth of each first sampling point in the plurality of first sampling points, and the transparency of each first sampling point in the plurality of first sampling points, the color of the first pixel point, the category of the first pixel point, and the depth of the first pixel point can be obtained based on the color of each first sampling point in the plurality of first sampling points, the category of each first sampling point in the plurality of first sampling points, the depth of each first sampling point in the plurality of first sampling points, and the transparency of each first sampling point in the plurality of first sampling points.
[0151] Then, for the remaining pixels except the first pixel in the first RGB image, the same operation as that performed on the first pixel can be performed on the remaining pixels, so the color of each pixel in the first RGB image, the category of each pixel in the first RGB image, and the depth of each pixel in the first RGB image can be obtained. This is equivalent to successfully obtaining the first RGB image, the first semantic image, and the first depth image.
[0152] Still as in the above example, the colors of these N sampling points, the categories of these N sampling points, the depths of these N sampling points, and the probability density of these N sampling points are obtained, which can be calculated to obtain the color r(u,v) of (u,v), the category s(u,v) of (u,v), and the depth d(u,v) of (u,v).
[0153] Then, for the remaining pixels except (u, v) in the RGB image taken by the camera in this pose, the same operation as that performed on (u, v) can be performed, so the color, category and depth of each pixel in the RGB image taken by the camera in this pose can be obtained, which is equivalent to obtaining the RGB image taken by the camera in this pose, the semantic map corresponding to the RGB image, and the depth map corresponding to the RGB image.
[0154] More specifically, the color of the first pixel, the category of the first pixel, the depth of the first pixel, and the transparency of the first pixel may be obtained in the following manner:
[0155] (4.1) After obtaining the color of each first sampling point in the plurality of first sampling points, the category of each first sampling point in the plurality of first sampling points, the depth of each first sampling point in the plurality of first sampling points, and the transparency of each first sampling point in the plurality of first sampling points, the color of each first sampling point in the plurality of first sampling points and the transparency of each first sampling point in the plurality of first sampling points may be calculated to obtain the color of the first pixel.
[0156] Still as in the above example, the color, category, depth and probability density of each sampling point in these N sampling points can be obtained. The following formula can be used to calculate r(u,v) of (u,v):
[0157]
[0158]
[0159] In the above formula, c(p i ) is the i-th sampling point p i Color, s(p i ) is p i Category, d(p i) is p i The depth, σ(p i ) is p i The probability density of i =d(p i+1 )-d(p i ), is the distance between the i+1th sampling point and the i-th sampling point.
[0160] (4.2) After obtaining the color of each of the plurality of first sampling points, the category of each of the plurality of first sampling points, the depth of each of the plurality of first sampling points, and the transparency of each of the plurality of first sampling points, the category of each of the plurality of first sampling points and the transparency of each of the plurality of first sampling points may be further calculated to obtain the category of the first pixel point.
[0161] Still as in the above example, the color, category, depth and probability density of each sampling point in these N sampling points can be obtained. The following formula can be used to calculate s(u,v) of (u,v):
[0162]
[0163] (4.3) After obtaining the color of each of the plurality of first sampling points, the category of each of the plurality of first sampling points, the depth of each of the plurality of first sampling points, and the transparency of each of the plurality of first sampling points, the depth of each of the plurality of first sampling points and the transparency of each of the plurality of first sampling points may be further calculated to obtain the depth of the first pixel.
[0164] Still as in the above example, the color, category, depth and probability density of each sampling point in these N sampling points can be obtained. The following formula can be used to calculate d(u,v) of (u,v):
[0165]
[0166] More specifically, the evaluation value of the first pixel can also be obtained in the following way:
[0167] After obtaining the color of each first sampling point in the plurality of first sampling points, the category of each first sampling point in the plurality of first sampling points, the depth of each first sampling point in the plurality of first sampling points, the transparency of each first sampling point in the plurality of first sampling points, and the depth of the first pixel point, the depth of each first sampling point in the plurality of first sampling points and the depth of the first pixel point can be calculated, so as to obtain an evaluation value (which can also be referred to as uncertainty) of the first pixel point, the evaluation value of the first pixel point being used to indicate the quality of the first pixel point, the higher the evaluation value, the lower the quality of the first pixel point, and the lower the evaluation value, the higher the quality of the first pixel point.
[0168] Then, for the remaining pixel points in the first RGB image except the first pixel point, the operations performed on the first pixel point can also be performed on the remaining pixel points, so that the evaluation values of the pixel points in the first RGB image can be finally obtained, that is, the first mask can be obtained.
[0169] Still as in the above example, after obtaining the color of each sampling point in the N sampling points, the category of each sampling point, the depth of each sampling point, and the probability density of each sampling point, the evaluation value u(u, v) of (u, v) can be calculated by the following formula:
[0170]
[0171] In the above formula, θ is a preset parameter. Then, for the remaining pixel points in the RGB image captured by the camera in the pose except (u, v), the operations performed on (u, v) can also be performed on the remaining pixel points, so that the evaluation values of the pixel points in the RGB image captured by the camera in the pose can be obtained, that is, the mask corresponding to the RGB image captured by the camera in the pose can be obtained.
[0172] 403. The second pose is processed by the first neural network model to obtain a second image captured by the camera, and the second image is converted to obtain a second three-dimensional point cloud.
[0173] After obtaining the second pose, the second pose can be input into the first neural network model to perform a series of processing on the second pose by the first neural network model, so as to simulate a second image captured by a camera in the second pose. Then, the second image captured by the camera can be converted to obtain a second three-dimensional point cloud.
[0174] Specifically, the second image can include a second RGB image captured by the camera, a second semantic image corresponding to the second RGB image, a second depth image corresponding to the second RGB image, and a second mask corresponding to the second RGB image. The second RGB image can include a plurality of pixels, the second semantic image can include a category of each pixel in the second RGB image, the second depth image can include a depth of each pixel in the second RGB image, and the second mask can include an evaluation value of each pixel in the second RGB image.
[0175] More specifically, the second image can be obtained by the following method:
[0176] Since the second RGB image includes a plurality of pixels, the operations performed by the embodiment on each pixel are similar. For the convenience of description, the following will be introduced illustratively with respect to any one pixel included in the second RGB image, which will be referred to as a second pixel. Then, the information obtaining process of the second pixel includes:
[0177] (1) After obtaining the second pose of the camera, the second pose can be input into the first neural network model. Then, the first neural network model can construct a second ray sent from the camera in the second pose to the second pixel, wherein the second ray can have the camera in the second pose as a starting point and the second pixel as a terminal point, or the second ray can have the camera in the second pose as a starting point and pass through the second pixel, which is not limited here.
[0178] (2) After obtaining the second ray, the second neural network model can sample the second ray to obtain a plurality of second sampling points on the second ray.
[0179] (3) After obtaining the plurality of second sampling points, the second neural network model can perform a series of processing on the plurality of second sampling points to obtain the color of each second sampling point in the plurality of second sampling points, the category of each second sampling point in the plurality of second sampling points, the depth of each second sampling point in the plurality of second sampling points, and the transparency of each second sampling point in the plurality of second sampling points (which can also be referred to as the probability density of each second sampling point in the plurality of second sampling points).
[0180] (4) After obtaining the color of each second sampling point in the plurality of second sampling points, the category of each second sampling point in the plurality of second sampling points, the depth of each second sampling point in the plurality of second sampling points, and the transparency of each second sampling point in the plurality of second sampling points, the color of the second pixel, the category of the second pixel, and the depth of the second pixel can be obtained based on the color of each second sampling point in the plurality of second sampling points, the category of each second sampling point in the plurality of second sampling points, the depth of each second sampling point in the plurality of second sampling points, and the transparency of each second sampling point in the plurality of second sampling points.
[0181] Then, for the rest of the pixel points in the second RGB image except the second pixel point, the operation performed on the second pixel point can also be performed on the rest of the pixel points, and thus the color of each pixel point in the second RGB image, the category of each pixel point in the second RGB image, and the depth of each pixel point in the second RGB image can be obtained, which is equivalent to successfully obtaining the second RGB image, the second semantic image, and the second depth image.
[0182] More specifically, the color of the second pixel point, the category of the second pixel point, the depth of the second pixel point, and the transparency of the second pixel point can also be obtained by the following way:
[0183] (4.1) After obtaining the color of each second sampling point in the plurality of second sampling points, the category of each second sampling point in the plurality of second sampling points, the depth of each second sampling point in the plurality of second sampling points, and the transparency of each second sampling point in the plurality of second sampling points, the color of each second sampling point in the plurality of second sampling points and the transparency of each second sampling point in the plurality of second sampling points can be calculated, and thus the color of the second pixel point can be obtained.
[0184] (4.2) After obtaining the color of each second sampling point in the plurality of second sampling points, the category of each second sampling point in the plurality of second sampling points, the depth of each second sampling point in the plurality of second sampling points, and the transparency of each second sampling point in the plurality of second sampling points, the category of each second sampling point in the plurality of second sampling points and the transparency of each second sampling point in the plurality of second sampling points can also be calculated, and thus the category of the second pixel point can be obtained.
[0185] (4.3) After obtaining the color of each second sampling point in the plurality of second sampling points, the category of each second sampling point in the plurality of second sampling points, the depth of each second sampling point in the plurality of second sampling points, and the transparency of each second sampling point in the plurality of second sampling points, the depth of each second sampling point in the plurality of second sampling points and the transparency of each second sampling point in the plurality of second sampling points can also be calculated, and thus the depth of the second pixel point can be obtained.
[0186] More specifically, the evaluation value of the second pixel point can also be obtained by the following way:
[0187] After obtaining the color of each of the plurality of second sampling points, the category of each of the plurality of second sampling points, the depth of each of the plurality of second sampling points, the transparency of each of the plurality of second sampling points, and the depth of the second pixel point, the depth of each of the plurality of second sampling points and the depth of the second pixel point can be calculated, so as to obtain an evaluation value (which can also be referred to as uncertainty) of the second pixel point, the evaluation value of the second pixel point being used to indicate the quality of the second pixel point, the higher the evaluation value, the lower the quality of the second pixel point, and the lower the evaluation value, the higher the quality of the second pixel point.
[0188] Then, for the remaining pixel points of the second RGB except the second pixel point, the operation performed on the second pixel point can also be performed on the remaining pixel points, so that the evaluation value of each pixel point in the second RGB can be finally obtained, that is, the second mask is obtained.
[0189] 404, fuse the first three-dimensional point cloud and the second three-dimensional point cloud to obtain a first four-dimensional point cloud.
[0190] After obtaining the first three-dimensional point cloud and the second three-dimensional point cloud, the first three-dimensional point cloud and the second three-dimensional point cloud can be fused (for example, feature matching, splicing, coordinate system conversion, and voxelization, etc.), so as to obtain the first four-dimensional point cloud. Then, the first four-dimensional point cloud can be used as subsequent data for model training.
[0191] Specifically, the first four-dimensional point cloud can be obtained in the following manner:
[0192] After obtaining the first three-dimensional point cloud, since the plurality of space points included in the first three-dimensional point cloud and the plurality of pixel points included in the first RGB are one-to-one corresponding (that is, the first three-dimensional point cloud is obtained by performing three-dimensional projection on each pixel point in the first RGB by using the parameters of the camera, the first semantic map, and the first depth map), the evaluation value of each pixel point in the first RGB can be used as the evaluation value of each space point in the first three-dimensional point cloud. From the first three-dimensional point cloud, the space points with evaluation values less than an evaluation threshold (the size of the threshold can be set according to actual requirements, which is not limited here) are removed, so as to obtain the first three-dimensional point cloud after denoising.
[0193] Similarly, after obtaining the second three-dimensional point cloud, since the plurality of spatial points contained in the second three-dimensional point cloud are in one-to-one correspondence with the plurality of pixel points contained in the second RGB image (that is, the second three-dimensional point cloud is obtained by performing three-dimensional projection on each pixel point in the second RGB image by using the parameters of the camera, the second semantic image, and the second depth image), the evaluation value of each pixel point in the second RGB image can be taken as the evaluation value of each spatial point in the second three-dimensional point cloud, and the spatial points with evaluation values less than the evaluation threshold value are removed from the second three-dimensional point cloud, so as to obtain the second three-dimensional point cloud after denoising.
[0194] After obtaining the first three-dimensional point cloud after denoising and the second three-dimensional point cloud after denoising, the first three-dimensional point cloud after denoising and the second three-dimensional point cloud after denoising can be fused (for example, feature matching, splicing, and voxelization, etc.), so as to obtain the first four-dimensional point cloud.
[0195] Still as the above example, as shown in Figure 6 Figure 6 A schematic diagram of point cloud acquisition provided by an embodiment of the present application, Figure 6 is obtained on the basis of Figure 5 After obtaining the RGB image, the semantic image, the depth image, and the mask film captured by the camera in the pose, projection can be completed by using the RGB image, the semantic image, and the depth image captured by the camera in the pose, so as to obtain the three-dimensional point cloud corresponding to the pose, and then the three-dimensional point cloud corresponding to the pose is denoised by using the mask film, so as to obtain the three-dimensional point cloud after denoising corresponding to the pose.
[0196] Similarly, based on the neural implicit field, the RGB image, the semantic image, the depth image, and the mask film captured by the camera in the remaining poses can also be obtained, and projection can be completed by using these images, so as to obtain the three-dimensional point cloud corresponding to the remaining poses, and then the three-dimensional point cloud corresponding to the remaining poses is denoised by using the mask film, so as to obtain the three-dimensional point cloud after denoising corresponding to the remaining poses.
[0197] Then, the three-dimensional point cloud after denoising corresponding to the pose and the three-dimensional point cloud after denoising corresponding to the remaining poses can be subjected to feature matching and splicing, so as to obtain the three-dimensional point cloud in the camera coordinate system, and then the three-dimensional point cloud in the camera coordinate system is converted into the three-dimensional point cloud in the world coordinate system, and then the three-dimensional point cloud in the world coordinate system is voxelized, so as to obtain the four-dimensional point cloud obtained by voxelization.
[0198] More specifically, the first three-dimensional point cloud can also be optimized in the following manner:
[0199] After obtaining the first image and the second image, the first image and the second image can also be input into a second neural network model to perform a series of processing on the first image and the second image through the second neural network model (for example, a GOD model, etc.) to obtain a second four-dimensional point cloud.
[0200] After obtaining the second four-dimensional point cloud, the second four-dimensional point cloud can be used to optimize the first four-dimensional point cloud (for example, by fusing the first four-dimensional point cloud and the second four-dimensional point cloud), thereby obtaining an optimized first four-dimensional point cloud.
[0201] Still like the above example, Figure 7 As shown ( Figure 7 A schematic diagram of point cloud optimization provided in an embodiment of the present application. Figure 7 is Figure 6 After obtaining the denoised 3D point cloud corresponding to that pose and the denoised 3D point clouds corresponding to the remaining poses, these point clouds can be feature matched and stitched together to obtain a 3D point cloud in the camera coordinate system. The 3D point cloud in the camera coordinate system can then be converted to a 3D point cloud in the world coordinate system. The 3D point cloud in the world coordinate system can then be voxelized to obtain a voxelized 4D point cloud.
[0202] In addition, the RGB image, semantic image, and depth image captured by the camera in this pose, as well as the RGB image, semantic image, and depth image captured by the camera in other poses, can be input into the GOD model to obtain the 4D point cloud predicted by the model. Then, the 4D point cloud obtained by voxelization and the 4D point cloud predicted by the model can be fused to obtain the optimized 4D point cloud.
[0203] It should be understood that in this embodiment, the denoising operation is only schematically introduced as being performed before the point cloud fusion operation. The denoising operation can also be performed during the point cloud fusion operation. For example, after obtaining the first three-dimensional point cloud and the second three-dimensional point cloud, the two three-dimensional point clouds can be first feature matched and spliced, and then the three-dimensional point cloud located in the camera coordinate system can be denoised to obtain the denoised three-dimensional point cloud located in the camera coordinate system, and then the coordinate system is transformed and voxelized to obtain the first four-dimensional point cloud.
[0204] In addition, the first neural network model provided in the embodiment of the present application can realize the volume rendering function, which can simulate the RGB images, semantic images and depth images taken by cameras in different postures, that is, the RGB images, semantic images and depth images under different viewing angles. The subsequent model trained based on the four-dimensional point cloud obtained from these images can more accurately understand the content captured by the camera, thereby having better performance. For example, Figure 8 As shown ( Figure 8A schematic diagram of an image in the field of autonomous driving provided in an embodiment of the present application) is provided, using a first neural network model (implicit neural field) to obtain RGB images, semantic maps and depth maps from different perspectives, wherein the semantic map (semantic segmentation) and depth map (depth estimation) are used to obtain a four-dimensional point cloud to train the model in the autonomous driving system, which can help the model of the autonomous driving system to understand the surrounding environment more accurately, thereby better planning the driving route and avoiding potential dangers. Specifically, semantic segmentation can assign each pixel in the RGB image to a different semantic category, such as roads, vehicles, pedestrians, etc., thereby helping the model of the autonomous driving system to better understand the road and the surrounding environment. Depth estimation can estimate the distance from each pixel to the camera, thereby helping the autonomous driving system to better understand the three-dimensional structure and distance relationship of the scene. This information can be used to generate high-precision maps, plan safer and more efficient driving routes, and when encountering obstacles or other dangerous situations, the model of the autonomous driving system can react more quickly and take appropriate measures to ensure the driving safety of the vehicle.
[0205] Furthermore, the embodiments of the present application can be compared with related technologies. The three-dimensional point cloud generated by the related technologies has certain noise points, while the embodiments of the present application can use the uncertainty obtained by volume rendering (that is, the aforementioned evaluation value) to denoise the obtained three-dimensional point cloud, thereby obtaining a denoised three-dimensional point cloud, which is conducive to constructing a four-dimensional point cloud with better quality and training a model with better performance. For example, Figure 9 As shown ( Figure 9 A schematic diagram of a three-dimensional point cloud provided in an embodiment of the present application) is provided. In the three-dimensional point cloud obtained by the related art (see Figure 9 The left half of the image has many noise points. The embodiment of the present application can effectively remove these noise points, thereby obtaining a three-dimensional point cloud with better quality (see Figure 9 the right half of the ).
[0206] Furthermore, the four-dimensional point cloud obtained by the embodiment of the present application can be compared with the four-dimensional point cloud collected by related technologies (for example, using a laser radar or other equipment), such as Figure 10 As shown ( Figure 10 A schematic diagram of the four-dimensional point cloud provided in an embodiment of the present application) shows that the three-dimensional scene accuracy reflected by the four-dimensional point clouds obtained by the two is comparable, and their three-dimensional overlap accuracy (Precision) is 74%, recall rate (Recall) is 84%, and intersection of union (IoU) is 64%.
[0207] In the embodiments of the present application, when a four-dimensional point cloud needs to be obtained, a first pose of a camera and a second pose of the camera can be obtained first, the first pose and the second pose being associated in time. Then, the first pose can be input into a first neural network model to process the first pose by the first neural network model, to obtain a first image captured by the camera, and to convert the first image to obtain a first three-dimensional point cloud. Then, the second pose can also be input into the first neural network model to process the second pose by the first neural network model, to obtain a second image captured by the camera, and to convert the second image to obtain a second three-dimensional point cloud. Finally, the first three-dimensional point cloud and the second three-dimensional point cloud can be fused to obtain a first four-dimensional point cloud. In the foregoing process, the neural network model (i.e., the first neural network model described above, for example, a neural implicit field) can be used to process different poses of the camera at different times (i.e., the first pose and the second pose described above) to obtain three-dimensional point clouds corresponding to different poses (which can also be understood as three-dimensional point clouds at different times, i.e., the first three-dimensional point cloud and the second three-dimensional point cloud described above), and then the three-dimensional point clouds corresponding to different poses are used to obtain a four-dimensional point cloud (i.e., the first four-dimensional point cloud described above). This way of obtaining a four-dimensional point cloud does not need to rely on a lidar, but only needs a camera, which can reduce the hardware cost required for obtaining a point cloud, and is conducive to the implementation of this way of obtaining a four-dimensional point cloud in some cost-limited scenarios.
[0208] Further, in the embodiments of the present application, based on the processing of the pose of the camera at a certain time by the neural network model, the RGB image captured by the camera at the pose, the corresponding semantic image and the depth image can be simulated. Since these images have better dense texture features, even if the content in the image is an extreme weather or a low-reflectivity object, various features of these situations can be reflected, so that a high-quality three-dimensional point cloud can be successfully obtained based on these images, and then a four-dimensional point cloud can be obtained. It can be seen that this way of obtaining a four-dimensional point cloud can be implemented in some scenarios with special conditions.
[0209] Further, in the embodiments of the present application, based on the processing of the pose of the camera at a certain time by the neural network model, the RGB image captured by the camera at the pose, the corresponding semantic image and the depth image can be simulated. Since these images have better dense texture features, even if the content in the image is an extreme weather or a low-reflectivity object, various features of these situations can be reflected, so that a high-quality three-dimensional point cloud can be successfully obtained based on these images, and then a four-dimensional point cloud can be obtained. It can be seen that this way of obtaining a four-dimensional point cloud can be implemented in some scenarios with special conditions.
[0210] Further, in the embodiments of the present application, after obtaining the four-dimensional point cloud, the RGB image, the corresponding semantic image and the depth image can be processed by another neural network model (i.e., the aforementioned second neural network model, such as the GOD model, etc.) to obtain a four-dimensional point cloud predicted by the model, so as to optimize the four-dimensional point cloud obtained based on the three-dimensional point cloud, thereby obtaining a four-dimensional point cloud with better quality.
[0211] The above is a detailed description of the point cloud acquisition method provided by the embodiments of the present application. The model training method provided by the embodiments of the present application will be introduced below. Figure 11 A flowchart of the model training method provided by the embodiments of the present application is shown in FIG. 2, and the method comprises the following steps. Figure 11
[0212] 1101. The pose of the camera is processed by the to-be-trained model to obtain an image captured by the camera, the image is used to acquire a three-dimensional point cloud, and the three-dimensional point cloud is used to acquire a four-dimensional point cloud.
[0213] In the embodiments, when the to-be-trained model needs to be trained, the to-be-trained model (e.g., a to-be-trained neural implicit field) can be acquired first. Then, a batch of training data can be acquired, which contains the pose of the camera and the real image captured by the camera in the pose, and the real image captured by the camera can include a real RGB image captured by the camera, a real semantic image corresponding to the real RGB image, and a real depth image corresponding to the real RGB image.
[0214] Then, the pose of the camera can be input to the to-be-trained model to process the pose of the camera by the to-be-trained model, so as to simulate a (predicted) image captured by the camera in the pose, which can be used to acquire a three-dimensional point cloud, and the three-dimensional point cloud can be used to acquire a four-dimensional point cloud.
[0215] In a possible implementation manner, the image captured by the camera includes a (predicted) RGB image captured by the camera, a (predicted) semantic image corresponding to the RGB image, and a (predicted) depth image corresponding to the RGB image, wherein the RGB image includes each pixel point, that is, the (predicted) color of each pixel point, the semantic image includes the (predicted) category of each pixel point in the RGB image, and the depth image includes the (predicted) depth of each pixel point in the RGB image.
[0216] In a possible implementation manner, the RGB image contains a target pixel point, the target pixel point being any one of the pixel points in the RGB image, and the camera pose is processed by the first neural network model to obtain the image captured by the camera, including: obtaining, by the to-be-trained model, a ray sent to the target pixel point by the camera in the pose; sampling, by the to-be-trained model, the ray to obtain a plurality of sampling points; processing, by the to-be-trained model, the plurality of sampling points to obtain (predicted) colors of the plurality of sampling points, (predicted) categories of the plurality of sampling points, (predicted) depths of the plurality of sampling points, and (predicted) transparencies of the plurality of sampling points; and obtaining, based on the colors of the plurality of sampling points, the categories of the plurality of sampling points, the depths of the plurality of sampling points, and the transparencies of the plurality of sampling points, a color of the target pixel point, a category of the target pixel point, and a depth of the target pixel point.
[0217] In a possible implementation manner, the obtaining, based on the colors of the plurality of sampling points, the categories of the plurality of sampling points, the depths of the plurality of sampling points, and the transparencies of the plurality of sampling points, the color of the target pixel point, the category of the target pixel point, and the depth of the target pixel point includes: calculating the colors of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the color of the target pixel point; calculating the categories of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the category of the target pixel point; and calculating the depths of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the depth of the target pixel point.
[0218] It should be noted that the description of step 1101 can refer to the related description of step 402 or step 403 in the foregoing embodiments, which will not be repeated here.
[0219] 1102. Obtain a target loss based on the image and the real image corresponding to the image, the target loss being used to indicate a difference between the image and the real image.
[0220] After obtaining the image captured by the camera, since the real image captured by the camera is known, the image captured by the camera and the real image captured by the camera can be processed to obtain a target loss, which can be used to indicate a difference between the image captured by the camera and the real image captured by the camera.
[0221] Specifically, after obtaining the RGB image captured by the camera, the semantic image, and the depth image, since the real RGB image, the real semantic image, and the real depth image are known, a target loss can be obtained based on these images. In the following, an arbitrary pixel point in the (predicted) RGB image captured by the camera is taken as an example for illustrative introduction, and the pixel point is referred to as a target pixel point. The real RGB image contains a real color of the target pixel point, the real semantic image contains a real category of the target pixel point, and the real depth image contains a real depth of the target pixel point. Then, the target loss can be obtained in the following manner:
[0222] The color of the target pixel point and the real color of the target pixel point are calculated by the first loss function, so as to obtain a first loss, the first loss being used to indicate the difference between the color of the target pixel point and the real color of the target pixel point.
[0223] The category of the target pixel point and the real category of the target pixel point are calculated by the second loss function, so as to obtain a second loss, the second loss being used to indicate the difference between the category of the target pixel point and the real category of the target pixel point.
[0224] The depth of the target pixel point and the real depth of the target pixel point are calculated by the third loss function, so as to obtain a third loss, the third loss being used to indicate the difference between the depth of the target pixel point and the real depth of the target pixel point.
[0225] The transparency distribution of the plurality of sampling points and the real transparency distribution (also known) of the plurality of sampling points are calculated by the fourth loss function, so as to obtain a fourth loss.
[0226] The first loss, the second loss, the third loss and the fourth loss are superimposed, so as to obtain a target loss.
[0227] For example, assuming that there is a neural implicit field to be trained, after inputting the pose of the camera at a certain moment to the neural implicit field to be trained, a predicted RGB image, a predicted semantic image and a predicted depth image photographed by the camera in the pose can be obtained. Since the real RGB image, the real semantic image and the real depth image photographed by the camera in the pose are known, the target loss can be calculated by the formula:
[0228] RL(u,v)=u(u,v)*MSE_Loss(r(u,v),r_gt(u,v))
[0229] DL(u,v)=u(u,v)*L2_Loss(d(u,v),d_gt(u,v))
[0230] SL(u,v)=u(u,v)*CE_Loss(s(u,v),s_gt(u,v))
[0231]
[0232] In the above formula, (u, v) is a certain pixel point in the predicted RGB image, p iFor a certain sample point in the N sample points on the ray sent to (u, v) in the pose of the camera, u (u, v) is the evaluation value of (u, v), r (u, v) is the predicted color of (u, v), r_gt (u, v) is the true color of (u, v) (obtained from the real RGB image), s (u, v) is the predicted category of (u, v), s_gt (u, v) is the true category of (u, v) (obtained from the real semantic image), d (u, v) is the predicted depth of (u, v), and d_gt (u, v) is the true depth of (u, v) (obtained from the real depth image). P (p i ) is the probability density distribution (i.e., the transparency distribution) of the N sample points, and Q (p i ) is the target probability density distribution (i.e., the true transparency distribution, for example, a Gaussian probability distribution, etc.) of the N sample points.
[0233] After obtaining the four losses RL (u, v), DL (u, v), SL (u, v), and UL (u, v), the target loss can be obtained by superimposing the four losses.
[0234] 1103、Based on the target loss, the parameters of the to-be-trained model are updated until the model training condition is met, and the first neural network model is obtained.
[0235] After obtaining the target loss, the parameters of the to-be-trained model can be updated using the target loss, and the to-be-trained model with the updated parameters can be further trained using the next batch of training data until the model training condition (for example, the target loss converges, etc.) is met, and the first neural network model (for example, a trained neural implicit field) in the embodiment shown in FIG. 10 can be obtained. Figure 4
[0236] In addition, the embodiments of the present application can also be compared with related technologies, mainly comparing the training speed of the neural network model and the accuracy of the depth estimation performed by the trained model, and the comparison results are shown in Table 1 and Table 2:
[0237] Table 1
[0238]
[0239]
[0240] Table 2
[0241] Method Time Accuracy Index One Accuracy Index Two Related Art Five 14 min 0.54 0.71 Embodiments of the Present Application 8 min 0.17 0.24
[0242] Based on Table 1 and Table 2, it can be known that the training speed and performance of the model trained by the embodiments of the present application are superior to those of related technologies.
[0243] The above is a detailed description of the point cloud acquisition method and the model training method provided by the embodiments of the present application. In the following, the point cloud acquisition device and the model training device provided by the embodiments of the present application will be introduced. Figure 12 A structural schematic diagram of the point cloud acquisition device provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the point cloud acquisition device includes: Figure 12
[0244] The acquisition module 1201 is configured to acquire a first pose of a camera and a second pose of the camera, the first pose being associated with the second pose in time.
[0245] The first processing module 1202 is configured to process the first pose by using a first neural network model to obtain a first image captured by the camera, and convert the first image to obtain a first three-dimensional point cloud.
[0246] The second processing module 1203 is configured to process the second pose by using the first neural network model to obtain a second image captured by the camera, and convert the second image to obtain a second three-dimensional point cloud.
[0247] The fusion module 1204 is configured to fuse the first three-dimensional point cloud and the second three-dimensional point cloud to obtain a first four-dimensional point cloud.
[0248] In a possible implementation manner, the first image includes a first RGB image captured by the camera, a first semantic image corresponding to the first RGB image, and a first depth image corresponding to the first RGB image, wherein the first semantic image includes categories of each pixel point in the first RGB image, and the first depth image includes depths of each pixel point in the first RGB image.
[0249] In a possible implementation manner, the second image includes a second RGB image captured by the camera, a second semantic image corresponding to the second RGB image, and a second depth image corresponding to the second RGB image, wherein the second semantic image includes categories of each pixel point in the second RGB image, and the second depth image includes depths of each pixel point in the second RGB image.
[0250] In a possible implementation, the first RGB image includes a first pixel point, the first pixel point being any one of the pixel points in the first RGB image, and the first processing module 1202 is configured to: acquire, by using the first neural network model, a first ray sent to the first pixel point by the camera in the first pose; sample, by using the first neural network model, the first ray to obtain a plurality of first sampling points; process, by using the first neural network model, the plurality of first sampling points to obtain colors of the plurality of first sampling points, categories of the plurality of first sampling points, depths of the plurality of first sampling points, and transparencies of the plurality of first sampling points; and acquire, based on the colors of the plurality of first sampling points, the categories of the plurality of first sampling points, the depths of the plurality of first sampling points, and the transparencies of the plurality of first sampling points, a color of the first pixel point, a category of the first pixel point, and a depth of the first pixel point.
[0251] In a possible implementation, the second RGB image includes a second pixel point, the second pixel point being any one of the pixel points in the second RGB image, and the second processing module 1203 is configured to: acquire, by using the first neural network model, a second ray sent to the second pixel point by the camera in the second pose; sample, by using the first neural network model, the second ray to obtain a plurality of second sampling points; process, by using the first neural network model, the plurality of second sampling points to obtain colors of the plurality of second sampling points, categories of the plurality of second sampling points, depths of the plurality of second sampling points, and transparencies of the plurality of second sampling points; and acquire, based on the colors of the plurality of second sampling points, the categories of the plurality of second sampling points, the depths of the plurality of second sampling points, and the transparencies of the plurality of second sampling points, a color of the second pixel point, a category of the second pixel point, and a depth of the second pixel point.
[0252] In a possible implementation, based on the colors of the plurality of first sampling points, the categories of the plurality of first sampling points, the depths of the plurality of first sampling points, and the transparencies of the plurality of first sampling points, the first processing module 1202 is configured to: calculate the colors of the plurality of first sampling points and the transparencies of the plurality of first sampling points to obtain a color of the first pixel point; calculate the categories of the plurality of first sampling points and the transparencies of the plurality of first sampling points to obtain a category of the first pixel point; and calculate the depths of the plurality of first sampling points and the transparencies of the plurality of first sampling points to obtain a depth of the first pixel point.
[0253] In a possible implementation, the second processing module 1203 is configured to: calculate the color of the second pixel point based on the color of the plurality of second sampling points, the category of the plurality of second sampling points, the depth of the plurality of second sampling points, and the transparency of the plurality of second sampling points; and calculate the category of the second pixel point based on the category of the plurality of second sampling points and the transparency of the plurality of second sampling points.
[0254] In a possible implementation, the apparatus further includes: a first evaluation module, configured to calculate the depth of the first pixel point based on the depth of the plurality of first sampling points and the depth of the first pixel point, to obtain an evaluation value of the first pixel point; a second evaluation module, configured to calculate the depth of the second pixel point based on the depth of the plurality of second sampling points and the depth of the second pixel point, to obtain an evaluation value of the second pixel point; and a fusion module 1204, configured to: take the evaluation value of each pixel point in the first RGB image as an evaluation value of each spatial point in the first three-dimensional point cloud, and remove spatial points with evaluation values less than an evaluation threshold from the first three-dimensional point cloud to obtain a denoised first three-dimensional point cloud; take the evaluation value of each pixel point in the second RGB image as an evaluation value of each spatial point in the second three-dimensional point cloud, and remove spatial points with evaluation values less than the evaluation threshold from the second three-dimensional point cloud to obtain a denoised second three-dimensional point cloud; and fuse the denoised first three-dimensional point cloud and the denoised second three-dimensional point cloud to obtain the first four-dimensional point cloud.
[0255] In a possible implementation, the apparatus further includes an optimization module, configured to: obtain the second four-dimensional point cloud by processing the first image and the second image through the second neural network model; and optimize the first four-dimensional point cloud based on the second four-dimensional point cloud to obtain an optimized first four-dimensional point cloud.
[0256] In a possible implementation, the first neural network model is a neural implicit field.
[0257] Figure 13 A structural schematic diagram of a model training apparatus provided by an embodiment of the present application is shown in FIG. 13. Figure 13 As shown in FIG. 13, the model training apparatus includes:
[0258] The processing module 1301 is configured to process the pose of the camera through the to-be-trained model to obtain an image captured by the camera, the image being used to acquire a three-dimensional point cloud, and the three-dimensional point cloud being used to acquire a four-dimensional point cloud.
[0259] The acquisition module 1302 is configured to acquire a target loss based on the image and a real image corresponding to the image, the target loss being used to indicate a difference between the image and the real image.
[0260] The training module 1303 is configured to update parameters of the to-be-trained model based on the target loss until a model training condition is met, to obtain the first neural network model.
[0261] In a possible implementation, the image includes an RGB image captured by a camera, a semantic image corresponding to the RGB image, and a depth image corresponding to the RGB image, where the semantic image includes categories of respective pixels in the RGB image, and the depth image includes depths of respective pixels in the RGB image.
[0262] In a possible implementation, the RGB image includes a target pixel, and the target pixel is any pixel in the RGB image. The processing module 1301 is configured to: obtain, by using the to-be-trained model, a ray sent by the camera in the pose to the target pixel; sample, by using the to-be-trained model, the ray to obtain a plurality of sampling points; process, by using the to-be-trained model, the plurality of sampling points to obtain colors of the plurality of sampling points, categories of the plurality of sampling points, depths of the plurality of sampling points, and transparencies of the plurality of sampling points; and obtain, based on the colors of the plurality of sampling points, the categories of the plurality of sampling points, the depths of the plurality of sampling points, and the transparencies of the plurality of sampling points, a color of the target pixel, a category of the target pixel, and a depth of the target pixel.
[0263] In a possible implementation, the processing module 1301 is configured to: calculate the colors of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the color of the target pixel; calculate the categories of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the category of the target pixel; and calculate the depths of the plurality of sampling points and the transparencies of the plurality of sampling points to obtain the depth of the target pixel.
[0264] In a possible implementation, the real image includes a real RGB image, a real semantic image, and a real depth image, the real RGB image includes a real color of the target pixel, the real semantic image includes a real category of the target pixel, and the real depth image includes a real depth of the target pixel. The obtaining module 1302 is configured to: obtain, based on the color of the target pixel and the real color of the target pixel, a first loss, where the first loss is used to indicate a difference between the color of the target pixel and the real color of the target pixel; obtain, based on the category of the target pixel and the real category of the target pixel, a second loss, where the second loss is used to indicate a difference between the category of the target pixel and the real category of the target pixel; obtain, based on the depth of the target pixel and the real depth of the target pixel, a third loss, where the third loss is used to indicate a difference between the depth of the target pixel and the real depth of the target pixel; and obtain, based on the first loss, the second loss, and the third loss, the target loss.
[0265] In one possible implementation, the acquisition module 1302 is further used to obtain a fourth loss based on the transparency distribution of multiple sampling points and the actual transparency distribution of multiple sampling points; the acquisition module 1302 is used to obtain a target loss based on the first loss, the second loss, the third loss and the fourth loss.
[0266] In one possible implementation, the first neural network model is a neural implicit field.
[0267] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and no further details will be given here.
[0268] The embodiment of the present application also relates to an execution device, Figure 14 A schematic diagram of the structure of the execution device provided in the embodiment of the present application. Figure 14 As shown, the execution device 1400 can be specifically a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Figure 11 The consumption prediction device described in the corresponding embodiment is used to implement Figure 5 The consumption prediction function in the corresponding embodiment. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403 and a memory 1404 (wherein the number of the processor 1403 in the execution device 1400 can be one or more, Figure 14 (taking one processor as an example), the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403 and the memory 1404 may be connected via a bus or other means.
[0269] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0270] The processor 1403 controls the operation of the execution device. In a specific application, various components of the execution device are coupled together by a bus system, which can include a data bus, a power bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, various buses are referred to as a bus system in the figure.
[0271] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with a signal processing capability. In the implementation process, the steps of the above method can be completed by hardware integrated logic circuits in the processor 1403 or by instructions in the form of software. The processor 1403 described above can be a general processor, a digital signal processor (DSP), a microprocessor or a microcontroller. It can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The processor 1403 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 1404, and the processor 1403 reads the information in the memory 1404 and combines the hardware to complete the steps of the above method.
[0272] The receiver 1401 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the execution device. The transmitter 1402 can be used to output digital or character information through the first interface; the transmitter 1402 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1402 can also include a display device such as a display screen.
[0273] In an embodiment of the present application, in one case, the processor 1403 is configured to obtain the image captured by the camera to obtain the three-dimensional point cloud, and further obtain the four-dimensional point cloud. Figure 4 In the corresponding embodiment, the first neural network model is used to obtain the image captured by the camera to obtain the three-dimensional point cloud, and further obtain the four-dimensional point cloud.
[0274] The present application also relates to a training device. Figure 15 A structural diagram of the training device provided in the embodiment of the present application. Figure 15 As shown, training device 1500 is implemented by one or more servers. Training device 1500 may vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1515 (e.g., one or more processors) and memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) storing application programs 1542 or data 1544. Memory 1532 and storage media 1530 may be either ephemeral or persistent storage. The program stored in storage media 1530 may include one or more modules (not shown), each of which may include a series of instruction operations on the training device. Furthermore, CPU 1515 may be configured to communicate with storage media 1530 to execute the series of instruction operations in storage media 1530 on training device 1500.
[0275] The training device 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input and output interfaces 1558; or, one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0276] Specifically, the training device can perform Figure 11 The model training method in the corresponding embodiment.
[0277] An embodiment of the present application also relates to a computer storage medium, which stores a program for signal processing. When the computer storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0278] An embodiment of the present application also relates to a computer program product, which stores instructions that, when executed by a computer, enable the computer to execute the steps executed by the aforementioned execution device, or enable the computer to execute the steps executed by the aforementioned training device.
[0279] The execution device, the training device or the terminal device provided by the embodiments of the present application can be a chip, which includes a processing unit, for example, a processor, and a communication unit, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute computer execution instructions stored in a storage unit, so that the chip in the execution device executes the data processing method described in the above embodiments, or so that the chip in the training device executes the data processing method described in the above embodiments. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0280] Specifically, refer to Figure 16 , Figure 16 A structural diagram of the chip provided by the embodiments of the present application is shown in FIG. 16. The chip can be a neural network processor NPU 1600, which is mounted on a host CPU (Host CPU) as a coprocessor and is assigned tasks by the Host CPU. The core part of the NPU is an operation circuit 1603, which extracts matrix data in a memory and performs multiplication operation under the control of a controller 1604.
[0281] In some implementations, the operation circuit 1603 internally includes a plurality of processing units (PEs). In some implementations, the operation circuit 1603 is a two-dimensional systolic array. The operation circuit 1603 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1603 is a general-purpose matrix processor.
[0282] For example, assuming that there are an input matrix A, a weight matrix B and an output matrix C. The operation circuit takes the corresponding data of the matrix B from the weight memory 1602 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1601 and performs matrix operation with the matrix B, and the partial result or final result of the obtained matrix is saved in an accumulator 1608.
[0283] Unified memory 1606 is used to store input and output data. Weight data is directly transferred to weight memory 1602 through the Direct Memory Access Controller (DMAC) 1605. Input data is also transferred to unified memory 1606 through the DMAC.
[0284] BIU stands for Bus Interface Unit 1614 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1609 .
[0285] The bus interface unit 1614 (BIU) is used for the instruction fetch memory 1609 to obtain instructions from the external memory, and is also used for the storage unit access controller 1605 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0286] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1606 or move weight data to the weight memory 1602 or move input data to the input memory 1601.
[0287] The vector calculation unit 1607 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1603, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of the predicted label plane.
[0288] In some implementations, the vector calculation unit 1607 can store the processed output vector to the unified memory 1606. For example, the vector calculation unit 1607 can apply a linear function or a nonlinear function to the output of the operation circuit 1603, such as linear interpolation of the predicted label plane extracted by the convolution layer, or another example is a vector of accumulated values to generate an activation value. In some implementations, the vector calculation unit 1607 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1603, for example, for use in subsequent layers in a neural network.
[0289] An instruction fetch buffer 1609 connected to the controller 1604 is used to store instructions used by the controller 1604;
[0290] The unified memory 1606, the input memory 1601, the weight memory 1602, and the instruction memory 1609 are on-chip memories. The external memory is private to the NPU hardware architecture.
[0291] Any processor mentioned in the above can be a general central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of the above programs.
[0292] It should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0293] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions for making a computer device (which can be a personal computer, a training device, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0294] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in the form of a computer program product in whole or in part.
[0295] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. A point cloud acquisition method, characterized in that, The method comprises: obtaining a first pose of a camera and a second pose of the camera, the first pose being associated in time with the second pose; processing the first pose by a first neural network model to obtain a first image captured by the camera, and converting the first image to obtain a first three-dimensional point cloud; processing the second pose by the first neural network model to obtain a second image captured by the camera, and converting the second image to obtain a second three-dimensional point cloud; fusing the first three-dimensional point cloud and the second three-dimensional point cloud to obtain a first four-dimensional point cloud.
2. The method of claim 1, wherein, The first image comprises a first red-green-blue (RGB) image captured by the camera, a first semantic image corresponding to the first RGB image, and a first depth image corresponding to the first RGB image, wherein the first semantic image comprises a category of each pixel point in the first RGB image, and the first depth image comprises a depth of each pixel point in the first RGB image.
3. The method of claim 2, wherein, The second image comprises a second RGB image captured by the camera, a second semantic image corresponding to the second RGB image, and a second depth image corresponding to the second RGB image, wherein the second semantic image comprises a category of each pixel point in the second RGB image, and the second depth image comprises a depth of each pixel point in the second RGB image.
4. The method of claim 3, wherein, The first RGB image comprises a first pixel point, the first pixel point being any one of the pixel points in the first RGB image, and the processing of the first pose by the first neural network model to obtain the first image captured by the camera comprises: obtaining, by the first neural network model, a first ray sent to the first pixel point by the camera in the first pose; sampling, by the first neural network model, the first ray to obtain a plurality of first sampling points; processing, by the first neural network model, the plurality of first sampling points to obtain colors of the plurality of first sampling points, categories of the plurality of first sampling points, depths of the plurality of first sampling points, and transparencies of the plurality of first sampling points; based on the colors of the plurality of first sampling points, the categories of the plurality of first sampling points, the depths of the plurality of first sampling points, and the transparencies of the plurality of first sampling points, obtaining a color of the first pixel point, a category of the first pixel point, and a depth of the first pixel point.
5. The method of claim 4, wherein, The second RGB image comprises a second pixel point, the second pixel point being any one of the pixel points in the second RGB image, and the processing of the second pose by the first neural network model to obtain the second image captured by the camera comprises: obtaining, by the first neural network model, a second ray sent to the second pixel point by the camera in the second pose; sampling, by the first neural network model, the second ray to obtain a plurality of second sampling points; obtaining the color, the category, the depth and the transparency of the second sampling points by processing the second sampling points through the first neural network model; obtaining the color, the category and the depth of the second pixel point based on the color, the category, the depth and the transparency of the second sampling points.
6. The method of claim 4, wherein, The obtaining the color, the category and the depth of the first pixel point based on the color, the category, the depth and the transparency of the first sampling points comprises: calculating the color and the transparency of the first sampling points to obtain the color of the first pixel point; calculating the category and the transparency of the first sampling points to obtain the category of the first pixel point; calculating the depth and the transparency of the first sampling points to obtain the depth of the first pixel point.
7. The method according to claim 4 or 5, characterized in that, The obtaining the color, the category and the depth of the second pixel point based on the color, the category, the depth and the transparency of the second sampling points comprises: calculating the color and the transparency of the second sampling points to obtain the color of the second pixel point; calculating the category and the transparency of the second sampling points to obtain the category of the second pixel point; calculating the depth and the transparency of the second sampling points to obtain the depth of the second pixel point.
8. The method according to any one of claims 5 to 7, characterized in that, The method further comprises: calculating the depth of the first sampling points and the depth of the first pixel point to obtain the evaluation value of the first pixel point; calculating the depth of the second sampling points and the depth of the second pixel point to obtain the evaluation value of the second pixel point; The fusing the first three-dimensional point cloud and the second three-dimensional point cloud to obtain a first four-dimensional point cloud comprises: taking the evaluation value of each pixel point in the first RGB image as the evaluation value of each spatial point in the first three-dimensional point cloud, removing spatial points with evaluation values less than an evaluation threshold from the first three-dimensional point cloud to obtain a denoised first three-dimensional point cloud; taking the evaluation value of each pixel point in the second RGB image as the evaluation value of each spatial point in the second three-dimensional point cloud, removing spatial points with evaluation values less than the evaluation threshold from the second three-dimensional point cloud to obtain a denoised second three-dimensional point cloud; fusing the denoised first three-dimensional point cloud and the denoised second three-dimensional point cloud to obtain a first four-dimensional point cloud.
9. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: obtaining the second four-dimensional point cloud by processing the first image and the second image through a second neural network model; optimizing the first four-dimensional point cloud based on the second four-dimensional point cloud to obtain an optimized first four-dimensional point cloud.
10. The method according to any one of claims 1 to 9, characterized in that, The first neural network model is a neural implicit field.
11. A model training method, comprising: The method comprises: processing the pose of the camera through a to-be-trained model to obtain an image captured by the camera, the image being used to obtain a three-dimensional point cloud, and the three-dimensional point cloud being used to obtain a four-dimensional point cloud; obtaining a target loss based on the image and a real image corresponding to the image, the target loss being used to indicate the difference between the image and the real image; updating the parameters of the to-be-trained model based on the target loss until a model training condition is met to obtain a first neural network model.
12. The method of claim 11, wherein, The image comprises an RGB image captured by the camera, a semantic image corresponding to the RGB image, and a depth image corresponding to the RGB image, wherein the semantic image comprises the category of each pixel point in the RGB image, and the depth image comprises the depth of each pixel point in the RGB image.
13. The method of claim 12, wherein, The RGB image comprises a target pixel point, and the target pixel point is any one of the pixel points in the RGB image. The processing of the pose through the first neural network model to obtain the image captured by the camera comprises: obtaining a ray sent by the camera in the pose to the target pixel point through a to-be-trained model; sampling the ray through the to-be-trained model to obtain a plurality of sampling points; processing the plurality of sampling points through the to-be-trained model to obtain the color, category, depth, and transparency of the plurality of sampling points; 14. The method of claim 13, wherein, obtaining the color, category, and depth of the target pixel point based on the color, category, depth, and transparency of the plurality of sampling points. The obtaining of the color, category, and depth of the target pixel point based on the color, category, depth, and transparency of the plurality of sampling points comprises: calculating the color and transparency of the plurality of sampling points to obtain the color of the target pixel point; calculating the category and transparency of the plurality of sampling points to obtain the category of the target pixel point; 15. The method according to claim 13 or 14, characterized in that, calculating the depth and transparency of the plurality of sampling points to obtain the depth of the target pixel point. The real image comprises a real RGB image, a real semantic image, and a real depth image, the real RGB image comprising the real color of the target pixel point, the real semantic image comprising the real category of the target pixel point, and the real depth image comprising the real depth of the target pixel point. The obtaining of the target loss based on the image and the real image corresponding to the image comprises: obtaining a first loss based on the color of the target pixel point and the real color of the target pixel point, the first loss being used to indicate a difference between the color of the target pixel point and the real color of the target pixel point; obtaining a second loss based on the category of the target pixel point and the real category of the target pixel point, the second loss being used to indicate a difference between the category of the target pixel point and the real category of the target pixel point; obtaining a third loss based on the depth of the target pixel point and the real depth of the target pixel point, the third loss being used to indicate a difference between the depth of the target pixel point and the real depth of the target pixel point; obtaining a target loss based on the first loss, the second loss and the third loss.
16. The method of claim 15, wherein, The obtaining the target loss based on the image and the real image corresponding to the image further includes: obtaining a fourth loss based on the transparency distribution of the plurality of sampling points and the real transparency distribution of the plurality of sampling points; The obtaining the target loss based on the first loss, the second loss and the third loss includes: obtaining the target loss based on the first loss, the second loss, the third loss and the fourth loss.
17. The method according to any one of claims 11 to 16, characterized in that, The first neural network model is a neural implicit field.
18. A point cloud acquisition apparatus, characterized by comprising: The device includes: an obtaining module, configured to obtain a first pose of a camera and a second pose of the camera, the first pose being associated in time with the second pose; a first processing module, configured to process the first pose by a first neural network model to obtain a first image captured by the camera, and convert the first image to obtain a first three-dimensional point cloud; a second processing module, configured to process the second pose by the first neural network model to obtain a second image captured by the camera, and convert the second image to obtain a second three-dimensional point cloud; a fusion module, configured to fuse the first three-dimensional point cloud and the second three-dimensional point cloud to obtain a first four-dimensional point cloud.
19. A model training apparatus, comprising: The device includes: a processing module, configured to process a pose of a camera by a to-be-trained model to obtain an image captured by the camera, the image being used to obtain a three-dimensional point cloud, and the three-dimensional point cloud being used to obtain a four-dimensional point cloud; an obtaining module, configured to obtain a target loss based on the image and a real image corresponding to the image, the target loss being used to indicate a difference between the image and the real image; a training module, configured to update parameters of the to-be-trained model based on the target loss until a model training condition is met to obtain a first neural network model.
20. A point cloud acquisition apparatus, comprising: The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, when the code is executed, the point cloud obtaining device executes the method in any one of claims 1 to 17.
21. A computer storage medium, comprising, The computer storage medium stores one or more instructions, which, when executed by one or more computers, cause the one or more computers to implement the method in any one of claims 1 to 17.
22. A computer program product, characterised in that, The computer program product stores instructions which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 17.