A method, apparatus, electronic device, and storage medium for generating point clouds

By generating bird's-eye view features and three-dimensional geometric representation data, and combining bicycle motion conditions to predict future point clouds, the problem of insufficient information utilization in autonomous driving is solved, and the accuracy of point cloud prediction and model reliability are improved.

CN117788277BActive Publication Date: 2025-07-08SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311816094.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-08
Estimated Expiration
2043-12-26

AI Technical Summary

Technical Problem

Existing vision processing technologies cannot fully utilize the semantic, geometric and dynamic timing information in image point cloud sequences in autonomous driving, resulting in insufficient point cloud prediction efficiency and accuracy.

Method used

By obtaining the historical timing image data of the target vehicle, generating bird's-eye view feature information, and using neural network models to extract three-dimensional geometric representation data, combining the bicycle motion conditions to predict future point clouds, achieving full utilization of information in the image point cloud sequence.

Benefits of technology

It improves the accuracy of point cloud prediction and the reliability of autonomous driving models, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117788277B_ABST
    Figure CN117788277B_ABST
Patent Text Reader

Abstract

The present invention discloses a point cloud generation method, apparatus, electronic device, and storage medium. The method includes: obtaining historical sequential image data of a target vehicle and determining bird's-eye view feature information of the historical sequential image data; collecting three-dimensional geometric representation data of the historical sequential image data within the bird's-eye view feature information; and predicting and generating a future point cloud according to the ego-motion condition of the target vehicle and the three-dimensional geometric representation data. The embodiments of the present invention implement visual image processing for an autonomous driving model, fully utilize semantic, geometric features, and dynamic sequential information within an image point cloud sequence, can improve the accuracy of point cloud prediction and generation, can improve the reliability of the autonomous driving model, and enhance the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, electronic device and storage medium for generating point clouds. Background Art

[0002] With the development of vision processing technology, autonomous driving is about to become a reality. Currently, vision processing technology mainly focuses on the research of general vision. However, there is still relatively little research on how to self-supervised pre-train an autonomous driving model, which means that common vision processing cannot simultaneously cover features such as semantics, geometry, and time series. There are still deficiencies in vision processing in aspects such as end-to-end perception, prediction, and planning. How to make full use of the rich semantic, geometric features, and dynamic time series information in the image point cloud sequence to improve the efficiency and accuracy of point cloud prediction has become an urgent problem to be solved in the industry. Summary of the Invention

[0003] The present invention provides a method, device, electronic device and storage medium for generating point clouds, so as to implement visual image processing for an autonomous driving model, make full use of semantic, geometric features, and dynamic time series information in the image point cloud sequence, improve the accuracy of point cloud prediction generation, improve the reliability of the autonomous driving model, and enhance the user experience.

[0004] According to one aspect of the present invention, a method for generating point clouds is provided, wherein the method includes:

[0005] Obtain historical time-series image data of a target vehicle, and determine bird's-eye view feature information of the historical time-series image data;

[0006] Collect three-dimensional geometric representation data of the historical time-series image data within the bird's-eye view feature information;

[0007] Predict and generate future point clouds according to the ego-motion conditions of the target vehicle and the three-dimensional geometric representation data.

[0008] According to another aspect of the present invention, a device for generating point clouds is provided, wherein the device includes:

[0009] An encoder module, configured to obtain historical time-series image data of a target vehicle, and determine bird's-eye view feature information of the historical time-series image data;

[0010] A latent space renderer module, configured to collect three-dimensional geometric representation data of the historical time-series image data within the bird's-eye view feature information;

[0011] A decoder module, configured to predict and generate future point clouds according to the ego-motion conditions of the target vehicle and the three-dimensional geometric representation data.

[0012] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the point cloud generation method according to any embodiment of the present invention.

[0016] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the point cloud generation method according to any embodiment of the present invention when executed.

[0017] The technical solution of the embodiment of the present invention can achieve visual image processing in the scenario of an autonomous driving model by obtaining historical sequential image data in a target vehicle, generating bird's-eye view feature information corresponding to the historical sequential image data, extracting three-dimensional geometric representation data corresponding to the historical sequential image data according to the bird's-eye view feature information, and predicting and generating future point clouds according to the three-dimensional geometric representation data under the condition of the self-vehicle movement of the target vehicle. It can make full use of semantic, geometric features and dynamic sequential information in the image point cloud sequence, improve the accuracy of point cloud prediction and generation, improve the reliability of the autonomous driving model, and enhance the user experience.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0020] Figure 1 is a flowchart of a point cloud generation method provided in Embodiment 1 of the present invention;

[0021] Figure 2 is a flowchart of another point cloud generation method provided in Embodiment 2 of the present invention;

[0022] Figure 3It is a schematic diagram of the structural framework of a point cloud generation method provided in Embodiment 3 of the present invention;

[0023] Figure 4 It is a schematic structural diagram of a latent space renderer provided in Embodiment 3 of the present invention;

[0024] Figure 5 It is a schematic structural diagram of a decoder provided in Embodiment 3 of the present invention;

[0025] Figure 6 It is a schematic structural diagram of a point cloud generation device provided in Embodiment 4 of the present invention;

[0026] Figure 7 It is a schematic structural diagram of an electronic device implementing the embodiments of the present invention. Detailed implementation manners

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0029] Embodiment 1

[0030] Figure 1 It is a flowchart of a point cloud generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to a situation. This method can be executed by a point cloud generation device, and the point cloud generation device can be implemented in the form of hardware and / or software. The point cloud generation device can be configured in a server or a server cluster. As Figure 1 shown, the method includes:

[0031] Step 110: Obtain the historical sequential image data of the target vehicle and determine the bird's-eye view feature information of the historical sequential image data.

[0032] Among them, the target vehicle can be the main body for autonomous driving. The target vehicle can collect surrounding environment data at different moments during the autonomous driving process. The historical sequential image data can be generated by organizing the surrounding environment data collected by the target vehicle. The historical sequential image data can be a data set. The image frames in the historical sequential image data can be arranged in chronological order. The historical sequential image data can reflect the driving state of the target vehicle in the past period of time. It can be understood that the time length in the historical sequential image data can be set according to specific business needs. The higher the accuracy requirement for autonomous driving, the longer or more fine-grained the time length of the image frames included in the historical sequential image data can be. The bird's-eye view feature information can be a bird's-eye view generated by processing two-dimensional images from one or more perspectives, which can improve the association of object and spatial relationships in the scene and facilitate the three-dimensional geometric feature processing of the historical sequential image data.

[0033] In an embodiment of the present invention, the historical sequential image data composed of the environmental image data collected by the target vehicle at different moments during the autonomous driving process can be extracted. The environmental image data collected at different moments in the historical sequential image data can be processed to convert one or more frames of two-dimensional environmental image data at different moments into a bird's-eye view, and the bird's-eye view can be used as the bird's-eye view feature information corresponding to the historical sequential image data respectively.

[0034] Step 120: Collect the three-dimensional geometric representation data of the historical sequential image data in the bird's-eye view feature information.

[0035] Among them, the three-dimensional geometric representation data can be feature data reflecting the three-dimensional geometric relationship in space. The three-dimensional geometric representation data can include the occupancy probability of the target vehicle in the grid, the geometric shape of the driving path of the target vehicle, etc.

[0036] In an embodiment of the present invention, three-dimensional geometric feature extraction can be performed on the extracted bird's-eye view feature information, and the extracted three-dimensional geometric features can be used as the three-dimensional geometric representation data. It can be understood that the historical sequential image data can correspond to multiple frames of bird's-eye view feature information, and various corresponding three-dimensional geometric representation data can exist for each frame of bird's-eye view feature information. Specifically, the three-dimensional geometric representation data can be obtained by extracting through a neural network model. For example, in some embodiments of the invention, the three-dimensional geometric representation data can be generated by processing with a long short-term memory model. The three-dimensional geometric representation data at different moments can be generated by processing the three-dimensional geometric representation data at the previous moment and the bird's-eye view feature information at the current moment.

[0037] Step 130: Predict and generate future point clouds based on the ego-motion conditions of the target vehicle and the three-dimensional geometric representation data.

[0038] Among them, the ego-motion conditions can be information on the operating states of the target vehicle at different times, and the ego-motion conditions can include, but are not limited to, information such as vehicle speed, vehicle driving direction, and vehicle position. The future point clouds can be point cloud data generated by prediction, and the future point clouds can represent the point cloud data of the state of the target vehicle at one or more future times.

[0039] In the embodiments of the present invention, information such as the vehicle speed, vehicle driving direction, and vehicle position of the target vehicle can be extracted as the ego-motion conditions. It can be understood that the ego-motion conditions can include the ego-motion conditions within a historical period of time, and can also include the ego-motion conditions for a future period of time predicted based on the historical ego-motion conditions. The future point clouds of the target vehicle can be predicted through the ego-motion conditions and the pre-acquired three-dimensional geometric representation data. The time at which the future point clouds are located can include the next frame or the next few frames of future times, and this time can be determined by the specific business scenario.

[0040] In the embodiments of the present invention, by acquiring historical sequential image data within the target vehicle, generating bird's-eye view feature information corresponding to the historical sequential image data, extracting three-dimensional geometric representation data corresponding to the historical sequential image data according to the bird's-eye view feature information, and predicting and generating future point clouds according to the three-dimensional geometric representation data under the ego-motion conditions of the target vehicle, visual image processing in the scenario of an autonomous driving model can be realized, the full utilization of semantic, geometric features, and dynamic sequential information in the image point cloud sequence can be realized, the accuracy of point cloud prediction and generation can be improved, the reliability of the autonomous driving model can be improved, and the user experience can be enhanced.

[0041] Embodiment Two

[0042] Figure 2 is a flowchart of another method for generating point clouds according to Embodiment Two of the present invention. The embodiments of the present invention are specific implementations based on the above-mentioned embodiments of the invention. Refer to Figure 2 , and the method provided by the embodiments of the present invention specifically includes the following steps:

[0043] Step 210: Extract the image frames collected by the target vehicle in chronological order to form historical sequential image data.

[0044] Among them, the image frames can be environmental images of the environment where the target vehicle is located, and the image frames can be images of the surrounding environment of the target vehicle.

[0045] In an embodiment of the present invention, the target vehicle can collect images of the surrounding environment as image frames in a regular or irregular manner during driving, and can sort the image frames in the chronological order of acquisition or generation of each image frame. It can be understood that there can be one or more image frames at the same moment, and the image frames sorted in chronological order can be used as historical chronological image data.

[0046] Step 220: Input the image frames corresponding to different moments in the historical chronological image data into a preset historical encoder in sequence to generate bird's-eye view feature information of the image frames at different moments, where the preset historical encoder includes a bottom-up bird's-eye view model, a top-down bird's-eye view model, a multi-modal bird's-eye view model, and a decoding bird's-eye view model.

[0047] Among them, the preset historical encoder can be an encoder that processes one or more frames of two-dimensional images into a bird's-eye view. The preset historical encoder can include one or more of a bottom-up bird's-eye view model, a top-down bird's-eye view model, a multi-modal bird's-eye view model, and a decoding bird's-eye view model. The bottom-up bird's-eye view model can include models such as Lift-Splat-Shoot model and BEVDet model. The top-down bird's-eye view model can include models such as DETR model and PETR model. The multi-modal bird's-eye view model can include bevfusion–ADLab, etc. The decoding bird's-eye view model can include models such as CenterNet model.

[0048] In an embodiment of the present invention, the image frames at different moments in the historical chronological image data can be separately input into the preset historical encoder for processing. It can be understood that the image frames at the same moment can be input into the preset historical encoder for processing simultaneously. The preset historical encoder can perform bird's-eye view conversion on the image frames, and can extract the bird's-eye view feature information of the image frames corresponding to different moments output by the preset historical encoder. In an embodiment of the present invention, the implementation manner of the preset historical encoder can be realized through a bottom-up bird's-eye view model, a top-down bird's-eye view model, a multi-modal bird's-eye view model, and a decoding bird's-eye view model.

[0049] Step 230: Invoke a preset latent space renderer to process the bird's-eye view feature information of each frame.

[0050] Among them, the preset latent space renderer can be a device that processes three-dimensional geometric features in the bird's-eye view feature information. The preset latent space renderer can extract the occupancy probability of each occupied grid in the bird's-eye view feature information as three-dimensional geometric representation data.

[0051] In the embodiments of the present invention, the bird's-eye view feature information of each frame can be respectively input into a preset latent space renderer for processing, so as to extract the three-dimensional geometric representation data of the bird's-eye view feature information of each frame in the preset latent space renderer. In some embodiments of the invention, the preset latent space renderer can have the characteristics of a long short-term memory model. The bird's-eye view feature information can be input into the preset latent space renderer in chronological order. The BEV query of the preset latent space renderer can retain the weight information of the bird's-eye view feature information of each frame for use in the subsequent processing of the bird's-eye view feature information, making the processing of the bird's-eye view feature information of each frame interdependent, thereby retaining the temporal characteristics of the bird's-eye view feature information, and thus improving the accuracy of future point cloud prediction.

[0052] Step 240: Extract the occupancy grid feature generated by the preset latent space renderer corresponding to the bird's-eye view feature information as the three-dimensional geometric representation data.

[0053] Among them, the occupancy grid feature can be the probability feature that the target vehicle occupies each grid in the bird's-eye view feature information. The occupancy grid feature can be generated by the preset latent space renderer. It can be understood that each bird's-eye view feature information can correspond to a set of occupancy grid features, and the set of occupancy grid features can include the occupancy probability of each grid in the bird's-eye view grid.

[0054] In the embodiments of the present invention, the occupancy grid feature generated by the preset latent space renderer processing the bird's-eye view feature information can be extracted, and the occupancy grid feature of each bird's-eye view feature information can be used as the corresponding three-dimensional geometric representation data.

[0055] Step 250: Call a preset multi-layer perceptron to determine the self-vehicle motion condition at the next moment based on the image frame at the current moment in the historical time-series image data.

[0056] Among them, the preset multi-layer perceptron can be a neural network model used to predict the future self-vehicle motion condition of the target vehicle. The preset multi-layer perceptron can be generated by training an artificial neural network. The network structure and weight parameters of the preset multi-layer perceptron are not limited here.

[0057] In the embodiments of the present invention, a pre-trained preset multi-layer perceptron can be obtained. The image frame at the current moment can be extracted from the historical time-series image data, and the preset multi-layer perceptron can be called to process the image frame at the current moment, so that the preset multi-layer perceptron predicts the self-vehicle motion condition at the next moment based on the image frame at the current moment. It can be understood that in some embodiments, the image frame input to the preset multi-layer perceptron is not limited to the current moment, but can also be the image frames within a period of time adjacent to the current moment.

[0058] Step 260: The bird's-eye view query corresponding to the three-dimensional geometric representation data at the current moment, the occupancy grid feature of the three-dimensional geometric representation data at the current moment, and the ego-vehicle motion condition are sequentially processed through a self-attention module, a temporal cross-attention module, and a feed-forward network for at least one iteration to generate future point clouds.

[0059] Among them, the bird's-eye view query can be intermediate parameter information retained during the generation process of three-dimensional geometric representation data at different moments. The bird's-eye view query can be generated during the processing of each frame of bird's-eye view feature information by a preset latent space renderer. For example, the bird's-eye view query at the current moment can be generated by the preset latent space renderer based on the bird's-eye view query at the previous moment and the bird's-eye view feature information at the current moment. The bird's-eye view query at the current moment can reflect the temporal characteristics of the bird's-eye view feature information at the current moment.

[0060] The self-attention module can be composed of a self-attention mechanism. When the self-attention base station processes sequential data, each element can establish an association with other elements in the sequence, and can adaptively capture the long-range dependence relationship between elements by calculating the relative importance between elements in the sequence. Specifically, the self-attention module can calculate the similarity between elements in the sequence and normalize these similarities into attention weights, and each element can be weighted and summed with the corresponding attention weight to obtain the output. The temporal cross-attention module can be a model that calculates the attention between a certain element in one sequence and all elements in another sequence. The above one sequence and another sequence can be the occupancy grid feature and the ego-vehicle motion condition provided in the embodiments of the present invention. The feed-forward network can be a pre-trained neural network model. The feed-forward network can predict the future point cloud, and the feed-forward network can include an input layer, a hidden layer, and an output layer.

[0061] In the embodiments of the present invention, the bird's-eye view query and the occupancy grid feature of the three-dimensional geometric representation data at the current moment, and the ego-vehicle motion condition at the current moment can be extracted. The ego-vehicle motion condition, the bird's-eye view query, and the occupancy grid feature can be sequentially processed through the self-attention module, the temporal cross-attention module, and the feed-forward network for multiple iterations, and the output result can be used as the future point cloud.

[0062] In an embodiment of the present invention, by sorting the image frames collected by a target vehicle in chronological order into historical chronological image data, inputting the image frames at different times in the historical chronological image data into a preset historical encoder for processing in sequence to generate bird's-eye view feature information, using a preset latent space renderer to process the bird's-eye view feature information to extract occupancy grid features as three-dimensional geometric representation data, calling a preset multi-layer perceptron to predict the motion condition of the ego vehicle at the next moment according to the image frame at the current moment, and performing multiple iterative processes on the bird's-eye view query, occupancy grid features at the current moment, and the motion condition of the ego vehicle at the next moment based on a self-attention module, a temporal cross-attention module, and a feed-forward network to generate future point clouds, so as to realize the full utilization of semantic, geometric features, and dynamic temporal information in the image point cloud sequence, improve the accuracy of point cloud prediction generation, improve the reliability of the autonomous driving model, and enhance the user experience.

[0063] Further, based on the above-mentioned invention embodiment, calling a preset latent space renderer to process each frame of bird's-eye view feature information includes:

[0064] Accumulating the conditional probabilities of different occupancy grids in the bird's-eye view feature information according to a preset conditional probability function; determining the conditional probability and the ray feature corresponding to the bird's-eye view feature information according to a preset feature expectation function; taking the weighted product of the ray feature and the conditional probability as the occupancy grid feature.

[0065] Among them, the preset conditional probability function can be a pre-configured conditional probability determination function. The preset feature expectation function can be to calculate the joint probability distribution of the rays from the origin of the bird's-eye view corresponding to the bird's-eye view feature information to each grid, and the ray feature can be the joint probability value determined by the preset feature expectation function.

[0066] In the embodiment of the present invention, the occupancy conditional probability of each grid in the bird's-eye view corresponding to the bird's-eye view feature information can be calculated by a preset conditional probability function, and the joint probability distribution of the rays from the origin of the bird's-eye view corresponding to the bird's-eye view feature information to each grid can be calculated by using a preset feature expectation function as the ray feature, and the weighted product of the ray feature and the conditional probability can be taken as the occupancy grid feature.

[0067] Further, based on the above-mentioned invention embodiment, the preset conditional probability function at least includes the following:

[0068]

[0069] Among them, represents the conditional probability of the path point j on the ray from the origin to the grid i in the grid i in the bird's-eye view feature information, p represents the independent occupancy probability of the grids of different bird's-eye view feature information, and the value of p is determined by neural network estimation;

[0070] The preset feature expectation function at least includes the following:

[0071]

[0072] Among them, represents the conditional probability of waypoint k on the ray from the origin to grid i within the bird's-eye view feature information of grid i, represents the bird's-eye view feature information of the k-th image frame.

[0073] Based on the above embodiments of the invention, the self-vehicle motion conditions include at least one of the following: vehicle driving direction, vehicle speed, and vehicle position.

[0074] In the embodiments of the present invention, the self-vehicle motion conditions may be information reflecting the motion state of the target vehicle at different times, and the self-vehicle motion conditions may include, but are not limited to, vehicle driving direction, vehicle speed, vehicle position, etc.

[0075] Embodiment III

[0076] A point cloud generation method provided by an embodiment of the present invention can achieve visual point cloud prediction. This method can be implemented through a historical encoder, a latent space renderer, and a decoder. The historical encoder can condense historical input images into bird's-eye view (BEV) features, the latent space renderer can transform BEV features into 3D geometric features, and the decoder can predict future point clouds according to the 3D geometric features. The embodiments of the present invention implement point cloud prediction and generation through the model ViDAR. See Figure 3 , ViDAR mainly includes three core components: 1. An encoder that can be pre-trained and used to extract BEV features from visual inputs. 2. A latent space renderer that can extract three-dimensional geometric representations from BEV features. 3. A decoder that can predict future BEV features in an autoregressive manner and project future BEV features into a three-dimensional occupancy probability network through a prediction head, and then parse the point cloud through the occupancy probability network. The logic of ViDAR is to determine the waypoint distance of the maximum occupancy response along each ray by emitting rays in different specified directions from the origin, and determine the position of the point through the distance and the ray direction, so as to predict the future point cloud.

[0077] See Figure 4 , the goal of the latent space renderer is to extract more discriminative and representative features of visual inputs. The latent space renderer can accumulate the conditional probabilities of the BEV grids corresponding to visual inputs through a conditional probability function, and then calculate the features of each ray through a feature expectation function. Finally, the ray features are weighted with their associated conditional probabilities to obtain the features of each occupied grid. Specifically, the data representation of the conditional probability function includes:

[0078]

[0079] Among them, i represents different BEV grids, j represents different waypoints on different rays pointing from the origin to BEV grid i, and p represents the independent occupancy probability of different BEV grids, which is obtained by neural network estimation. The embodiment of the present invention can convert the occupancy probability of each independent BEV grid into a joint probability distribution of the previous waypoint probabilities through a conditional probability function.

[0080] In the embodiment of the present invention, after determining the joint probability distribution, the ray feature can be determined through a feature expectation function, and the data representation of the feature expectation function can include:

[0081]

[0082] Among them, represents the conditional probability of grid i in the bird's-eye view feature information and waypoint k on the ray pointing from the origin to grid i, represents the bird's-eye view feature information of the k-th image frame.

[0083] The embodiment of the present invention can obtain the ray feature corresponding to each ray triggered by the origin through the feature expectation function. Then the ray feature and the conditional probability can be weighted to obtain the final occupancy grid feature:

[0084]

[0085] In the embodiment of the present invention, the decoder proposed by ViDAR can decode the future point cloud from the three-dimensional geometric representation. Refer to Figure 5 , the decoder can be an iteratively used transformer, and can continuously predict the BEV feature of the next frame in an autoregressive manner from the result of the previous frame. Among them, in the t-th iteration, it first encodes the desired self-vehicle motion condition of the next frame into a high-dimensional representation through a multi-layer perceptron (MLP). The self-vehicle motion condition can include the future direction, position, speed, etc. of the self-vehicle, and then adds the high-dimensional representation to a series of BEV queries (Future BEV Queries) as input. A deformable self-attention module, a temporal cross-attention module, and a feed-forward network can be used to predict the future point cloud based on the self-vehicle motion condition and the previous frame BEV feature.

[0086] In the embodiments of the present invention, experiments are conducted on the nuScenes dataset to verify the performance of the point cloud generation method. The nuScenes dataset covers complex urban scenarios, including temporal segments of 1000 different autonomous driving scenarios, multi-view images from 6 cameras, point clouds of a 32-line radar, and annotations at 2Hz. The effects of ViDAR implemented by the point cloud generation method provided in the embodiments of the present invention and previous point cloud generation methods on predicting future point clouds are compared on the nuScenes dataset. At the same time, the comparison results of different downstream pre-training are shown in the following table:

[0087] Table 1 Future Point Cloud Prediction

[0088]

[0089]

[0090] The results of future point cloud prediction are shown in Table 1. The evaluation metric is the Chamfer-Distance between the predicted point cloud and the actual point cloud. The smaller this metric is, the better the prediction effect. As shown in Table 1 above, when using visual image input, the method provided in the embodiments of the present invention obtains better point cloud prediction results compared with the existing state-of-the-art methods.

[0091] Table 2 Comparison of Downstream Perception Tasks

[0092]

[0093] The embodiments of the present invention show the results on autonomous driving downstream tasks through Table 2. The evaluation metrics are NDS and mAP. Among them, the metric NDS is the detection score on the nuScenes dataset, and mAP is a commonly used metric to measure the detection correctness. The higher these two metrics are, the better the downstream effect. As shown in Table 2 above, ViDAR implemented by the method provided in the embodiments of the present invention far exceeds the existing pre-training schemes.

[0094] Embodiment 4

[0095] Figure 6 It is a schematic structural diagram of a point cloud generation device according to Embodiment 4 of the present invention. As Figure 6 shown, the device includes: an encoder module 301, a latent space renderer module 302, and a decoder module 303.

[0096] The encoder module 301 is used to obtain the historical temporal image data of the target vehicle and determine the bird's-eye view feature information of the historical temporal image data.

[0097] The latent space renderer module 302 is used to collect three-dimensional geometric representation data of the historical time-series image data within the bird's-eye view feature information.

[0098] The decoder module 303 is used to predict and generate future point clouds according to the ego-motion conditions of the target vehicle and the three-dimensional geometric representation data.

[0099] In the embodiment of the present invention, the encoder module acquires historical time-series image data in the target vehicle and generates bird's-eye view feature information corresponding to the historical time-series image data. The latent space renderer module extracts three-dimensional geometric representation data of the corresponding historical time-series image data according to the bird's-eye view feature information. The decoder module predicts and generates future point clouds according to the three-dimensional geometric representation data under the ego-motion conditions of the target vehicle, which can realize visual image processing in the scenario of the autonomous driving model, can make full use of semantic, geometric features and dynamic time-series information in the image point cloud sequence, can improve the accuracy of point cloud prediction and generation, can improve the reliability of the autonomous driving model, and enhance the user experience.

[0100] In some embodiments of the invention, the encoder module 301 includes:

[0101] The time-series processing unit is used to extract the image frames collected by the target vehicle in chronological order to form the historical time-series image data.

[0102] The bird's-eye view processing unit is used to sequentially input the image frames corresponding to different moments in the historical time-series image data into a preset historical encoder to generate the bird's-eye view feature information of the image frames at different moments, where the preset historical encoder includes a bottom-up bird's-eye view model, a top-down bird's-eye view model, a multi-modal bird's-eye view model, and a decoding bird's-eye view model.

[0103] In some embodiments of the invention, the latent space renderer module 302 includes:

[0104] The bird's-eye view processing unit is used to call a preset latent space renderer to process each frame of the bird's-eye view feature information.

[0105] The geometry extraction unit is used to extract the occupancy grid features generated by the preset latent space renderer corresponding to the bird's-eye view feature information as the three-dimensional geometric representation data.

[0106] In some embodiments of the invention, the bird's-eye view processing unit is specifically used to: accumulate the conditional probabilities of different occupancy grids in the bird's-eye view feature information according to a preset conditional probability function; determine the conditional probabilities and the ray features corresponding to the bird's-eye view feature information according to a preset feature expectation function; and use the weighted product of the ray features and the conditional probabilities as the occupancy grid features.

[0107] In some embodiments of the invention, the preset conditional probability function at least includes the following:

[0108]

[0109] Wherein, the represents the conditional probability of waypoint j on the ray from the origin to grid i within the bird's-eye view feature information of grid i, p represents the independent occupancy probability of grids with different bird's-eye view feature information, and the value of p is determined by neural network estimation;

[0110] The preset feature expectation function at least includes the following:

[0111]

[0112] Wherein, the represents the conditional probability of waypoint k on the ray from the origin to grid i within the bird's-eye view feature information of grid i, represents the bird's-eye view feature information of the k-th image frame.

[0113] In some other embodiments of the invention, the decoder module 303 is specifically configured to: call a preset multi-layer perceptron to determine the self-vehicle motion condition at the next moment based on the image frame at the current moment in the historical time-series image data; perform at least one iteration process on the bird's-eye view query corresponding to the three-dimensional geometric representation data at the current moment, the occupancy grid feature of the three-dimensional geometric representation data at the current moment, and the self-vehicle motion condition through a self-attention module, a temporal cross-attention module, and a forward network to generate the future point cloud.

[0114] Based on the above embodiments of the invention, the self-vehicle motion condition includes at least one of the following: vehicle driving direction, vehicle speed, and vehicle position.

[0115] The point cloud generation device provided by the embodiments of the present invention can execute the point cloud generation method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0116] Embodiment Six

[0117] Figure 7It is a schematic structural diagram of the electronic device implementing the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0118] As Figure 7 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0119] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0120] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the point cloud generation method.

[0121] In some embodiments, the point cloud generation method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the point cloud generation method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the point cloud generation method by any other suitable means (e.g., by means of firmware).

[0122] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0123] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0124] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0125] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0126] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0127] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0128] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0129] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating point clouds, characterized in that, The method includes: Obtaining historical sequential image data of a target vehicle and determining bird's-eye view feature information of the historical sequential image data; Collecting three-dimensional geometric representation data of the historical sequential image data within the bird's-eye view feature information, including: invoking a preset latent space renderer to process each frame of the bird's-eye view feature information; extracting an occupancy grid feature generated by the preset latent space renderer corresponding to the bird's-eye view feature information as the three-dimensional geometric representation data; wherein, invoking the preset latent space renderer to process each frame of the bird's-eye view feature information includes: accumulating the conditional probabilities of different occupancy grids within the bird's-eye view feature information according to a preset conditional probability function; determining the conditional probabilities and ray features corresponding to the bird's-eye view feature information according to a preset feature expectation function; taking the weighted product of the ray features and the conditional probabilities as the occupancy grid feature; Predicting and generating future point clouds according to the ego-vehicle motion conditions of the target vehicle and the three-dimensional geometric representation data, including: invoking a preset multi-layer perceptron to determine the ego-vehicle motion conditions at the next moment based on the image frame at the current moment within the historical sequential image data; sequentially passing the bird's-eye view query corresponding to the three-dimensional geometric representation data at the current moment, the occupancy grid feature of the three-dimensional geometric representation data at the current moment, and the ego-vehicle motion conditions through a self-attention module, a temporal cross-attention module, and a forward network for at least one iteration to generate the future point clouds.

2. The method according to claim 1, wherein The obtaining of the historical sequential image data and the determination of the bird's-eye view feature information of the historical sequential image data include: Extracting image frames collected by the target vehicle in chronological order to form the historical sequential image data; Sequentially inputting the image frames corresponding to different moments within the historical sequential image data into a preset historical encoder to generate the bird's-eye view feature information of the image frames at different moments, wherein the preset historical encoder includes a bottom-up bird's-eye view model, a top-down bird's-eye view model, a multi-modal bird's-eye view model, and a decoding bird's-eye view model.

3. The method according to claim 1, wherein The preset conditional probability function at least includes the following: Among them, the represents the conditional probability of waypoint j on the ray from the origin to grid i within the grid i of the bird's-eye view feature information, p (i,j) represents the independent occupancy probability of the grids of different bird's-eye view feature information, p (i,j) The value of is determined by neural network estimation; The preset feature expectation function at least includes the following: Among them, the represents the conditional probability of the waypoint k on the ray from the origin to the grid i within the bird's-eye view feature information of the grid i, represents the bird's-eye view feature information of the k-th image frame.

4. The method according to claim 1, wherein The ego-vehicle motion conditions include at least one of the following: vehicle driving direction, vehicle speed, and vehicle position.

5. A point cloud generation device, characterized in that, The device includes: An encoder module, configured to obtain historical sequential image data of a target vehicle and determine bird's-eye view feature information of the historical sequential image data; Latent space renderer module, configured to collect three-dimensional geometric representation data of the historical time-series image data within the bird's-eye view feature information, including: invoking a preset latent space renderer to process each frame of the bird's-eye view feature information; extracting the occupancy grid features generated by the preset latent space renderer corresponding to the bird's-eye view feature information as the three-dimensional geometric representation data; wherein, the invoking of the preset latent space renderer to process each frame of the bird's-eye view feature information includes: accumulating the conditional probabilities of different occupancy grids within the bird's-eye view feature information according to a preset conditional probability function; determining the conditional probabilities and the ray features corresponding to the bird's-eye view feature information according to a preset feature expectation function; taking the weighted product of the ray features and the conditional probabilities as the occupancy grid features. Decoder module, configured to predict and generate future point clouds according to the ego-motion conditions of the target vehicle and the three-dimensional geometric representation data, including: invoking a preset multi-layer perceptron to determine the ego-motion conditions at the next moment based on the image frame at the current moment within the historical time-series image data; sequentially passing the bird's-eye view query corresponding to the three-dimensional geometric representation data at the current moment, the occupancy grid features of the three-dimensional geometric representation data at the current moment, and the ego-motion conditions through a self-attention module, a temporal cross-attention module, and a feed-forward network for at least one iteration to generate the future point clouds.

6. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the point cloud generation method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the point cloud generation method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Image processing method, device and equipment and computer readable storage medium

    CN114723955A

  • Perception fusion system, electronic equipment and storage medium

    CN116664997A