Point cloud generation method, device and equipment
By using the verified simulator and the internal and external parameter training model of the binocular camera and lidar, the parallax features are extracted and optimized to generate point clouds, which solves the problems of point cloud sparsity and incompleteness, improves the integrity and density of point cloud generation, and enhances the reliability of environmental perception.
Patent Information
- Application Number
- CN202510865114.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In existing technologies, point cloud generation technology suffers from sparsity, incompleteness, and insufficient adaptability to dynamic scenarios, which affects the reliability and robustness of environmental perception algorithms in fields such as autonomous driving and industrial robots.
The initial model is trained by generating a training data set using a verified simulator. The internal and external parameters of the binocular camera and lidar are combined to extract multi-dimensional parallax features and perform optimization processing to generate depth map data to generate point clouds.
It improves the integrity and density of point clouds, enhances the generalization ability of point cloud generation models, and improves the accuracy and robustness of environmental perception.
Smart Images

Figure CN120374701B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a point cloud generation method, device, and equipment. Background Art
[0002] In environmental perception tasks such as autonomous driving, industrial robotics, and intelligent service robots, high-precision, high-density point clouds are essential for achieving core functions such as obstacle detection and semantic segmentation. Currently, point cloud generation technology primarily relies on direct acquisition from sensors such as LiDAR and depth cameras. However, due to limitations in hardware performance and environmental complexity, the generated raw point clouds often suffer from sparsity, incompleteness, and insufficient adaptability to dynamic scenarios, severely limiting the reliability and robustness of downstream perception algorithms. Summary of the Invention
[0003] The embodiments of the present application provide a point cloud generation method, apparatus, and device that can effectively improve the integrity and density of the point cloud.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides a point cloud generation method, the method comprising:
[0006] A point cloud generation model is used to extract features from the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by collecting data of a calibration scene using a binocular camera and a laser radar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained using the first data; the second data is obtained by simulating data collection of the binocular camera and the laser radar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained using the second data;
[0007] The point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data;
[0008] The point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data;
[0009] Depth map data is generated based on the optimized disparity data, and a point cloud is generated based on the depth map data.
[0010] In a second aspect, an embodiment of the present application provides a point cloud generation device, comprising a first generation unit and a second generation unit;
[0011] The first generating unit is used to extract features of the data to be processed by using the point cloud generation model to obtain feature data to be processed; and to extract multi-dimensional disparity features of the feature data to be processed by using the point cloud generation model to obtain disparity feature data; and to optimize the disparity features of the disparity feature data by using the point cloud generation model to obtain optimized disparity data; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training the initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data and the second internal and external parameters; the first data is obtained by collecting data of the calibration scene by using a binocular camera and a laser radar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained by using the first data; the second data is obtained by simulating the data collection of the binocular camera and the laser radar in the calibration scene by using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained by using the second data;
[0012] The second generating unit is configured to generate depth map data according to the optimized disparity data, and generate a point cloud according to the depth map data.
[0013] In a third aspect, an embodiment of the present application provides a point cloud generation device, comprising a processor and a memory storing processor-executable instructions; when the instructions are executed by the processor, the above-mentioned point cloud generation method is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides an electronic device, including a point cloud generating device, and an electronic device for constructing map data based on the point cloud generated by the point cloud generating device; and performing corresponding operations based on the map data.
[0015] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned point cloud generation method is implemented.
[0016] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the above-mentioned point cloud generation method.
[0017] The embodiment of the present application provides a point cloud generation method, device and equipment, wherein the point cloud generation device can use a point cloud generation model to extract features from the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data and the second internal and external parameters; the first data is obtained by collecting data from a calibration scene using a binocular camera and a laser radar; the first internal and external parameters represent the The internal and external parameters of the binocular camera and the lidar are obtained using the first data; the second data is obtained by simulating data acquisition of the binocular camera and the lidar in a calibration scene using a simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; the point cloud generation model is used to extract multi-dimensional disparity features of the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain optimized disparity data; depth map data is generated based on the optimized disparity data, and a point cloud is generated based on the depth map data.
[0018] It can be seen that the present application can pre-calibrate the simulator to be verified, including the first data obtained by collecting data from a real calibration scene using a binocular camera and a laser radar, the first internal and external parameters of the binocular camera and the laser radar, the second data obtained by simulating the data collection scene of the binocular camera and the laser radar using the simulator to be verified, and the internal and external parameters of the binocular camera and the laser radar calculated using the second data to complete the verification, and whether the simulator to be verified has a good simulation effect, that is, whether the simulator to be verified can be used to obtain data that is basically the same as the real calibration scene, so as to complete the verification based on the calibration results. The simulator generated data used for the test is used to train the initial model to obtain a point cloud generation model, which can improve the generalization ability of the point cloud generation model. Then, when generating a point cloud, the point cloud generation device can use the constructed point cloud generation model to first extract features of the data to be processed, and then extract multi-dimensional disparity features of the feature data to be processed, so as to obtain disparity features of different dimensions. Then, the disparity features of different dimensions will be optimized at different scales to further optimize the disparity features, thereby generating depth map data based on the optimized disparity data, and generating a point cloud based on the depth map data, which can effectively improve the integrity and density of the generated point cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Various other advantages and benefits will become apparent to those skilled in the art by reading the detailed description of the preferred embodiment below.The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present application.
[0020] Figure 1 This is a schematic diagram of the implementation process of the point cloud generation method proposed in the embodiment of the present application;
[0021] Figure 2 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 1 ;
[0022] Figure 3 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 2 ;
[0023] Figure 4 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 3 ;
[0024] Figure 5 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 4 ;
[0025] Figure 6 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 5 ;
[0026] Figure 7 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 6 ;
[0027] Figure 8 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 7 ;
[0028] Figure 9 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 8 ;
[0029] Figure 10 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 9 ;
[0030] Figure 11 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 10 ;
[0031] Figure 12 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 10 one;
[0032] Figure 13 Schematic diagram of the implementation of the point cloud generation method proposed in this application embodiment Figure 10 two;
[0033] Figure 14A schematic diagram of the implementation flow of the simulator verification method proposed in the embodiment of the present application;
[0034] Figure 15 Schematic diagram of the structure of the point cloud generation device proposed in this embodiment Figure 1 ;
[0035] Figure 16 Schematic diagram of the structure of the point cloud generation device proposed in this embodiment Figure 2 . DETAILED DESCRIPTION
[0036] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to explain the related applications and are not intended to limit the applications. It should also be noted that for ease of description, only the portions relevant to the related applications are shown in the drawings.
[0037] With the continuous development of artificial intelligence (AI), robots are evolving from their traditional limited functions of inspection, delivery, and sweeping to general-purpose humanoid robots, primarily embodied in intelligence. This allows robots to play a productive role in a wider range of human scenarios, including retail, industry, and the home. Embodied intelligence gives robots a more powerful "brain," enabling them to perceive, learn, and dynamically interact with their physical environment like humans. This places extremely high demands on robots to accurately perceive their surroundings, requiring AI algorithms to accurately analyze the position and posture of objects surrounding the robot from sensor data.
[0038] Currently, point cloud generation technology mainly relies on direct acquisition by sensors such as LiDAR and depth cameras. However, due to limitations in hardware performance and environmental complexity, the generated raw point clouds generally suffer from sparsity, incompleteness, and insufficient adaptability to dynamic scenes, which seriously restricts the reliability and robustness of downstream perception algorithms.
[0039] Stereo matching technology calculates depth through disparity, and its principle can be expressed as , where Z is the depth, is the baseline distance, is the focal length, is parallax; binocular depth estimation is a commonly used technology in robot perception, which mainly uses binocular images to calculate depth information, thereby reconstructing or restoring the environment in three dimensions. Traditional algorithms often manually select or design image features based on task requirements, and then use the selected or designed image features for depth estimation. These algorithms are more susceptible to the influence of high texture, uniform areas, and target legibility. In recent years, due to the rapid development of artificial intelligence technology, more and more methods have adopted deep neural network methods to solve the problem of binocular depth estimation. However, most methods are often only trained on relatively fixed small datasets. These datasets are usually collected from real scenes or small-scale synthesis. At the same time, although deep neural networks have made great progress, the depth estimation networks in the past are still relatively small. Models trained based on this method often only have certain generalization capabilities in a single scene domain or dataset, and have poor cross-domain generalization capabilities. Once the test set or scene domain changes, the robustness of the algorithm will be greatly reduced, resulting in the robot being unable to accurately perceive the environment in three dimensions. At the same time, although lidar directly provides three-dimensional point cloud information of the scene, the data is often noisy, the ranging accuracy is greatly affected by the environment, and there is a lack of texture information of the target object. Therefore, in the field of three-dimensional perception, the processing of real scene datasets is often very difficult, generally takes a long time, and the resulting datasets are small in size and of average quality, which leads to a bottleneck in further improving the performance of depth estimation methods.
[0040] In order to solve the problems of sparsity, incompleteness and lack of adaptability to dynamic scenes in current point clouds, the embodiments of the present application provide a point cloud generation method, device and equipment; the point cloud generation device can use the point cloud generation model to extract features of the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training the initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data and the second internal and external parameters; the first data is obtained by calibrating the scene using a binocular camera and a lidar The method is obtained by collecting data from the first row; the first internal and external parameters characterize the internal and external parameters of the binocular camera and the laser radar obtained by using the first data; the second data is obtained by simulating the data collection of the binocular camera and the laser radar in the calibration scene using the simulator to be verified; the second internal and external parameters characterize the internal and external parameters of the binocular camera and the laser radar obtained by using the second data; the point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain optimized disparity data; depth map data is generated according to the optimized disparity data, and a point cloud is generated according to the depth map data.
[0041] It can be seen from this that the present application can pre-calibrate the simulator to be verified, including the first data obtained by collecting data from a real calibration scene using a binocular camera and a laser radar, the first internal and external parameters of the binocular camera and the laser radar, the second data obtained by simulating the data collection scene of the binocular camera and the laser radar using the simulator to be verified, and the internal and external parameters of the binocular camera and the laser radar calculated using the second data to complete the verification. It can be used to verify whether the simulator to be verified has a good simulation effect, that is, whether it is possible to use the simulator to be verified to obtain data that is basically the same as the real calibration scene, thereby The initial model is trained based on the verified simulator generated data to obtain a point cloud generation model to improve the generalization ability of the point cloud generation model; then, when generating a point cloud, the point cloud generation device can use the constructed point cloud generation model to first extract features of the processed data, and then extract multi-dimensional disparity features of the processed feature data, so as to obtain disparity features of different dimensions, and then optimize the disparity features of different dimensions at different scales to further optimize the disparity features, thereby generating depth map data according to the optimized disparity data, and generating a point cloud based on the depth map data, which can effectively improve the integrity and density of the generated point cloud.
[0042] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0043] An embodiment of the present application provides a point cloud generation method, which is applied to a point cloud generation device, such as Figure 1 As shown, the point cloud generating method of the point cloud generating device may include the following steps:
[0044] Step 101: Use a point cloud generation model to extract features from the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator.
[0045] In an embodiment of the present application, the point cloud generation device may first obtain data to be processed. The data to be processed may include image data, or image data and radar data; wherein the image data may be obtained by a camera, and the radar data may be obtained by a lidar.
[0046] In some embodiments of the present application, the image data may be obtained by capturing images of the environment in which the point cloud generating device is located through a camera; the radar data may be obtained by scanning the environment in which the point cloud generating device is located through a laser radar.
[0047] In an embodiment of the present application, the point cloud generating device may be a device in an electronic device, used to generate a point cloud; the electronic device may include a binocular camera and a laser radar, and the binocular camera may include a left camera and a right camera. For example, the electronic device may be a robot equipped with a binocular camera and a laser radar, so that the point cloud generating device can generate a point cloud by acquiring image data collected by the binocular camera carried by the robot, or it can generate a point cloud by acquiring image data collected by the binocular camera and radar data collected by the laser radar carried by the robot.
[0048] In an embodiment of the present application, the image data acquired by the binocular camera may include image data of a left camera and image data of a right camera; the type of the binocular camera is not limited in this application, for example, it may be a binocular red, green, and blue (RGB) camera, or a binocular infrared (IR) camera.
[0049] In an embodiment of the present application, the calibrated simulator is obtained by calibrating the simulator to be calibrated based on the first data, the first internal and external parameters, the second data and the second internal and external parameters; the first data is obtained by collecting data of a calibration scene using a binocular camera and a laser radar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar calculated using the first data; the second data is obtained by simulating the data collection of the binocular camera and the laser radar in the calibration scene using the simulator to be calibrated; the second internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar calculated using the second data.
[0050] In some embodiments of the present application, when the simulator to be verified is verified based on the first data, the first internal and external parameters, the second data and the second internal and external parameters, the first data and the second data can be visualized, and then the visualized first data and the visualized second data can be compared to obtain a first comparison result; at the same time, the error calculation of the first internal and external parameters and the second internal and external parameters can be performed to obtain error information between the first internal and external parameters and the second internal and external parameters; the error calculation can also be performed on the coordinate information of the feature points of the feature object in the first data and the second data respectively to obtain the error result of the feature point coordinates; thereby, the verification result of the simulator to be verified is determined based on the first comparison result, the error information and the error result of the feature point coordinates, and the verification result can be used to indicate whether the simulator to be verified has passed the verification.
[0051] In some embodiments of the present application, during verification, the feature objects are various objects in the calibration scene, and the feature points may be key points of these objects; for example, inner corner points of the calibration plate.
[0052] In some embodiments of the present application, if the verification result of the simulator to be verified is failure to pass the verification, the simulator can be adjusted and then the adjusted simulator can be verified to determine whether the adjusted simulator can pass the verification.
[0053] In some embodiments of the present application, if the verification result of the simulator to be verified is that it passes the verification, it can be determined that the simulator is a simulator that passes the verification.
[0054] In some embodiments of the present application, there may be multiple simulators to be verified; that is, different simulators may be verified separately to determine the verification results of each simulator, and then the simulators that pass the verification are selected from the verification results of each simulator for subsequent generation of training data sets, and the simulators that fail the verification are adjusted so that the adjusted simulators pass the verification, or the simulators that fail the verification are directly eliminated.
[0055] In some embodiments of the present application, different types of objects can be set in the calibration scene, and these different types of objects can be used to represent different types of scenes. For example, for some scenes with objects with highly reflective surfaces, highly reflective objects can be placed in the calibration scene, and their surfaces can have highly reflective points and / or highly reflective strips, so as to detect the data synthesis effect of the simulator to be verified on such highly reflective surface object scenes. When the verification result is that the verification is passed, it means that the simulator is suitable for the scene corresponding to the object set in the calibration scene.
[0056] In some embodiments of the present application, a binocular camera and a lidar are used to collect data of the calibration scene, that is, the binocular camera collects image data of the calibration scene, and the lidar collects radar data of the calibration scene; that is, the first data may include image data of the calibration scene collected by the binocular camera, and point cloud of the calibration scene collected by the lidar.
[0057] In some embodiments of the present application, the first intrinsic and extrinsic parameters can be calculated using the calibration relationship and the first data; the first intrinsic and extrinsic parameters include the first intrinsic parameters and the first extrinsic parameters of the binocular camera, and the first intrinsic parameters and the first extrinsic parameters of the lidar.
[0058] In some embodiments of the present application, the second data may include a synthetic image and a synthetic point cloud obtained by simulating the data acquisition of a binocular camera and a lidar in a calibration scene by the simulator to be verified, wherein the synthetic image is an image obtained by simulating and emulating the data acquisition of a binocular camera in a calibration scene by the simulator to be verified, and the synthetic point cloud is a point cloud obtained by simulating and emulating the data acquisition of a lidar in a calibration scene by the simulator to be verified.
[0059] In some embodiments of the present application, second intrinsic and extrinsic parameters can be calculated based on the first data, where the second intrinsic and extrinsic parameters include second intrinsic and second extrinsic parameters of the binocular camera, and second intrinsic and second extrinsic parameters of the lidar.
[0060] In some embodiments of the present application, the first comparison result can be used to indicate the difference between the image synthesized by the simulator to be verified and the image of the real calibration scene captured by the binocular camera, as well as the difference between the point cloud synthesized by the simulator to be verified and the point cloud of the real calibration scene captured by the lidar; the error information can be used to indicate the error between the second internal and external parameters calculated based on the second data and the first internal and external parameters calculated based on the first data.
[0061] In some embodiments of the present application, the first comparison result can be understood as a qualitative assessment, which can be used to determine whether there are problems such as offset and stratification between the image and point cloud corresponding to the real scene through the visualized synthetic image and synthetic point cloud; the error information and the error result of the feature point coordinates can be understood as a quantitative assessment, which determines whether the degree of deviation of these data meets the preset error conditions through the error value between the specific second internal and external parameters and the first internal and external parameters, and the numerical values of the feature point coordinates in the first data and the second data respectively. If the preset error conditions are met, the error is considered to be small, otherwise the error is considered to be large.
[0062] In some embodiments of the present application, when the first comparison result shows that the difference meets the preset difference condition and the error information meets the preset error condition, the simulator to be verified can be determined as a simulator that has passed the verification; wherein, the present application does not specifically limit the preset difference condition and the preset error condition, for example, the preset difference condition may be that there is no offset and stratification, and the preset error condition may include that the error value between the second internal and external parameter and the first internal and external parameter is less than the first error threshold, and the error result of the feature point coordinates is less than the second error threshold.
[0063] In some embodiments of the present application, a preset difference condition can be determined based on the deviation allowed in the actual application field, and a preset error condition can be determined based on the sensor errors of the binocular camera and the lidar. For example, if the actual application field of the point cloud generation method is the navigation equipment field, and the geometric deviation in this field allows an offset of less than 0.1mm, then the preset difference condition can be an offset of less than 0.1mm. If the offset is less than 0.1mm, it is considered to meet the preset difference condition. If the focal length error in the sensor internal and external parameter errors of the binocular camera is less than 0.1%, then the preset error condition can include a focal length error of less than 0.1%. It can be understood that in the embodiments of the present application, the point cloud generation model is used to extract features from the processed data, which can be image feature extraction from image data, or point cloud feature extraction from radar data while extracting image features from image data.
[0064] In some embodiments of the present application, the point cloud generation model may include a monocular depth estimation network, a multi-scale feature extraction network, and a point cloud feature extraction network; wherein the monocular depth estimation network can be used to perform depth estimation on the image data corresponding to the left camera and the right camera respectively; the multi-scale feature extraction network can be used to perform multi-scale feature extraction on the image data corresponding to the left camera and the right camera respectively; the point cloud feature extraction network can be used to extract point cloud features from radar data.
[0065] In an embodiment of the present application, the multi-scale feature extraction network performs multi-scale feature extraction on image data, which means extracting representations of the image data at different resolutions and different levels, so that the image data can be understood from different perspectives; that is, "scale" can be understood as resolution and level.
[0066] In some embodiments of the present application, when a point cloud generation device uses a point cloud generation model to perform feature extraction on the data to be processed and obtains the feature data to be processed, it can use a monocular depth estimation network to extract the monocular depth feature information of the image data, and use a multi-scale feature extraction network to extract the multi-scale depth feature information of the image data when the data to be processed includes image data and radar data; and use a point cloud feature extraction network to extract the point cloud features of the radar data; thereby determining the monocular depth feature information, multi-scale depth feature information and point cloud features as the feature data to be processed.
[0067] In some embodiments of the present application, when a point cloud generation device uses a point cloud generation model to perform feature extraction on the data to be processed to obtain feature data to be processed, it can use a monocular depth estimation network to extract monocular depth feature information of the image data when the data to be processed includes image data, and use a multi-scale feature extraction network to extract multi-scale depth feature information of the image data, and determine the monocular depth feature information and the multi-scale depth feature information as the feature data to be processed.
[0068] It can be understood that in an embodiment of the present application, when extracting image features, it can include using a monocular depth estimation network to extract monocular depth feature information of the left camera image data, and using a multi-scale feature extraction network to extract multi-scale depth feature information of the left camera image data, thereby using the monocular depth feature information of the left camera image data and the multi-scale depth feature information of the left camera image data as the feature data to be processed of the left camera image data; at the same time, it can also include using a monocular depth estimation network to extract monocular depth feature information of the right camera image data, and using a multi-scale feature extraction network to extract multi-scale depth feature information of the right camera image data, thereby using the monocular depth feature information of the right camera image data and the multi-scale depth feature information of the right camera image data as the feature data to be processed of the right camera image data.
[0069] In an embodiment of the present application, since the feature data to be processed may include depth features corresponding to the image data of the binocular camera and point cloud features corresponding to the radar data of the lidar, the above-mentioned feature extraction process realizes a cross-modal feature extraction, which can complement the advantages of visually rich color and texture information and accurate ranging capabilities of point clouds, and make up for the weaknesses of single-modality sensors, such as the lack of ranging capabilities when only images are obtained through binocular cameras, and the lack of target color and texture information when only point clouds are obtained through lidar.
[0070] In an embodiment of the present application, the feature data to be processed may be a hybrid cost volume feature obtained by fusing the depth feature information of the image data, i.e., monocular depth feature information and multi-scale depth feature information, and point cloud features; wherein the cost volume may be understood as a three-dimensional data structure for storing pixel matching costs under different disparity assumptions, where the pixel matching cost refers to a quantified similarity value between pixels in the left camera image data and pixels in the right camera image data, and the three-dimensional dimensions of the three-dimensional data structure include a disparity assumption dimension, a feature map height dimension, and a feature map width dimension.
[0071] In some embodiments of the present application, when the depth feature information and point cloud features of the image data are fused to obtain the final feature data to be processed, the point cloud features can be projected onto a two-dimensional image coordinate system for fusion with the image features; the advantages of lidar include active detection and accurate ranging, but lack of semantic information, sparse point clouds in the distance, and the quality of point clouds is greatly affected by the environment; binocular cameras perform poorly in weak texture areas and dark light conditions. By fusing the features of these two modalities, the stereo matching effect can be improved to generate a more complete and denser point cloud.
[0072] For example, the principle of projecting point cloud features into a two-dimensional image coordinate system can be expressed as the following formula:
[0073] (1)
[0074] in, , , is the world coordinate of the point cloud, , Represents the image feature coordinates obtained after transformation, and Represents the rotation matrix and displacement of the point cloud from the world coordinate system to the camera coordinate system, is the camera internal parameter, is the depth of the point cloud in the camera coordinate system, is the scaling factor between the converted image features and the original image features.
[0075] Step 102: Use the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data.
[0076] In an embodiment of the present application, after the point cloud generation device uses the point cloud generation model to extract features from the data to be processed and obtains the feature data to be processed, it can use the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed and obtain disparity feature data.
[0077] In some embodiments of the present application, the point cloud generation model may include a global feature extractor and a hybrid convolutional feature extractor.
[0078] In an embodiment of the present application, the global feature extractor can be a feature extractor based on Transformer that includes an attention mechanism; the global feature extractor can include a self-attention layer, a first sum normalization layer, a feedforward neural network, and a second sum normalization layer.
[0079] In some embodiments of the present application, a point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed. When the disparity feature data is obtained, a global feature extractor can be used to perform convolution downsampling processing on the feature data to be processed to obtain the downsampled disparity features, and global feature extraction processing can be performed based on the downsampled disparity features and position coding information corresponding to the data to be processed to obtain a first disparity feature; a hybrid convolution feature extractor is used to perform multi-scale local disparity feature extraction processing on the feature data to be processed to obtain multi-scale local disparity feature information, and hybrid convolution feature extraction processing is performed based on the multi-scale local disparity feature information to obtain a second disparity feature; and the disparity feature data is determined based on the first disparity feature and the second disparity feature.
[0080] In some embodiments of the present application, disparity feature data is determined based on the first disparity feature and the second disparity feature, which can be understood as taking the first disparity feature and the second disparity feature as disparity feature data; in the actual operation process of determining the disparity feature data, the first disparity feature and the second disparity feature can be spliced to obtain the disparity feature data.
[0081] In the embodiments of the present application, the method for obtaining the position coding information corresponding to the data to be processed is not limited in this application. For example, the position coding information can be determined based on the position information of the binocular camera, or based on the radar data of the lidar.
[0082] In an embodiment of the present application, the global feature extraction process may include a process of cyclically performing multiple feature extractions on the input first disparity feature and position encoding information using a self-attention layer, a first sum normalization layer, a feedforward neural network, and a second sum normalization layer, and a process of three-dimensional convolution upsampling; for example, Figure 2 As shown, after the feature data to be processed is down-sampled by three-dimensional convolution 21, the down-sampled disparity feature 22 can be obtained. The down-sampled disparity feature can be understood as a coarse-grained disparity feature, which can significantly reduce computing power; then the down-sampled disparity feature 22 and the three-dimensional position encoding information 23 corresponding to the data to be processed are input into the self-attention layer 24, the first sum normalization layer 25, the feedforward neural network 26 and the second sum normalization layer 27 in sequence, and the processing of these four network layers is cyclically executed N times, and then after up-sampling 28 of three-dimensional convolution, the first disparity feature 29 is output; where N≥2.
[0083] In the embodiment of the present application, the specific value of N in the loop executed N times is not limited in this application. In actual application, the value of N can be determined by comprehensively considering the requirements for computing efficiency and the model accuracy of the global feature extractor.
[0084] In an embodiment of the present application, a global feature extractor can utilize a global self-attention mechanism to perform global feature fusion, which has the advantage of being able to extract global context modeling information. Since the computational complexity of global information is often large, three-dimensional convolution downsampling is used to obtain coarse-grained disparity features, and then the coarse-grained disparity features are processed, which can significantly reduce computing power requirements. At the same time, the hybrid convolution feature extractor focuses more on the perception of local information. The feature extraction based on the global feature extractor and the hybrid convolution feature extractor can be greatly enriched to improve the point cloud generation effect.
[0085] In some embodiments of the present application, when hybrid convolution feature extraction processing is performed based on multi-scale local disparity feature information to obtain the second disparity feature, three-dimensional convolution processing, spatial convolution processing, and disparity convolution processing can be performed based on the local disparity feature information at the target scale in the multi-scale local disparity feature information to obtain third feature information; then, the multi-scale image features of the image data are fused using attention weights to obtain fourth feature information; thereby, the second disparity feature is determined based on the third feature information and the fourth feature information.
[0086] In an embodiment of the present application, the hybrid convolutional feature extractor can be constructed based on a convolutional neural network, for example, the hybrid convolutional feature extractor can be constructed based on the Hourglass network structure or the U-Net network structure.
[0087] In the embodiment of the present application, the target scale may be any one of multiple scales.
[0088] In an embodiment of the present application, when determining the second disparity feature based on the third feature information and the fourth feature information, the third feature information and the fourth feature information can be fused to obtain a fusion result, and a hybrid convolution feature extractor can be used to continue to perform subsequent scale feature extraction on the fusion result, including a downsampling or upsampling process, to finally obtain the second disparity feature.
[0089] For example, Figure 3 As shown, the hybrid convolution feature extractor can extract multi-scale disparity features from the feature data to be processed 33 by downsampling 31 and upsampling 32, that is, obtain multi-scale local disparity features to obtain second disparity features 34; wherein, in the process of extracting the multi-scale disparity features, the features of a certain scale can be extracted as follows: Figure 4 The processing shown in FIG. 3 includes sequentially performing three-dimensional convolution processing 36, spatial convolution processing 37, and disparity convolution processing 38 on the local disparity feature information 35 at the target scale to obtain third feature information 39, and using multi-scale attention weights 310 to process the multi-scale image features 311 to obtain fourth feature information 312, and then fusing the third feature information and the fourth feature information 313. It can be understood that when performing the following processing on the features of a certain scale, Figure 4 After the processing shown, the hybrid convolutional feature extractor can be used to continue to perform subsequent downsampling and / or upsampling processes on the fusion result to obtain a second disparity feature; through the hybrid convolutional feature extractor, the computational efficiency can be significantly improved while effectively extracting the disparity feature.
[0090] In the embodiments of the present application, the multi-scale image features may be obtained by extracting multi-scale image features from image data. The method for extracting the multi-scale image features is not limited in this application.
[0091] In an embodiment of the present application, three-dimensional convolution processing is a process of performing convolution simultaneously in spatial and temporal dimensions, where space includes length and width; spatial convolution processing represents the process of convolution of disparity features in spatial dimension; and parallax convolution processing represents the process of convolution of disparity features in parallax dimension.
[0092] For example, the three-dimensional convolution is expressed as K_s×K_s×K_d, and the three-dimensional convolution can be decomposed into spatial convolution K_s×K_s×1 and parallax convolution 1×1×K_d, so as to realize spatial convolution processing by spatial convolution and realize parallax convolution processing by parallax convolution. The execution order of spatial convolution processing and parallax convolution processing is not limited in this application; for example, the spatial convolution processing is performed first and the parallax convolution processing is performed later. Figure 5 As shown, the disparity feature 41 can be subjected to spatial convolution processing 42 and then to parallax convolution processing 43 to obtain a disparity feature 44 that has undergone spatial convolution and parallax convolution. This can realize the extraction of disparity features in different dimensions, greatly reduce the amount of calculation parameters and computing power, and thus improve computing efficiency.
[0093] Step 103: Optimize the multi-scale disparity features of the disparity feature data using the point cloud generation model to obtain optimized disparity data.
[0094] In an embodiment of the present application, the point cloud generation device uses the point cloud generation model to extract multi-dimensional disparity features from the feature data to be processed. After obtaining the disparity feature data, the point cloud generation device can use the point cloud generation model to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data.
[0095] In some embodiments of the present application, the point cloud generation model may include an optimization iterator for performing optimization processing of multi-scale disparity features on disparity feature data.
[0096] In an embodiment of the present application, the optimization iterator can be a model constructed based on a convolutional gated recurrent unit (ConvGRU) and includes a multi-level hidden state update and attention selection mechanism. Based on the optimization iterator, the disparity features can be further refined and optimized.
[0097] In some embodiments of the present application, when a point cloud generation device uses a point cloud generation model to optimize multi-scale disparity features of disparity feature data to obtain optimized disparity data, it can use an optimization iterator to fuse the image features and disparity feature data of the image data in the processed data to obtain fused features; then determine the disparity correction value based on the first feature information at different scales in the fused features; and then optimize the initial disparity value based on the disparity correction value to obtain optimized disparity data.
[0098] In an embodiment of the present application, the fused features may be presented in the form of a correlation coefficient feature pyramid; the fused features may include first feature information at multiple scales.
[0099] In an embodiment of the present application, the process of optimizing the disparity features by the optimization iterator can be understood as determining a multi-scale correlation coefficient feature pyramid after fusing the image features and the disparity feature data, and then calculating the disparity correction value for the first feature information at each scale in the multi-scale correlation coefficient feature pyramid, thereby obtaining the optimized disparity data based on the disparity correction value and the disparity initial value, and completing the optimization of the disparity features.
[0100] In some embodiments of the present application, when the initial disparity value is optimized according to the disparity correction value to obtain optimized disparity data, the disparity correction value corresponding to each scale can be superimposed with the initial disparity value and then passed to the next scale for further correction until the disparity correction value corresponding to the last scale is obtained as the disparity correction result, which is superimposed with the initial disparity value to obtain optimized disparity data.
[0101] In an embodiment of the present application, the initial disparity value may be calculated based on disparity feature data. The method for obtaining the initial disparity value is not limited in this application. For example, the value may be obtained by performing feature extraction, cost calculation, and cost aggregation on the disparity feature data.
[0102] For example, Figure 6 As shown, by fusing image features, including image features 51 of the left camera, image features 52 of the right camera, and disparity feature data 53, a fused feature 54 is obtained in the form of a correlation coefficient feature pyramid. Then, the disparity correction value at the current scale is calculated based on the first feature information at each scale in the fused feature. The disparity correction value at the next scale is calculated based on the result of superimposing the disparity correction value at the current scale with the initial disparity value 55 and the first feature information of the next scale, and so on, until the disparity correction value 56 of the last scale is obtained. The disparity correction value 56 of the last scale is superimposed with the initial disparity value 55 to obtain optimized disparity data 57.
[0103] In some embodiments of the present application, when the point cloud generation device determines the disparity correction value based on the first feature information at different scales in the fused features, it can perform feature downsampling of multiple resolutions on the first feature information to obtain multiple second feature information; then fuse the multiple second feature information to obtain fused feature information; and then determine the disparity correction value based on the fused feature information.
[0104] In some embodiments of the present application, the first feature information may be downsampled by 32 times, 16 times, and 8 times, to obtain resolutions of the original resolution of the first feature information. 、 as well as That is, the multiple resolutions may include the original resolution of the first feature information. 、 as well as ; Then, the second feature information obtained after the three downsampling is fused to obtain fused feature information.
[0105] For example, if feature updates are only performed at a fixed resolution, the image receptive field will be relatively small even if the updates are repeated multiple times, which will have a negative impact on targets with weak textures and large scales. Figure 7 As shown, when performing the calculation of the disparity correction value at a certain scale, the embodiment of the present application can first perform downsampling 62 of 32 times the resolution, downsampling 63 of 16 times the resolution, and downsampling 64 of 8 times the resolution on the first feature information 61 at the scale, and fuse the second feature information obtained after the above-mentioned downsampling at different magnifications to obtain fused feature information 65; thereby, the disparity correction value at the scale can be calculated based on the fused feature information at the scale.
[0106] That is to say, in the embodiments of the present application, the disparity features can be refined by optimizing the iterator, and local details can be gradually corrected through multi-stage correction, and finally a precisely optimized disparity can be obtained in the last layer, taking into account both computational efficiency and precision optimization.
[0107] Step 104: Generate depth map data based on the optimized disparity data, and generate a point cloud based on the depth map data.
[0108] In an embodiment of the present application, the point cloud generation device uses a point cloud generation model to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data. Then, the device can generate depth map data based on the optimized disparity data and generate a point cloud based on the depth map data.
[0109] In an embodiment of the present application, the point cloud generated in the above manner has good completeness and density, accurately represents the spatial environment, and has richer representation of details. Therefore, when subsequent environmental perception tasks are performed based on such a point cloud, the reliability and robustness of environmental perception can be greatly improved.
[0110] In some embodiments of the present application, after obtaining the point cloud, map data can be constructed based on the point cloud, and then related operations such as road planning, positioning or navigation can be implemented based on the map data.
[0111] In some embodiments of the present application, before the point cloud generation device uses the point cloud generation model to extract features from the data to be processed to obtain the feature data to be processed, the following steps may be further included:
[0112] Step 105: Acquire first scene data determined through three-dimensional modeling.
[0113] In an embodiment of the present application, the point cloud generating device can obtain first scene data determined by three-dimensional modeling.
[0114] In an embodiment of the present application, the first scene data is scene data generated by three-dimensional modeling; for example, three-dimensional scene modeling can be implemented by three-dimensional modeling software to obtain the first scene data.
[0115] Step 106: Acquire second scene data; wherein the second scene data is scene data obtained by reconstructing a simulated scene based on the first image information.
[0116] In an embodiment of the present application, the point cloud generating device can obtain second scene data; wherein the second scene data is scene data obtained by reconstructing a simulated scene based on the first image information.
[0117] In the embodiments of the present application, the specific method of reconstructing the simulation scene is not limited in this application. For example, the simulation scene reconstruction can be achieved through any method such as Multi-View Stereo Reconstruction (MVS) reconstruction, Neural Radiance Fields (NeRF), 3D Gaussian Splatting (3DGS) or 4D Gaussian reconstruction, Simultaneous Localization and Mapping (SLAM) combined with visual reconstruction and data twins. These methods can quickly convert real-world scenes into three-dimensional scene data, thereby realizing the construction of low-cost, large-scale simulation scene data.
[0118] In the embodiments of the present application, through three-dimensional modeling and simulation scene reconstruction, scene data covering a variety of scenarios such as retail, industry, home, and autonomous driving can be generated, including indoor and outdoor, dynamic and static, weak texture, transparent objects and highly reflective surface objects, etc., and then when sensor data is subsequently synthesized based on these scene data, tens of millions of high-fidelity stereo image data and point clouds covering a variety of scenes can be generated.
[0119] Step 107: Perform sensor data synthesis on the first scene data and the second scene data based on the verified simulator and pre-configured parameters to obtain synthesized image data and a synthesized point cloud; wherein the pre-configured parameters include simulation configuration parameters of the camera and radar, and environmental parameters.
[0120] In an embodiment of the present application, the point cloud generation device obtains first scene data determined by three-dimensional modeling, and obtains second scene data; wherein the second scene data is scene data obtained by reconstructing a simulated scene based on the first image information, and then the first scene data and the second scene data can be synthesized by sensor data based on a verified simulator and pre-configured parameters to obtain synthesized image data and a synthesized point cloud; wherein the pre-configured parameters include simulation configuration parameters of the camera and radar, and environmental parameters.
[0121] In an embodiment of the present application, a verified simulator may be used to parse and present scene data.
[0122] In an embodiment of the present application, the simulation configuration parameters of the camera may include the camera intrinsic parameters, baseline, focal length and other parameters of the binocular camera; the simulation configuration parameters of the radar may include the number of beams, angular resolution, point cloud resolution and ranging accuracy; the environmental parameters may include lighting, material, object layout and dynamic physical simulation parameters.
[0123] In the embodiments of the present application, sensor data synthesis can be understood as a synthesis of imaging data, that is, synthesizing images and point clouds at different perspectives in a scene to obtain synthesized image data and synthesized point clouds.
[0124] Step 108: Determine a training data set based on the synthesized image data and the synthesized point cloud.
[0125] In an embodiment of the present application, after the point cloud generation device performs sensor data synthesis on the first scene data and the second scene data based on a verified simulator and pre-configured parameters to obtain synthesized image data and a synthesized point cloud, it can determine a training data set based on the synthesized image data and the synthesized point cloud.
[0126] In some embodiments of the present application, when the point cloud generation device determines a training data set based on the synthesized image data and the synthesized point cloud, it can use a data quality assessment model to classify the synthesized image data and the synthesized point cloud to obtain classified image data and classified point cloud; and determine the training data set based on the image data with qualified classification results in the classified image data and the point cloud with qualified classification results in the classified point cloud.
[0127] In an embodiment of the present application, in order to ensure that the data in the training data set meets the requirements of generating a dense point cloud, a data quality assessment model can be used to classify the synthesized image data and the synthesized point cloud, so as to determine the training data set based on the classification results of qualified image data and qualified point cloud; that is, the data quality assessment model can be a classification model.
[0128] In some embodiments of the present application, in addition to determining image data with qualified classification results and point clouds with qualified classification results as training data sets, image data and point clouds that have been manually verified can also be obtained. Manually verified image data and point clouds refer to image data and point clouds obtained by manually verifying the classified image data and classified point clouds again. In this way, fuzzy or unreliable samples can be identified and eliminated, thereby determining this part of qualified data that has been verified twice as a training data set, which can further improve the quality of the data set.
[0129] In some embodiments of the present application, the data quality assessment model can also be updated using data classified as qualified and data classified as unqualified to obtain an updated data quality assessment model that can more accurately classify qualified and unqualified data.
[0130] For example, Figure 8 As shown, after obtaining the synthesized data 71, wherein the synthesized data includes synthesized image data and synthesized point cloud, the synthesized data can be classified 72 using a data quality assessment model to obtain qualified data 73 and unqualified data 74. For example, qualified data can be data with the image centered and the object displayed completely, and unqualified data can be data with a display angle that is too deviated or unclear, etc.; the classified data can then be verified, for example, by manual verification, to obtain verified qualified data 75, so as to determine a training data set 76 based on the verified qualified data; in addition, both qualified data and unqualified data in the classified data can be used to iteratively update the data quality assessment model.
[0131] Step 109: Use the training data set to train the initial model to obtain a point cloud generation model.
[0132] In an embodiment of the present application, after determining a training data set based on the synthesized image data and the synthesized point cloud, the point cloud generation device can use the training data set to train the initial model to obtain a point cloud generation model.
[0133] In an embodiment of the present application, the initial model represents an untrained model for point cloud generation, and the initial model may include an untrained initial global feature extractor, an initial hybrid convolutional feature extractor, an initial optimization iterator, an initial monocular depth estimation network, an initial multi-scale feature extraction network, and an initial point cloud feature extraction network.
[0134] In an embodiment of the present application, multi-source sensor data is synthesized based on rich-source synthetic scenes, and the sources of synthetic scenes include pure simulation environment construction achieved through three-dimensional modeling, and scene migration from real to simulation (Real2Sim) achieved through simulation scene reconstruction; different scene construction methods through the above-mentioned methods can greatly increase data richness; multi-source sensors can support the synthesis of binocular RGB cameras and infrared cameras, as well as lidar point clouds, thereby generating image data and radar data with varying resolutions and field of view based on rich sensor configurations, greatly improving data richness.
[0135] In an embodiment of the present application, by training the initial model with the above-mentioned massive simulated synthetic data, it is possible to effectively solve the problem of degraded cross-domain performance of traditional stereo matching methods and multi-source sensor data fusion in real complex scenes, so that the final point cloud generation model can maintain strong generalization capabilities in rich real-world scenes and different sensor combinations or configurations.
[0136] An embodiment of the present application provides a point cloud generation method, in which a point cloud generation device can use a point cloud generation model to extract features of data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by using a binocular camera and a laser radar to collect data of a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained using the first data; the second data is obtained by simulating data collection of the binocular camera and the laser radar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained using the second data; the point cloud generation model is used to extract multi-dimensional disparity features of the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize multi-scale disparity features of the disparity feature data to obtain optimized disparity data; depth map data is generated according to the optimized disparity data, and a point cloud is generated according to the depth map data.
[0137] Based on the above embodiment, in another embodiment of the present application, illustratively, as Figure 9 As shown, after the binocular camera images collected from the real environment, including the left camera image 81 and the right camera image 82, are input into the point cloud generation model, the monocular depth estimation network 83 and the multi-scale feature extraction network 84 in the point cloud generation model can be used to extract depth features of the binocular camera images, and then the extracted depth features, including the left camera depth features 85 and the right camera depth features 86, are fused to obtain a mixed cost volume feature 87, and then the mixed cost volume feature 87 is input into the global feature extractor 88 and the mixed convolution feature extractor 89 to extract multi-dimensional disparity features, and then the obtained disparity feature data 810 is input into the optimization iterator 811 to optimize the disparity features to obtain optimized disparity data 812, so that depth map data can be generated according to the optimized disparity data, and then a point cloud can be generated according to the depth map data.
[0138] For example, Figure 10As shown, after obtaining binocular camera images from a real environment, including images 91 of a left camera and images 92 of a right camera, and radar data 93 obtained by a lidar, the above data can be input into a point cloud generation model, and a monocular depth estimation network 94 and a multi-scale feature extraction network 95 in the point cloud generation model are used to extract depth features from the images. At the same time, a point cloud feature extraction network 96 in the point cloud generation model is used to extract point cloud features from the radar data. The extracted depth features, including the depth features 97 of the left camera and the depth features 98 of the right camera, and the point cloud features 99 are fused to obtain a mixed cost volume feature 910. The mixed cost volume feature is then input into a global feature extractor 911 and a mixed convolutional feature extractor 912 for multi-dimensional disparity feature extraction. The obtained disparity feature data 913 is then input into an optimization iterator 914 to optimize the disparity feature to obtain an optimized disparity feature 915. Thus, depth map data can be generated based on the optimized disparity data, and then a point cloud can be generated based on the depth map data.
[0139] In some embodiments of the present application, Figure 11 As shown, the first scene data 111 can be obtained through three-dimensional modeling, and the simulation scene can be reconstructed through Real2Sim to obtain the second scene data 112, wherein the specific method of implementing Real2Sim is not limited in this application, such as MVS reconstruction, NeRF, three-dimensional Gaussian splashing, etc., and then the first scene data and the second scene data are imported into the simulator 113, and the binocular camera, lidar and environmental parameters are configured 114, and synthesized data 115 is obtained based on random poses. The synthesized data may include synthesized image data and synthesized point clouds.
[0140] For example, Figure 12 As shown, after the binocular camera images collected from the real environment, including the left camera image 121 and the right camera image 122, are input into the point cloud generation model, the monocular depth estimation network and the multi-scale feature extraction network in the point cloud generation model can be used to extract depth features 123 from the binocular camera images, and then the extracted depth features, including the depth features of the left camera and the depth features of the right camera, are fused to obtain a mixed cost volume feature 124, and then multi-dimensional disparity features are extracted from the mixed cost volume feature 125, including inputting the mixed cost volume feature into a global feature extractor and a mixed convolution feature extractor to obtain multi-dimensional disparity features, and then multi-scale disparity feature optimization processing 126 is performed on the disparity feature, including inputting the obtained disparity feature into an optimization iterator to optimize the multi-dimensional disparity feature to obtain optimized disparity data, so that depth map data 127 can be generated according to the optimized disparity data, and then a point cloud 128 is generated according to the depth map data 127.
[0141] For example, Figure 13 As shown, after collecting binocular camera images from the real environment, including the left camera image 131 and the right camera image 132, and the radar data 133 obtained by the lidar, the above data can be input into the point cloud generation model, and the monocular depth estimation network and the multi-scale feature extraction network in the point cloud generation model are used to extract depth features 134 from the image. At the same time, the point cloud feature extraction network in the point cloud generation model is used to extract point cloud features 135 from the radar data. Then, the extracted depth features, including the depth features of the left camera and the depth features of the right camera, as well as the point cloud features, are combined into a point cloud feature extraction network. After fusion, a mixed cost volume feature 136 is obtained, and then multi-dimensional disparity feature extraction 137 is performed on the mixed cost volume feature, including inputting the mixed cost volume feature into a global feature extractor and a mixed convolution feature extractor to obtain multi-dimensional disparity feature, and then multi-scale disparity feature optimization processing 138 is performed on the disparity feature, including inputting the disparity feature into an optimization iterator to optimize the disparity feature to obtain the optimized disparity feature; thereby, depth map data 139 can be generated based on the optimized disparity data, and then point cloud 1310 and point cloud 1311 can be generated based on the depth map data.
[0142] In some embodiments of the present application, when it is necessary to enhance the generalization ability of the point cloud generation model in a specific type of scene, a small amount of real data collection can be used to fine-tune the model. During fine-tuning, images and laser point clouds are used as the original input of the model. Synthetic data is extracted from high-precision three-dimensional scene reconstruction data based on the real-time posture of the sensor in space to obtain the true value of model training. Only samples with high data accuracy and general point cloud thickening effect are added to the model fine-tuning. In this way, the model performance will be further improved in specific scenarios; that is, the point cloud generation model trained based on massive synthetic data in the embodiment of the present application has high generalization in real scene data. When it is necessary to further improve the point cloud thickening performance in a certain type of scene, a small amount of real machine data fine-tuning can be used to further improve the performance of the point cloud generation model.
[0143] In an embodiment of the present application, the point cloud generation model includes a monocular depth estimation network, a multi-scale feature extraction network, a point cloud feature extraction network, a global feature extractor, a hybrid convolutional feature extractor, and an optimization iterator. By proposing a point cloud generation model with the above structure, the present application can maintain good generalization capabilities in any real complex scene and generate a point cloud with good integrity and density.
[0144] In some embodiments of the present application, the global feature extractor can be a feature extractor based on Transformer that includes an attention mechanism; the global feature extractor can include a self-attention layer, a first sum normalization layer, a feedforward neural network, and a second sum normalization layer; the global feature extractor can perform global feature fusion, which has the advantage of being able to extract global context modeling information, but since the amount of global information calculation of Transformer is often large, the present application adopts a method of processing the coarse-grained disparity features obtained after downsampling, which can significantly reduce computing power; however, this part focuses on global feature extraction and has a relatively weak perception of local information. Therefore, the overall architecture needs to be combined with a hybrid convolutional feature extractor and an optimized iterator to further enrich feature extraction.
[0145] In some embodiments of the present application, a hybrid convolutional feature extractor can automatically extract features, has the advantages of weight parameter sharing and local connection, preserves spatial structural relationships while extracting features, and can efficiently process large-scale data. Because disparity features have an additional dimension of disparity compared to general image features, the hybrid convolutional feature extractor of the present application fuses multi-scale features through downsampling and then upsampling. When extracting features at a certain scale, it uses a partial three-dimensional convolution plus a two-dimensional spatial convolution plus a one-dimensional disparity convolution to extract features. Feature extraction is also performed through an attention mechanism guided by multi-scale image features. Compared to the current related stereo matching network that uses three-dimensional convolution for feature extraction, it can effectively extract disparity features while significantly improving computational efficiency.
[0146] In some embodiments of the present application, an optimization iterator can be used to optimize the disparity features. The process of optimizing the disparity features by the optimization iterator can be understood as determining a multi-scale correlation coefficient feature pyramid after fusing the image features and the disparity feature data, and then calculating the disparity correction value for the first feature information at each scale in the multi-scale correlation coefficient feature pyramid, thereby obtaining the optimized disparity data based on the disparity correction value and the disparity initial value, and completing the optimization of the disparity features; in addition, the optimization iterator has the characteristics of multi-level hidden state updates, and can simultaneously perform feature updates on 8x, 16x, and 32x down-sampling features, so that the optimization iterator can produce good responses to areas of different sizes and can adapt to large-scale, highly random, and diverse synthetic data.
[0147] In some embodiments of the present application, robot scene interaction and refined perception are important prerequisites for the robot to conduct physical interaction and good control. The robot's accurate three-dimensional structural perception of the target scene is achieved through a point cloud generation model generalized based on massive synthetic data; multimodal data fusion technology can use the different sensor advantages of cameras and lidars to identify the accurate position and posture information of the target object, thereby laying a solid perception foundation for the robot to perform a series of actions such as obstacle avoidance, grasping, placement, and assembly. By enhancing environmental perception capabilities and operation planning accuracy, the point cloud generation model can become a key support for robot control to move from the laboratory to practical applications. In the field of autonomous driving, the point cloud generation model can also improve the perception capabilities of autonomous driving vehicles. By generalizing the massive simulation data training model in real scenes, it can help the autonomous driving field improve the use of stereo matching, overcome the bottleneck problem of data scarcity, and improve three-dimensional scene perception capabilities.
[0148] In some embodiments of the present application, the effect of point cloud thickening depends on a point cloud generation model trained based on large-scale synthetic data, and the synthetic data depends on the generation of the simulator. Therefore, the degree of alignment between the simulator synthetic data and the real machine collected data will affect the final point cloud thickening effect; therefore, the present application proposes a method for calibrating the simulator to improve the generation effect of synthetic data, thereby ensuring the point cloud thickening effect; for example, Figure 14 As shown, a calibration scene can be built in a real environment (step 1301), and then the binocular camera and the laser radar are used to collect first data in the calibration scene (step 1302). Then, the first internal and external parameters of the binocular camera and the laser radar are calculated using the first data (step 1303). The calibration scene can also be simulated based on the first internal and external parameters and the simulator to be verified to synthesize the image and the point cloud to obtain second data (step 1304). Then, the second internal and external parameters of the binocular camera and the laser radar are calculated based on the second data (step 1305). Then, qualitative and quantitative evaluations are performed based on the first data, the second data, the first internal and external parameters, and the second internal and external parameters to complete the verification of the simulator to be verified (step 1306).
[0149] For example, in the first step, in order to construct the same scene for the real machine and the simulation, a calibration room, i.e., a calibration scene, can be built in the real world first; the calibration room can be equipped with multiple calibration plates, QR codes, radar high reflectors, such as high reflective points or high reflective strips, and some objects of fixed shapes, such as cubes, cuboids, etc. The positions and postures of different targets in the real machine environment are calculated by calibration methods, and then targets of the same size are constructed in the simulator, and the targets are placed in the same posture as in the real environment; in the second step, real machine data can be collected and the internal and external parameters of the sensor can be calculated, including collecting data of different objects in a real calibration room through sensors such as binocular cameras and lidars, i.e., first data, and then calculating the internal parameters of the sensor and the external parameters of the camera and lidar in different frame data, i.e., first internal and external parameters, through the calibration relationship; in the third step, the sensor data is synthesized through the simulator, including simulating the data of different objects collected by binocular cameras and lidars in the real calibration room in the simulator to render and synthesize different frames. The image and laser point cloud are used to obtain the second data, and the internal and external parameters of the binocular camera and lidar simulated in the simulator are calculated again based on the second data, that is, the second internal and external parameters; the fourth step is to evaluate the degree of alignment between the simulation rendering result and the real machine acquisition data, that is, to use the first data, the second data, the first internal and external parameters and the second internal and external parameters to perform qualitative and quantitative evaluations. For example, in the qualitative evaluation, the visualized first and second data can be combined to evaluate whether the data are consistent, for example, it can be determined whether there is offset, stratification, etc.; the quantitative evaluation can evaluate the coordinate errors of the feature points of the feature objects in the first and second data, such as the coordinate errors of the corner points in the calibration plate and the key points of other feature objects. It can also be evaluated whether the second internal and external parameters calculated by the simulation are consistent with the first internal and external parameters of the real sensor, or whether the error size is within a reasonable range; thereby, the results of the qualitative and quantitative evaluations can be combined to determine whether the simulation data and the real machine data can be effectively aligned, that is, to determine whether the simulator can effectively simulate the real data.
[0150] In an embodiment of the present application, the average distance alignment error of the verified simulator is only 0.2 mm, and the average angle alignment error is only 0.01 degrees; thus, the verified simulator is used to generate massive scene data, and used to train the initial model to obtain a point cloud generation model, which can effectively improve the generalization ability of the point cloud generation model, so that a point cloud with good integrity and density can be generated in any scenario.
[0151] Exemplarily, the method of the embodiments of the present application can be applied to a variety of practical scenarios, such as robot-scene interaction. Refined perception is an important prerequisite for the robot to conduct physical interaction and good control. The robot's accurate three-dimensional structural perception of the target scene is achieved through point cloud thickening technology based on generalization of massive synthetic data. Multimodal data fusion technology can utilize the different sensor advantages of cameras and lidars to identify the accurate position and posture information of the target object, thereby laying a solid perception foundation for the robot to perform a series of actions such as obstacle avoidance, grasping, placement, and assembly. Point cloud thickening technology can become a key support for robot control to move from the laboratory to practical application by enhancing environmental perception capabilities and operation planning accuracy; it can also be applied to the field of autonomous driving. Multimodal point cloud thickening technology can also improve the perception capabilities of autonomous driving vehicles. By achieving the generalization of massive simulation data training models in real scenes, it can help the autonomous driving field improve the use of stereo matching, overcome the bottleneck problem of data scarcity, and enhance the three-dimensional scene perception capability; it can also be applied to scene three-dimensional reconstruction. The method proposed in this application can be used for three-dimensional reconstruction. For example, the laser point cloud input is modified into thickened point cloud and color data as new input, which can simultaneously utilize three-dimensional data and target color texture information, thereby improving the three-dimensional reconstruction effect.
[0152] In summary, the embodiment of the present application uses alignment verification of simulation-real machine data, massive data synthesis, and a network architecture that is easy to learn features. The point cloud generation model trained can have good simulation-real machine migration capabilities, and can effectively generate point clouds when real machine data is used as input, thereby enhancing the integrity and density of the point cloud.
[0153] The embodiment of the present application provides a point cloud generation method, wherein a point cloud generation device can obtain data to be processed; a point cloud generation model is used to extract features from the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by collecting data from a calibration scene using a binocular camera and a laser radar; the first internal and external parameters characterize the use of The internal and external parameters of the binocular camera and the lidar are calculated using the first data; the second data is obtained by simulating data acquisition of the binocular camera and the lidar in a calibration scene using a simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar calculated using the second data; the point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain optimized disparity data; depth map data is generated based on the optimized disparity data, and a point cloud is generated based on the depth map data.
[0154] Based on the above embodiment, in another embodiment of the present application, a point cloud generating device is provided, such as Figure 15 As shown, the point cloud generating device 1 may include a first generating unit 11 , a second generating unit 12 and a training unit 13 .
[0155] The first generation unit 11 can be used to use the point cloud generation model to extract features of the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training the initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data and the second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data of a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; and the point cloud generation model is used to extract multi-dimensional disparity features of the feature data to be processed to obtain disparity feature data; and the point cloud generation model is used to perform disparity feature optimization processing on the disparity feature data to obtain optimized disparity data.
[0156] The second generating unit 12 may be configured to generate depth map data according to the optimized disparity data, and generate a point cloud according to the depth map data.
[0157] In some embodiments of the present application, the point cloud generation model may include a global feature extractor and a hybrid convolution feature extractor; the first generation unit 11 can also be used to use the global feature extractor to perform convolution downsampling processing on the feature data to be processed to obtain the downsampled disparity feature, and perform global feature extraction processing based on the downsampled disparity feature and the position coding information corresponding to the data to be processed to obtain the first disparity feature; and use the hybrid convolution feature extractor to perform multi-scale local disparity feature extraction processing on the feature data to be processed to obtain multi-scale local disparity feature information, and perform hybrid convolution feature extraction processing based on the multi-scale local disparity feature information to obtain the second disparity feature; and determine the disparity feature data based on the first disparity feature and the second disparity feature.
[0158] In some embodiments of the present application, the point cloud generation model may include an optimization iterator; the first generation unit 11 can also be used to use the optimization iterator to fuse the image features and disparity feature data of the image data in the processed data to obtain fused features; and determine the disparity correction value based on the first feature information at different scales in the fused features; and optimize the initial disparity value based on the disparity correction value to obtain optimized disparity data.
[0159] In some embodiments of the present application, the first generation unit 11 can also be used to perform three-dimensional convolution processing, spatial convolution processing, and disparity convolution processing based on the local disparity feature information at the target scale in the multi-scale local disparity feature information to obtain third feature information; and use attention weights to fuse the multi-scale image features of the image data to obtain fourth feature information; and determine the second disparity feature based on the third feature information and the fourth feature information.
[0160] In some embodiments of the present application, the first generation unit 11 can also be used to perform feature downsampling of multiple resolutions on the first feature information to obtain multiple second feature information; and to fuse the multiple second feature information to obtain fused feature information; and to determine the disparity correction value based on the fused feature information.
[0161] In some embodiments of the present application, the point cloud generation model may include a monocular depth estimation network, a multi-scale feature extraction network, and a point cloud feature extraction network; the first generation unit 11 can also be used to, when the data to be processed includes image data and radar data, use the monocular depth estimation network to extract monocular depth feature information of the image data, and use the multi-scale feature extraction network to extract multi-scale depth feature information of the image data; and use the point cloud feature extraction network to extract point cloud features of the radar data; and determine the monocular depth feature information, multi-scale depth feature information, and point cloud features as feature data to be processed; and when the data to be processed includes image data, use the monocular depth estimation network to extract monocular depth feature information of the image data, use the multi-scale feature extraction network to extract multi-scale depth feature information of the image data, and determine the monocular depth feature information and multi-scale depth feature information as feature data to be processed.
[0162] The training unit 13 can be used to obtain first scene data determined by three-dimensional modeling; and obtain second scene data; wherein the second scene data is scene data obtained by reconstructing a simulated scene based on the first image information; and synthesize sensor data of the first scene data and the second scene data based on a verified simulator and pre-configured parameters to obtain synthesized image data and a synthesized point cloud; wherein the pre-configured parameters include simulation configuration parameters of the camera and radar, and environmental parameters; and determine a training data set based on the synthesized image data and the synthesized point cloud; and use the training data set to train the initial model to obtain a point cloud generation model.
[0163] In the embodiments of the present application, further, Figure 16 Schematic diagram of the structure of the point cloud generation device proposed in this embodiment Figure 2 ,like Figure 16 As shown, the point cloud generation device 1 proposed in the embodiment of the present application may also include a processor 14 and a memory 15 storing executable instructions of the processor 14; further, the point cloud generation device 1 may also include a communication interface 16, and a bus 17 for connecting the processor 14, the memory 15 and the communication interface 16.
[0164] In the embodiments of the present application, the processor 14 may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that for different devices, the electronic device used to implement the functions of the processor may also be other, and this embodiment of the present application does not specifically limit this. The point cloud generation device 1 may also include a memory 15, which may be connected to the processor 14. The memory 15 is used to store executable program code, which includes computer operating instructions. The memory 15 may include high-speed RAM memory or non-volatile memory, such as at least two disk drives.
[0165] In the embodiment of the present application, the bus 17 is used to connect the communication interface 16, the processor 14 and the memory 15, and to facilitate mutual communication between these devices.
[0166] In the embodiment of the present application, the memory 15 is used to store instructions and data.
[0167] Furthermore, in an embodiment of the present application, the above-mentioned processor 14 is used to use a point cloud generation model to extract features from the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training the initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by using a binocular camera and a lidar to collect data from a calibration scene; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by simulating the data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; the point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain optimized disparity data; depth map data is generated according to the optimized disparity data, and a point cloud is generated according to the depth map data.
[0168] In practical applications, the memory 15 may be a volatile memory, such as a random-access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 14.
[0169] In addition, the functional modules in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or software functional modules.
[0170] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0171] An embodiment of the present application provides a point cloud generation device for extracting features from data to be processed using a point cloud generation model to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on first data, first internal and external parameters, second data, and second internal and external parameters; the first data is obtained by collecting data from a calibration scene using a binocular camera and a lidar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by simulating data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; the point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data; the point cloud generation model is used to optimize multi-scale disparity features of the disparity feature data to obtain optimized disparity data; depth map data is generated based on the optimized disparity data, and a point cloud is generated based on the depth map data.
[0172] It can be seen that the present application can pre-calibrate the simulator to be verified, including the first data obtained by collecting data from a real calibration scene using a binocular camera and a laser radar, the first internal and external parameters of the binocular camera and the laser radar, the second data obtained by simulating the data collection scene of the binocular camera and the laser radar using the simulator to be verified, and the internal and external parameters of the binocular camera and the laser radar calculated using the second data to complete the verification, and whether the simulator to be verified has a good simulation effect, that is, whether the simulator to be verified can be used to obtain data that is basically the same as the real calibration scene, so as to complete the verification based on the calibration results. The simulator generated data used for the test is used to train the initial model to obtain a point cloud generation model, which can improve the generalization ability of the point cloud generation model. Then, when generating a point cloud, the point cloud generation device can use the constructed point cloud generation model to first extract features of the data to be processed, and then extract multi-dimensional disparity features of the feature data to be processed, so as to obtain disparity features of different dimensions. Then, the disparity features of different dimensions will be optimized at different scales to further optimize the disparity features, thereby generating depth map data based on the optimized disparity data, and generating a point cloud based on the depth map data, which can effectively improve the integrity and density of the generated point cloud.
[0173] Specifically, the program instructions corresponding to a point cloud generation method in this embodiment can be stored on a storage medium such as a CD, a hard disk, or a USB flash drive. When the program instructions corresponding to a point cloud generation method in the storage medium are read or executed by a point cloud generation device, the following steps are included:
[0174] A point cloud generation model is used to extract features from the data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by collecting data of a calibration scene using a binocular camera and a laser radar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained using the first data; the second data is obtained by simulating data collection of the binocular camera and the laser radar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the laser radar obtained using the second data;
[0175] The point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data;
[0176] The point cloud generation model is used to optimize the multi-scale disparity features of the disparity feature data to obtain the optimized disparity data;
[0177] Depth map data is generated based on the optimized disparity data, and a point cloud is generated based on the depth map data.
[0178] In an embodiment of the present application, an electronic device is provided. The electronic device may include a point cloud generating device, which may be used to execute the aforementioned point cloud generating method; the electronic device may be used to construct map data based on the point cloud generated by the point cloud generating device; and perform corresponding operations based on the map data.
[0179] For example, the electronic device can be a robot, and the point cloud generation device can be deployed in the robot, so that the robot can use the point cloud to build a map of the current environment, and plan the optimal path according to the map, or perform operations such as obstacle avoidance according to the map; the electronic device can also be a vehicle computer in a vehicle, and the point cloud generation device can be deployed in the vehicle computer, so that the vehicle computer can use the point cloud generated by the point cloud generation device to build a map to achieve high-precision positioning and navigation of the vehicle.
[0180] In some embodiments of the present application, the electronic device may further include a binocular camera and a laser radar. The binocular camera may be used to acquire image data, and the laser radar may be used to acquire radar data.
[0181] Exemplarily, the electronic device may be a sweeping robot including a binocular camera, a laser radar, and a point cloud generating device.
[0182] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage) containing computer-usable program code.
[0183] The present application is described with reference to the implementation flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the flowchart. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0184] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which is implemented in the implementation flow diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process described in the flowchart. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0186] The above embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Any equivalent substitution or modification made by those skilled in the art based on the present invention is within the protection scope of the present invention.
Claims
1. A point cloud generation method, characterized in that: The method comprises: A point cloud generation model is used to extract features from data to be processed to obtain feature data to be processed; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by collecting data of a calibration scene using a binocular camera and a lidar; the first internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the first data; the second data is obtained by simulating data collection of the binocular camera and the lidar in the calibration scene using the simulator to be verified; the second internal and external parameters represent the internal and external parameters of the binocular camera and the lidar obtained using the second data; Extracting multi-dimensional disparity features from the feature data to be processed using the point cloud generation model to obtain disparity feature data; Optimizing multi-scale disparity features of the disparity feature data using the point cloud generation model to obtain optimized disparity data; Depth map data is generated according to the optimized disparity data, and a point cloud is generated according to the depth map data.
2. The point cloud generation method according to claim 1, characterized in that: The point cloud generation model includes a global feature extractor and a hybrid convolution feature extractor; the point cloud generation model is used to extract multi-dimensional disparity features from the feature data to be processed to obtain disparity feature data, including: Performing convolution downsampling processing on the feature data to be processed using the global feature extractor to obtain a downsampled disparity feature, and performing global feature extraction processing based on the downsampled disparity feature and position encoding information corresponding to the data to be processed to obtain a first disparity feature; Performing multi-scale local disparity feature extraction processing on the feature data to be processed using the hybrid convolution feature extractor to obtain multi-scale local disparity feature information, and performing hybrid convolution feature extraction processing based on the multi-scale local disparity feature information to obtain a second disparity feature; The disparity feature data is determined according to the first disparity feature and the second disparity feature.
3. The point cloud generation method according to claim 2, characterized in that: The point cloud generation model includes an optimization iterator; and the optimization processing of multi-scale disparity features of the disparity feature data using the point cloud generation model to obtain optimized disparity data includes: Using the optimization iterator, the image features of the image data in the data to be processed and the disparity feature data are fused to obtain fused features; determining a disparity correction value according to first feature information at different scales in the fused features; The initial parallax value is optimized according to the parallax correction value to obtain the optimized parallax data.
4. The point cloud generation method according to claim 3, characterized in that: The performing a hybrid convolution feature extraction process according to the multi-scale local disparity feature information to obtain a second disparity feature includes: performing three-dimensional convolution processing, spatial convolution processing, and parallax convolution processing on the local disparity feature information at the target scale in the multi-scale local disparity feature information to obtain third feature information; fusing the multi-scale image features of the image data using attention weights to obtain fourth feature information; The second disparity feature is determined based on the third feature information and the fourth feature information.
5. The point cloud generation method according to claim 4, characterized in that: The determining of the disparity correction value according to the first feature information at different scales in the fused features includes: Performing feature downsampling of multiple resolutions on the first feature information to obtain multiple second feature information; fusing the plurality of second feature information to obtain fused feature information; The disparity correction value is determined according to the fused feature information.
6. The point cloud generation method according to any one of claims 1 to 5, characterized in that: The point cloud generation model includes a monocular depth estimation network, a multi-scale feature extraction network and a point cloud feature extraction network; The point cloud generation model is used to extract features from the data to be processed to obtain feature data to be processed, including: When the data to be processed includes image data and radar data, the monocular depth feature information of the image data is extracted using the monocular depth estimation network, and the multi-scale depth feature information of the image data is extracted using the multi-scale feature extraction network; extracting point cloud features of the radar data using the point cloud feature extraction network; Determining the monocular depth feature information, the multi-scale depth feature information, and the point cloud feature as the feature data to be processed; In a case where the data to be processed includes image data, the monocular depth feature information of the image data is extracted using the monocular depth estimation network, and the multi-scale depth feature information of the image data is extracted using the multi-scale feature extraction network, and the monocular depth feature information and the multi-scale depth feature information are determined as the feature data to be processed.
7. The point cloud generation method according to any one of claims 1 to 5, characterized in that: The method further comprises: Acquiring first scene data determined by three-dimensional modeling; Acquire second scene data; wherein the second scene data is scene data obtained by reconstructing a simulated scene based on the first image information; Performing sensor data synthesis on the first scene data and the second scene data based on the verified simulator and pre-configured parameters to obtain synthesized image data and a synthesized point cloud; wherein the pre-configured parameters include simulation configuration parameters of the camera and the radar, and environmental parameters; determining the training dataset based on the synthesized image data and the synthesized point cloud; The initial model is trained using the training data set to obtain the point cloud generation model.
8. A point cloud generating device, characterized in that: The point cloud generating device includes a first generating unit and a second generating unit; The first generating unit is configured to extract features from the data to be processed using a point cloud generation model to obtain feature data to be processed; and to extract multi-dimensional disparity features from the feature data to be processed using the point cloud generation model to obtain disparity feature data; and to optimize the disparity features of the disparity feature data using the point cloud generation model to obtain optimized disparity data; wherein the data to be processed includes at least image data; the point cloud generation model is obtained by training an initial model using a training data set generated by a verified simulator; the verified simulator is obtained by verifying the simulator to be verified based on the first data, the first internal and external parameters, the second data, and the second internal and external parameters; the first data is obtained by collecting data from a calibration scene using a binocular camera and a laser radar; the first internal and external parameters characterize the internal and external parameters of the binocular camera and the laser radar obtained using the first data; the second data is obtained by simulating data collection of the binocular camera and the laser radar in the calibration scene using the simulator to be verified; the second internal and external parameters characterize the internal and external parameters of the binocular camera and the laser radar obtained using the second data; The second generating unit is configured to generate depth map data according to the optimized disparity data, and generate a point cloud according to the depth map data.
9. A point cloud generating device, characterized in that: The point cloud generation device includes a processor and a memory storing processor-executable instructions; when the instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The electronic device comprises the point cloud generating device according to claim 9; The electronic device is used to construct map data based on the point cloud generated by the point cloud generating device; And performing corresponding operations according to the map data.
Citation Information
Patent Citations
Visual positioning method and device and computer readable medium
CN110568447A